forge + Plori Router on 50 SWE-bench Verified tasks
On a committed 50-task subset, forge + the Plori Router resolved 26 of 50 tasks. The run cost $28.54 and finished in 2h 12m at concurrency 8. Every attempted task and every failure remains below.
SWE-bench Verified · 2026-08-09 UTC · swebench 4.1.0 · single run · control-plane v135
Every attempted task
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08
- 09
- 10
- 11
- 12
- 13
- 14
- 15
- 16
- 17
- 18
- 19
- 20
- 21
- 22
- 23
- 24
- 25
- 26
- 27
- 28
- 29
- 30
- 31
- 32
- 33
- 34
- 35
- 36
- 37
- 38
- 39
- 40
- 41
- 42
- 43
- 44
- 45
- 46
- 47
- 48
- 49
- 50
Ordered by committed task ID. The four warning marks keep their official outcomes and identify the run-specific setup defect described below.
- Resolved
- 26/50
- Total cost
- $28.54
- Suite wall-clock
- 2h12m
Where the money and time went
The model and the computer are separate line items. Concurrency changes elapsed time, not the amount of work performed.
Cost composition
$28.54- Model
- $20.58
- Infrastructure
- $7.96
Infrastructure cost is derived from measured wall time at the public run-time rate. It is not read from this account's charged credits, which would be lower while included run time remains.
Time compression
7.2× effective15h 55m of agent work finished in 2h 12m with 8 lanes. This is the runtime result: the model choice cannot explain suite-level parallelism.
What the Router actually served
The run pinned the Plori Router sentinel, not a concrete checkpoint. Publishing the call distribution answers which models really ran.
Per-task evidence
Every attempted instance stays in the denominator. A missing patch, an official unresolved verdict and a harness error remain different outcomes.
Showing 50 of 50 attempted tasks
Swipe sideways for cost, wall time, models and failures.
| Task | Outcome | Tokens | Model | Infra | Total | Wall | Models served | Failure |
|---|---|---|---|---|---|---|---|---|
| astropy__astropy-13236 | Not scored | 166,401 | $0.1917 | $0.1744 | $0.3660 | 21m55s | deepseek-v4-flash-0731 | run error |
| astropy__astropy-13398 | Resolved | 1,135,701 | $0.3750 | $0.1050 | $0.4800 | 13m36s | deepseek-v4-flash-0731 | None |
| astropy__astropy-13453 | Not scored | 1,059,411 | $0.4250 | $0.0902 | $0.5152 | 11m49s | deepseek-v4-flash-0731 | patch empty |
| django__django-10554 | Not scored | 118,807 | $0.1417 | $0.0916 | $0.2333 | 11m59s | deepseek-v4-flash-0731 | patch empty |
| django__django-10880 | Resolved | 253,684 | $0.2083 | $0.0994 | $0.3078 | 12m56s | deepseek-v4-flash-0731 | None |
| django__django-11066 | Resolved | 260,640 | $0.6083 | $0.1463 | $0.7547 | 18m34s | kimi-k3, deepseek-v4-flash-0731 | None |
| django__django-11179 | Not scoredSetup affected | 119,087 | $0.1583 | $0.2499 | $0.4083 | 30m00s | deepseek-v4-flash-0731, gpt-5.6-luna | timeout · salvaged |
| django__django-11490 | Resolved | 1,030,464 | $0.4917 | $0.1971 | $0.6888 | 24m39s | deepseek-v4-flash-0731 | None |
| django__django-11964 | Resolved | 726,643 | $0.4333 | $0.1557 | $0.5891 | 19m41s | deepseek-v4-flash-0731 | None |
| django__django-11999 | Resolved | 961,189 | $0.5583 | $0.2154 | $0.7737 | 26m51s | deepseek-v4-flash-0731 | run error · salvaged |
| django__django-12155 | Resolved | 758,286 | $0.4000 | $0.0763 | $0.4763 | 9m09s | deepseek-v4-flash-0731 | None |
| django__django-12193 | Resolved | 607,447 | $0.4000 | $0.0946 | $0.4946 | 11m21s | deepseek-v4-flash-0731 | None |
| django__django-12209 | Resolved | 364,141 | $0.2500 | $0.1316 | $0.3816 | 16m47s | deepseek-v4-flash-0731 | None |
| django__django-12273 | Not scored | 407,615 | $0.2750 | $0.2500 | $0.5250 | 30m00s | gpt-5.6-luna, deepseek-v4-flash-0731 | timeout |
| django__django-12663 | Not scoredSetup affected | 640,370 | $0.4000 | $0.2499 | $0.6499 | 30m00s | deepseek-v4-flash-0731 | timeout |
| django__django-12754 | Resolved | 925,673 | $0.4833 | $0.2499 | $0.7332 | 30m59s | deepseek-v4-flash-0731 | timeout · salvaged |
| django__django-13028 | Not scored | 86,712 | $0.1250 | $0.0514 | $0.1764 | 6m10s | deepseek-v4-flash-0731, gpt-5.6-luna | run error |
| django__django-13112 | Not scored | 1,059,201 | $0.6250 | $0.1519 | $0.7769 | 18m14s | deepseek-v4-flash-0731 | patch empty |
| django__django-13297 | Not scored | 1,050,347 | $0.5083 | $0.1863 | $0.6947 | 22m22s | deepseek-v4-flash-0731 | patch empty |
| django__django-13344 | Not scored | 255,937 | $0.2667 | $0.1502 | $0.4168 | 18m01s | deepseek-v4-flash-0731 | patch empty |
| django__django-13820 | Resolved | 414,956 | $0.2417 | $0.0690 | $0.3107 | 8m17s | gpt-5.6-luna, deepseek-v4-flash-0731 | None |
| django__django-14034 | Resolved | 445,583 | $0.3333 | $0.0986 | $0.4320 | 12m50s | deepseek-v4-flash-0731 | None |
| django__django-14155 | Not scored | 134,759 | $0.1833 | $0.1036 | $0.2870 | 12m26s | deepseek-v4-flash-0731 | patch empty |
| django__django-14787 | Not scored | 36,557 | $0.0750 | $0.2499 | $0.3249 | 30m59s | gpt-5.6-luna, deepseek-v4-flash-0731 | timeout |
| django__django-16255 | Not scored | 224,136 | $0.2417 | $0.2500 | $0.4916 | 30m00s | deepseek-v4-flash-0731 | timeout |
| django__django-16485 | Resolved | 670,767 | $0.3917 | $0.0847 | $0.4764 | 10m10s | gpt-5.6-luna, deepseek-v4-flash-0731 | None |
| django__django-16661 | Resolved | 185,091 | $0.1917 | $0.1022 | $0.2938 | 12m15s | deepseek-v4-flash-0731 | None |
| django__django-9296 | Not scoredSetup affected | 689,663 | $0.4833 | $0.2410 | $0.7244 | 29m55s | deepseek-v4-flash-0731 | run error |
| matplotlib__matplotlib-20676 | Not scored | 1,592,237 | $0.7417 | $0.2400 | $0.9817 | 29m48s | deepseek-v4-flash-0731 | patch empty |
| matplotlib__matplotlib-26466 | Scored, unresolved | 955,773 | $0.4500 | $0.2500 | $0.7000 | 30m00s | deepseek-v4-flash-0731, gpt-5.6-luna | timeout · salvaged |
| psf__requests-2931 | Resolved | 786,059 | $0.4500 | $0.0707 | $0.5207 | 8m29s | deepseek-v4-flash-0731 | None |
| pydata__xarray-4695 | Resolved | 879,817 | $0.4083 | $0.1016 | $0.5100 | 12m12s | deepseek-v4-flash-0731 | None |
| pydata__xarray-6721 | Scored, unresolved | 1,054,441 | $0.5917 | $0.0789 | $0.6705 | 9m28s | deepseek-v4-flash-0731 | None |
| scikit-learn__scikit-learn-25232 | Resolved | 1,080,897 | $0.5167 | $0.2298 | $0.7465 | 28m34s | deepseek-v4-flash-0731 | None |
| scikit-learn__scikit-learn-25747 | Resolved | 1,094,949 | $0.2667 | $0.1077 | $0.3744 | 13m56s | gpt-5.6-luna, deepseek-v4-flash-0731 | None |
| scikit-learn__scikit-learn-25931 | Not scored | 1,059,723 | $0.5500 | $0.1653 | $0.7153 | 20m50s | deepseek-v4-flash-0731 | patch empty |
| scikit-learn__scikit-learn-25973 | Resolved | 1,210,146 | $0.5417 | $0.1395 | $0.6812 | 17m45s | deepseek-v4-flash-0731 | None |
| sphinx-doc__sphinx-10323 | Resolved | 1,045,758 | $0.4333 | $0.1643 | $0.5977 | 20m43s | deepseek-v4-flash-0731 | None |
| sphinx-doc__sphinx-10673 | Resolved | 1,081,144 | $0.3000 | $0.1123 | $0.4123 | 13m28s | deepseek-v4-flash-0731 | None |
| sphinx-doc__sphinx-11445 | Not scored | 766,221 | $0.4250 | $0.2499 | $0.6749 | 30m59s | deepseek-v4-flash-0731 | timeout |
| sphinx-doc__sphinx-11510 | Not scoredSetup affected | 1,157,995 | $0.5250 | $0.1718 | $0.6968 | 21m37s | deepseek-v4-flash-0731 | patch empty |
| sphinx-doc__sphinx-9698 | Resolved | 278,085 | $0.1833 | $0.0956 | $0.2790 | 11m29s | gpt-5.6-luna, deepseek-v4-flash-0731 | None |
| sympy__sympy-12096 | Not scored | 1,044,726 | $0.5583 | $0.2057 | $0.7641 | 25m41s | deepseek-v4-flash-0731 | patch empty |
| sympy__sympy-14531 | Not scored | 432,290 | $0.8667 | $0.2499 | $1.1166 | 30m59s | kimi-k3, deepseek-v4-flash-0731 | timeout |
| sympy__sympy-14711 | Resolved | 797,201 | $0.4083 | $0.1261 | $0.5345 | 15m08s | deepseek-v4-flash-0731 | None |
| sympy__sympy-17655 | Resolved | 1,011,968 | $0.4667 | $0.1588 | $0.6255 | 19m04s | deepseek-v4-flash-0731 | None |
| sympy__sympy-19783 | Resolved | 1,067,385 | $0.5667 | $0.2256 | $0.7923 | 27m04s | deepseek-v4-flash-0731 | None |
| sympy__sympy-21596 | Not scored | 1,083,424 | $0.5250 | $0.1741 | $0.6991 | 21m53s | deepseek-v4-flash-0731 | patch empty |
| sympy__sympy-21930 | Scored, unresolved | 1,044,596 | $0.9083 | $0.1630 | $1.0713 | 20m34s | kimi-k3, deepseek-v4-flash-0731 | None |
| sympy__sympy-23413 | Resolved | 1,018,609 | $0.4333 | $0.1601 | $0.5935 | 19m13s | gpt-5.6-luna | None |
Methodology
The scoring harness and the running agent are separate. plori produced patches; the official SWE-bench harness evaluated them.
- System
- forge + Plori Router
- Router
- plori-auto · paid candidate pool
- Platform
- control-plane v135
- Suite
- SWE-bench Verified · subset-50-seed20260804
- Official harness
- swebench 4.1.0
- Execution
- patch-only · 25 min server deadline
- Concurrency
- 8 tasks in flight
- Run shape
- single run · 1 repeat
- First durable event
- p50 6.1s · p90 9.7s · n=50
The run started with six pre-warmed agent nodes and held eight nodes for eight concurrent tasks, one task per node.
Each task started in a fresh agent. The agent received the issue statement and base commit, edited the repository in its plori computer, and exported a patch. The committed driver then joined exact per-task usage and server timestamps with the official harness verdict. Tasks with no patch remain attempted but unscored.
Model cost comes from model-only credits at the 120-credits-per-USD peg. Infrastructure cost is wall minutes × 1 credit per minute ÷ 120, independent of any included run-time allowance. The two figures are added only after they have been published separately.
Limitations
Single run.
Single run. Treat 52% as low-50s rather than an exact capability estimate; no repeated-run variance is available for this configuration.
Four tasks were affected by an incomplete checkout
One retained patch proves the repository checkout stopped before its alphabetical tail; three unretrievable oversized patches share the same signature. The published 52% is the result this run produced, not a measurement of the agent's ceiling on those four tasks.
One submitted patch errored in the official harness
django__django-11179 was submitted but not scored. Harness errors are reported separately from unresolved patches.
A 50-task subset, not the full leaderboard
The subset is stratified by the dataset's difficulty labels and fixed by a committed seed. It is useful for this product's repeatable measurement, not a substitute for SWE-bench Verified's full 500-instance leaderboard.
No competitor comparison
This page publishes plori's own accuracy, cost and time under one stated method. It makes no claim about another agent's cost or capability.
Inspect the artifact and reproduce the run
The page is a static projection of one committed result. It makes no live API request, so a published number cannot change underneath its methodology.