forge + Plori Router on 50 SWE-bench Verified tasks

On a committed 50-task subset, forge + the Plori Router resolved 26 of 50 tasks. The run cost $28.54 and finished in 2h 12m at concurrency 8. Every attempted task and every failure remains below.

SWE-bench Verified · 2026-08-09 UTC · swebench 4.1.0 · single run · control-plane v135

Every attempted task

Resolved Scored, unresolved Not scored Setup affected
  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10
  11. 11
  12. 12
  13. 13
  14. 14
  15. 15
  16. 16
  17. 17
  18. 18
  19. 19
  20. 20
  21. 21
  22. 22
  23. 23
  24. 24
  25. 25
  26. 26
  27. 27
  28. 28
  29. 29
  30. 30
  31. 31
  32. 32
  33. 33
  34. 34
  35. 35
  36. 36
  37. 37
  38. 38
  39. 39
  40. 40
  41. 41
  42. 42
  43. 43
  44. 44
  45. 45
  46. 46
  47. 47
  48. 48
  49. 49
  50. 50

Ordered by committed task ID. The four warning marks keep their official outcomes and identify the run-specific setup defect described below.

Resolved
26/50
Total cost
$28.54
Suite wall-clock
2h12m

Where the money and time went

The model and the computer are separate line items. Concurrency changes elapsed time, not the amount of work performed.

Cost composition

$28.54
72.1%
27.9%
Model
$20.58
Infrastructure
$7.96

Infrastructure cost is derived from measured wall time at the public run-time rate. It is not read from this account's charged credits, which would be lower while included run time remains.

Time compression

7.2× effective
Summed task time15h 55m
Suite wall-clock2h 12m

15h 55m of agent work finished in 2h 12m with 8 lanes. This is the runtime result: the model choice cannot explain suite-level parallelism.

What the Router actually served

The run pinned the Plori Router sentinel, not a concrete checkpoint. Publishing the call distribution answers which models really ran.

Resolved modelCallsShare
deepseek/deepseek-v4-flash-0731
2,01686.8%
openai/gpt-5.6-luna
2249.7%
moonshotai/kimi-k3
823.5%

Per-task evidence

Every attempted instance stays in the denominator. A missing patch, an official unresolved verdict and a harness error remain different outcomes.

Showing 50 of 50 attempted tasks

Swipe sideways for cost, wall time, models and failures.

Per-task SWE-bench result, model and infrastructure cost, tokens, wall time, and failure reason
TaskOutcomeTokensModelInfraTotalWallModels servedFailure
astropy__astropy-13236Not scored166,401$0.1917$0.1744$0.366021m55sdeepseek-v4-flash-0731run error
astropy__astropy-13398Resolved1,135,701$0.3750$0.1050$0.480013m36sdeepseek-v4-flash-0731None
astropy__astropy-13453Not scored1,059,411$0.4250$0.0902$0.515211m49sdeepseek-v4-flash-0731patch empty
django__django-10554Not scored118,807$0.1417$0.0916$0.233311m59sdeepseek-v4-flash-0731patch empty
django__django-10880Resolved253,684$0.2083$0.0994$0.307812m56sdeepseek-v4-flash-0731None
django__django-11066Resolved260,640$0.6083$0.1463$0.754718m34skimi-k3, deepseek-v4-flash-0731None
django__django-11179Not scoredSetup affected119,087$0.1583$0.2499$0.408330m00sdeepseek-v4-flash-0731, gpt-5.6-lunatimeout · salvaged
django__django-11490Resolved1,030,464$0.4917$0.1971$0.688824m39sdeepseek-v4-flash-0731None
django__django-11964Resolved726,643$0.4333$0.1557$0.589119m41sdeepseek-v4-flash-0731None
django__django-11999Resolved961,189$0.5583$0.2154$0.773726m51sdeepseek-v4-flash-0731run error · salvaged
django__django-12155Resolved758,286$0.4000$0.0763$0.47639m09sdeepseek-v4-flash-0731None
django__django-12193Resolved607,447$0.4000$0.0946$0.494611m21sdeepseek-v4-flash-0731None
django__django-12209Resolved364,141$0.2500$0.1316$0.381616m47sdeepseek-v4-flash-0731None
django__django-12273Not scored407,615$0.2750$0.2500$0.525030m00sgpt-5.6-luna, deepseek-v4-flash-0731timeout
django__django-12663Not scoredSetup affected640,370$0.4000$0.2499$0.649930m00sdeepseek-v4-flash-0731timeout
django__django-12754Resolved925,673$0.4833$0.2499$0.733230m59sdeepseek-v4-flash-0731timeout · salvaged
django__django-13028Not scored86,712$0.1250$0.0514$0.17646m10sdeepseek-v4-flash-0731, gpt-5.6-lunarun error
django__django-13112Not scored1,059,201$0.6250$0.1519$0.776918m14sdeepseek-v4-flash-0731patch empty
django__django-13297Not scored1,050,347$0.5083$0.1863$0.694722m22sdeepseek-v4-flash-0731patch empty
django__django-13344Not scored255,937$0.2667$0.1502$0.416818m01sdeepseek-v4-flash-0731patch empty
django__django-13820Resolved414,956$0.2417$0.0690$0.31078m17sgpt-5.6-luna, deepseek-v4-flash-0731None
django__django-14034Resolved445,583$0.3333$0.0986$0.432012m50sdeepseek-v4-flash-0731None
django__django-14155Not scored134,759$0.1833$0.1036$0.287012m26sdeepseek-v4-flash-0731patch empty
django__django-14787Not scored36,557$0.0750$0.2499$0.324930m59sgpt-5.6-luna, deepseek-v4-flash-0731timeout
django__django-16255Not scored224,136$0.2417$0.2500$0.491630m00sdeepseek-v4-flash-0731timeout
django__django-16485Resolved670,767$0.3917$0.0847$0.476410m10sgpt-5.6-luna, deepseek-v4-flash-0731None
django__django-16661Resolved185,091$0.1917$0.1022$0.293812m15sdeepseek-v4-flash-0731None
django__django-9296Not scoredSetup affected689,663$0.4833$0.2410$0.724429m55sdeepseek-v4-flash-0731run error
matplotlib__matplotlib-20676Not scored1,592,237$0.7417$0.2400$0.981729m48sdeepseek-v4-flash-0731patch empty
matplotlib__matplotlib-26466Scored, unresolved955,773$0.4500$0.2500$0.700030m00sdeepseek-v4-flash-0731, gpt-5.6-lunatimeout · salvaged
psf__requests-2931Resolved786,059$0.4500$0.0707$0.52078m29sdeepseek-v4-flash-0731None
pydata__xarray-4695Resolved879,817$0.4083$0.1016$0.510012m12sdeepseek-v4-flash-0731None
pydata__xarray-6721Scored, unresolved1,054,441$0.5917$0.0789$0.67059m28sdeepseek-v4-flash-0731None
scikit-learn__scikit-learn-25232Resolved1,080,897$0.5167$0.2298$0.746528m34sdeepseek-v4-flash-0731None
scikit-learn__scikit-learn-25747Resolved1,094,949$0.2667$0.1077$0.374413m56sgpt-5.6-luna, deepseek-v4-flash-0731None
scikit-learn__scikit-learn-25931Not scored1,059,723$0.5500$0.1653$0.715320m50sdeepseek-v4-flash-0731patch empty
scikit-learn__scikit-learn-25973Resolved1,210,146$0.5417$0.1395$0.681217m45sdeepseek-v4-flash-0731None
sphinx-doc__sphinx-10323Resolved1,045,758$0.4333$0.1643$0.597720m43sdeepseek-v4-flash-0731None
sphinx-doc__sphinx-10673Resolved1,081,144$0.3000$0.1123$0.412313m28sdeepseek-v4-flash-0731None
sphinx-doc__sphinx-11445Not scored766,221$0.4250$0.2499$0.674930m59sdeepseek-v4-flash-0731timeout
sphinx-doc__sphinx-11510Not scoredSetup affected1,157,995$0.5250$0.1718$0.696821m37sdeepseek-v4-flash-0731patch empty
sphinx-doc__sphinx-9698Resolved278,085$0.1833$0.0956$0.279011m29sgpt-5.6-luna, deepseek-v4-flash-0731None
sympy__sympy-12096Not scored1,044,726$0.5583$0.2057$0.764125m41sdeepseek-v4-flash-0731patch empty
sympy__sympy-14531Not scored432,290$0.8667$0.2499$1.116630m59skimi-k3, deepseek-v4-flash-0731timeout
sympy__sympy-14711Resolved797,201$0.4083$0.1261$0.534515m08sdeepseek-v4-flash-0731None
sympy__sympy-17655Resolved1,011,968$0.4667$0.1588$0.625519m04sdeepseek-v4-flash-0731None
sympy__sympy-19783Resolved1,067,385$0.5667$0.2256$0.792327m04sdeepseek-v4-flash-0731None
sympy__sympy-21596Not scored1,083,424$0.5250$0.1741$0.699121m53sdeepseek-v4-flash-0731patch empty
sympy__sympy-21930Scored, unresolved1,044,596$0.9083$0.1630$1.071320m34skimi-k3, deepseek-v4-flash-0731None
sympy__sympy-23413Resolved1,018,609$0.4333$0.1601$0.593519m13sgpt-5.6-lunaNone

Methodology

The scoring harness and the running agent are separate. plori produced patches; the official SWE-bench harness evaluated them.

System
forge + Plori Router
Router
plori-auto · paid candidate pool
Platform
control-plane v135
Suite
SWE-bench Verified · subset-50-seed20260804
Official harness
swebench 4.1.0
Execution
patch-only · 25 min server deadline
Concurrency
8 tasks in flight
Run shape
single run · 1 repeat
First durable event
p50 6.1s · p90 9.7s · n=50

The run started with six pre-warmed agent nodes and held eight nodes for eight concurrent tasks, one task per node.

Each task started in a fresh agent. The agent received the issue statement and base commit, edited the repository in its plori computer, and exported a patch. The committed driver then joined exact per-task usage and server timestamps with the official harness verdict. Tasks with no patch remain attempted but unscored.

Model cost comes from model-only credits at the 120-credits-per-USD peg. Infrastructure cost is wall minutes × 1 credit per minute ÷ 120, independent of any included run-time allowance. The two figures are added only after they have been published separately.

Limitations

Single run.

Single run. Treat 52% as low-50s rather than an exact capability estimate; no repeated-run variance is available for this configuration.

Four tasks were affected by an incomplete checkout

One retained patch proves the repository checkout stopped before its alphabetical tail; three unretrievable oversized patches share the same signature. The published 52% is the result this run produced, not a measurement of the agent's ceiling on those four tasks.

One submitted patch errored in the official harness

django__django-11179 was submitted but not scored. Harness errors are reported separately from unresolved patches.

A 50-task subset, not the full leaderboard

The subset is stratified by the dataset's difficulty labels and fixed by a committed seed. It is useful for this product's repeatable measurement, not a substitute for SWE-bench Verified's full 500-instance leaderboard.

No competitor comparison

This page publishes plori's own accuracy, cost and time under one stated method. It makes no claim about another agent's cost or capability.

Inspect the artifact and reproduce the run

The page is a static projection of one committed result. It makes no live API request, so a published number cannot change underneath its methodology.