A six-task experiment inside FrontierSWE v2 kept Claude Opus 5 working for an average 16.50 hours under the Proximus harness, against 5.34 hours in Claude Code. GPT-5.6 averaged 10.41 hours under Proximus and 1.01 hours in Codex.
Scores rose on average too, although the score gains are not statistically resolved across only six tasks. The clearer finding is behavioral. A nominal 20-hour budget does not mean the agent will use 20 hours. The surrounding software decides whether the model remembers its progress, knows the time and can preserve a good answer before attempting another one.
For buyers evaluating difficult engineering work, the procurement unit is therefore not the model alone. It is the model, harness and time-control system.
The benchmark tests persistence as well as code
FrontierSWE v2 contains 34 implementation, optimization, scientific-computing, visual-reasoning and AI-research tasks. Each model-task pairing gets up to 20 hours and five trials. Models run at maximum reasoning effort. The headline result is the mean of partial-credit scores across tasks.
Claude Fable 5.1 scored 56.29 percent. GPT-5.6 scored 32.2 percent, a gap of 24.09 percentage points. The Fable row is not pure Fable throughout: the benchmark used Opus 5 as a fallback on tasks blocked by content filters. Nor can the scores be compared directly with FrontierSWE v1, which used a different task set and method.
The six-task harness swap asks a narrower and more useful question. It holds the model family fixed and changes the agent shell. Across the six published task means, Proximus raised Opus 5 from 39.12 to 50.98 percent, an 11.87-point gain. It raised GPT-5.6 from 27.37 to 32.17 percent, a 4.80-point gain. The runtime changes were larger: 5.34 to 16.50 hours for Opus, and 1.01 to 10.41 hours for GPT.
The task rows prevent a cleaner sales pitch. Opus scored higher with Proximus on all six, but three gains were 3.4 points or less. GPT scored 0.9 points lower on Astronomy and 0.3 points lower on SPICE. Its gains concentrated in Remotion, Snooker and music diarization. Proximus used more time on every task for both models.
A paired calculation using the six rounded task means puts the uncertainty in view. The 95 percent t interval for the Opus score difference runs from negative 1.16 to positive 24.89 percentage points. For GPT it runs from negative 1.77 to positive 11.37 points. Both cross zero. The corresponding time intervals remain positive: 9.20 to 13.12 additional hours for Opus and 3.72 to 15.07 for GPT.
Those are cross-task estimates, not official trial-level confidence intervals. FrontierSWE does not publish the native-harness trial scores needed for a fuller comparison. Six selected tasks are also too few to establish a general productivity effect. The evidence supports a claim about time use and a directional score result, not proof that Proximus makes production teams more productive.
More time is not productivity by itself. Proximus made Opus use 3.09 times as much clock for an 11.87-point mean gain. It made GPT use 10.30 times as much for a 4.80-point gain. The pilot does not publish comparable token or dollar totals for the native conditions. That trade can be rational on a valuable research problem and wasteful on a routine patch. Internal task economics still decide whether the extra iteration pays.
Proximus changes the stopping conditions
The benchmark authors describe Proximus as a minimal extension of mini-swe-agent. When context fills, the same model summarizes the trajectory. A persistent PROGRESS.md remains in the workspace to hold details that compaction may lose. The harness provides vision for inspecting plots, frames and rendered output.
Its submission tool targets premature stopping. A model can record the current clean workspace as a candidate, receive an explicit remaining-time cue and continue working. Later experiments cannot erase the saved candidate. A second immediate submission ends the run; any other action resumes it.
That bundle addresses two distinct failures: forgetting prior work and protecting a working checkpoint. The experiment does not separate the contribution of compaction, the progress file, vision, time reminders or submission semantics. Proximus was also designed by the benchmark maker for these tasks. A feature-level causal claim would require ablations.
Reproduction remains incomplete. The public repository exposes task material and verifier code, but calls itself a work in progress and says public Docker images and a self-contained runner will follow. It does not contain the Proximus implementation or a trial-level export of the native-harness pilot.
The scores do track concrete artifacts. Astronomy grades localization completeness, geometric accuracy, image registration, mosaic fidelity and runtime. Remotion grades visual and audio closeness on hidden inputs. TORCS grades driving on held-out tracks, although one performance anchor comes from an internal reference lap. Partial credit can show useful progress. It cannot show maintainability, review burden or whether a team would merge the result.
Epoch AI’s methodology page is an external summary, not a reproduction. Epoch says it sources the leaderboard result and data export from FrontierSWE. Neither page links an independent rerun of the harness ranking.
Buy the system that does the work
Engineering leaders should require harness disclosures alongside model names. A credible evaluation freezes both versions, reports time and cost distributions, records what survives compaction, and shows task-level outcomes rather than one mean. Buyers should rerun representative internal jobs with their own repositories, tests and review standards.
Investors should read the 56.29-to-32.2 model gap within the same qualification. It measures models inside Proximus. It does not establish the same gap between off-the-shelf Claude Code and Codex deployments.
FrontierSWE’s most useful result is that the shell changes agent behavior before model intelligence settles the task. A 20-hour limit is only a ceiling. Without memory, time awareness and a safe checkpoint, the agent may hand back the change after one hour.
Discussion
Sign in to join the discussion.
No comments yet. Be the first to share your thoughts.