Capability
What ForgeOS can do, measured — and what the measurements refuse to claim.
- In short
- Every number here is read from a record in the repository, and every record names the receipt it came from and the day it was measured.
- A number whose receipt is not on disk is refused before it can be recorded, so nothing on this page is remembered rather than measured.
- Where a measurement is not good enough to publish, the record says so — and so does this page, in the benchmark below.
87.9%
of briefs produced a verified artifact, on free models
±7pt noise band · 2026-09-03
3.0%
reported success on an artifact that does not hold up
The binding constraint — the bar refuses production above 2%
0
axe violations across every commercial page
Real browser, two widths · 2026-09-02
11
starter kinds driven to a verified artifact
Against 126 with a recipe — catalogued is not proven
The standing record
Six dimensions, each with a floor that only moves by measuring.
The floor never moves down by editing. It is kept beside the latest reading, so a good week cannot hide a bad one; and each dimension declares which direction is better, so a rising false-success rate is a regression rather than applause.
Build resolve rate
87.9% of briefs produced a verified artifact on the free model fleet. Sampled; the noise band is ±7 points, so a single run cannot move this. race-eval-b4 · 2026-09-03
False-success rate
3.0% of goals reported success with an artifact that does not hold up. This is the binding constraint: the release bar refuses production above 2%, so this number, not the pass rate, is what stands between ForgeOS and calling itself production-ready. race-eval-b4 · 2026-09-03
Build kinds proven
11 distinct starter kinds the governed corpus drives to a verified artifact. The number that only grows by building. build-eval corpus · 2026-09-04
Technologies with a recipe
126 technologies ForgeOS has a real command for, whether or not a given host has the toolchain installed. build-path coverage · 2026-09-04
Technologies catalogued
325 technologies ForgeOS can reason about and plan for. Catalogued is not proven: only 11 kinds above have been built to a verified artifact. build-path coverage · 2026-09-04
Accessibility violations
0 axe violations across every commercial page at two widths, measured in a real browser. serving-scale audit · 2026-09-02
The benchmark, as run
Three arms measured on one 50-task corpus. Not publishable as a comparison.
The protocol declares five arms and requires every declared arm to be run before the receipt may be called a comparison. Two have not been run, and the receipt says so in its own words rather than leaving those arms out of the table. One of the two was withdrawn on 2026-09-10 after its row failed to reproduce — the correction is described in the table below, on the arm it concerns.
schema forgeos.coding-agent-benchmark/v1
tasks_seen 50 minimum_tasks_per_arm 50
publishable false
claim Not publishable: arms never run: claude_code, codex.
Frontier agent, standalone
Not run — and this page used to say 100%. The receipt was rebuilt on 2026-09-10 because that row did not reproduce from the files in this repository. What had been recorded was a model answering each task directly with no agent loop, filed under the name of an agent product the protocol declares; its $0.00 and 0.792s were the grading loop's own runtime, not an agent's. Renaming an arm is not relabelling a measurement, it is substituting one. Withdrawing it also withdrew the gap-to-frontier dimension from the record above, which was this arm's rate minus ours and cannot be computed without it. Not one ForgeOS row changed: the three arms below are the same 150 rows as before.
ForgeOS, free model fleet
92% resolved · 92% hidden tests passed · 8% needed a person · $0.00 per task. The arm a new workspace runs on the day it is created.
ForgeOS, routed free-then-paid
88% resolved · 88% hidden tests passed · 12% needed a person · $0.0004 per task. Below the free fleet by policy, not defect: free candidates are exhausted and the walk stops on diminishing returns before escalating to a paid account.
ForgeOS, recipes only, no model
56% resolved · 44% needed a person · $0.00. What the build recipes achieve with no model at all — the floor a model has to beat to be worth its cost.
Codex
Not run. Declared in the protocol, never measured, never imputed. This single unmeasured arm is why the receipt is marked not publishable, and it stays that way until the arm is run.
What this is not. This is not a SWE-bench score, and ForgeOS publishes none. The corpus is ours: 50 tasks we wrote, run by us. It is offered as the record of a measurement, with its receipt, not as a leaderboard position. An honest number about our own corpus is worth more to a buyer than a flattering one about somebody else's.
How to read it
Three things a careful reader should know before trusting any of it.
Sampled numbers have noise
The resolve rate and the false-success rate are sampled, and the record carries their noise band (±7 and ±3 points). A run inside the band is the same result, and no single run is ever attributed to a single change.
The floor is not the latest
Each dimension keeps its best and its latest. The floor only rises when a measurement lands; a deploy that raised its own bar would be marking its own homework, so it reports and never writes.
Unbacked is refused, not rounded
A value whose receipt is not on disk is refused outright. There is no such thing on this page as a number somebody remembered.
What blocks production
Not the pass rate. A 3.0% false-success rate against a 2% bar is the one number that has to move for ForgeOS to call a build production-ready without a person reading it.
Run it on your own brief.
A workspace takes ten seconds, costs nothing, and comes with the free fleet measured above. The thread tells you what was verified and what was not.