ForgeOSby Medina19
Sign in Create account

Capability

What ForgeOS can do, measured — and what the measurements refuse to claim.

In short
Every number here is read from a record in the repository, and every record names the receipt it came from and the day it was measured.
A number whose receipt is not on disk is refused before it can be recorded, so nothing on this page is remembered rather than measured.
Where a measurement is not good enough to publish, the record says so — and so does this page, in the benchmark below.

87.9%

of briefs produced a verified artifact, on free models

±7pt noise band · 2026-09-03

3.0%

reported success on an artifact that does not hold up

The binding constraint — the bar refuses production above 2%

0

axe violations across every commercial page

Real browser, two widths · 2026-09-02

11

starter kinds driven to a verified artifact

Against 126 with a recipe — catalogued is not proven

The standing record

Six dimensions, each with a floor that only moves by measuring.

The floor never moves down by editing. It is kept beside the latest reading, so a good week cannot hide a bad one; and each dimension declares which direction is better, so a rising false-success rate is a regression rather than applause.

Build resolve rate

87.9% of briefs produced a verified artifact on the free model fleet. Sampled; the noise band is ±7 points, so a single run cannot move this. race-eval-b4 · 2026-09-03

False-success rate

3.0% of goals reported success with an artifact that does not hold up. This is the binding constraint: the release bar refuses production above 2%, so this number, not the pass rate, is what stands between ForgeOS and calling itself production-ready. race-eval-b4 · 2026-09-03

Build kinds proven

11 distinct starter kinds the governed corpus drives to a verified artifact. The number that only grows by building. build-eval corpus · 2026-09-04

Technologies with a recipe

126 technologies ForgeOS has a real command for, whether or not a given host has the toolchain installed. build-path coverage · 2026-09-04

Technologies catalogued

325 technologies ForgeOS can reason about and plan for. Catalogued is not proven: only 11 kinds above have been built to a verified artifact. build-path coverage · 2026-09-04

Accessibility violations

0 axe violations across every commercial page at two widths, measured in a real browser. serving-scale audit · 2026-09-02

The benchmark, as run

Three arms measured on one 50-task corpus. Not publishable as a comparison.

The protocol declares five arms and requires every declared arm to be run before the receipt may be called a comparison. Two have not been run, and the receipt says so in its own words rather than leaving those arms out of the table. One of the two was withdrawn on 2026-09-10 after its row failed to reproduce — the correction is described in the table below, on the arm it concerns.

Frontier agent, standalone

Not run — and this page used to say 100%. The receipt was rebuilt on 2026-09-10 because that row did not reproduce from the files in this repository. What had been recorded was a model answering each task directly with no agent loop, filed under the name of an agent product the protocol declares; its $0.00 and 0.792s were the grading loop's own runtime, not an agent's. Renaming an arm is not relabelling a measurement, it is substituting one. Withdrawing it also withdrew the gap-to-frontier dimension from the record above, which was this arm's rate minus ours and cannot be computed without it. Not one ForgeOS row changed: the three arms below are the same 150 rows as before.

ForgeOS, free model fleet

92% resolved · 92% hidden tests passed · 8% needed a person · $0.00 per task. The arm a new workspace runs on the day it is created.

ForgeOS, routed free-then-paid

88% resolved · 88% hidden tests passed · 12% needed a person · $0.0004 per task. Below the free fleet by policy, not defect: free candidates are exhausted and the walk stops on diminishing returns before escalating to a paid account.

ForgeOS, recipes only, no model

56% resolved · 44% needed a person · $0.00. What the build recipes achieve with no model at all — the floor a model has to beat to be worth its cost.

Codex

Not run. Declared in the protocol, never measured, never imputed. This single unmeasured arm is why the receipt is marked not publishable, and it stays that way until the arm is run.

What this is not. This is not a SWE-bench score, and ForgeOS publishes none. The corpus is ours: 50 tasks we wrote, run by us. It is offered as the record of a measurement, with its receipt, not as a leaderboard position. An honest number about our own corpus is worth more to a buyer than a flattering one about somebody else's.

How to read it

Three things a careful reader should know before trusting any of it.

Sampled numbers have noise

The resolve rate and the false-success rate are sampled, and the record carries their noise band (±7 and ±3 points). A run inside the band is the same result, and no single run is ever attributed to a single change.

The floor is not the latest

Each dimension keeps its best and its latest. The floor only rises when a measurement lands; a deploy that raised its own bar would be marking its own homework, so it reports and never writes.

Unbacked is refused, not rounded

A value whose receipt is not on disk is refused outright. There is no such thing on this page as a number somebody remembered.

What blocks production

Not the pass rate. A 3.0% false-success rate against a 2% bar is the one number that has to move for ForgeOS to call a build production-ready without a person reading it.

Run it on your own brief.

A workspace takes ten seconds, costs nothing, and comes with the free fleet measured above. The thread tells you what was verified and what was not.