SWE-bench and friends, without the broken-environment tax
Promigence runs SWE-bench-class benchmarks on environments verified before the instance starts and reproducible from a hash afterwards. Published work has found roughly a third of the hardest instances broken at the environment level, a tax paid out of every score computed on them.
Who this is for: Teams publishing or consuming SWE-bench, SWE-rebench, Terminal-Bench, SetupBench and similar results.
warm fork on the box
at 200 concurrent, all served
end to end, through the CLI
Three things that are true of almost every team doing this
A third of the hardest instances do not set up
Breakage in public task sets is concentrated in exactly the instances that separate good agents from great ones, so scores computed across them are not wrong by a random amount.
Nobody can reproduce your run
A benchmark defined by a repo and a script measures the repo as much as the agent. Rebuild it six months later and you are comparing two different things.
The harness is fine; the fleet is the bottleneck
Most harnesses are fine. What fails at 500 instances is image pulls, concurrency limits and cold setup repeated per instance.
With Promigence underneath
One verified snapshot per instance family
Built once, verified once, forked per instance. Instances that cannot build show up as a list at build time, not as mysterious agent failures.
A hash you can put in the paper
snap_…identifies a bit-identical environment. Anyone re-running against that hash gets what you had.Drop-in behind your existing harness
Adapters for common eval harnesses (pre-release), plus a compatibility SDK class so existing sandbox call sites keep working.
// Most existing sandbox call sites keep working through the compatibility class.
import { Sandbox } from "promigence";
const sandbox = await Sandbox.create("swebench-django"); // a snapshot you built
await sandbox.commands.run("git apply /tmp/model_patch.diff");
const result = await sandbox.commands.run("pytest -q tests/");
await sandbox.kill(); // ← the episode ends here; billing stops
Language: typescriptRun your own workload on it, free
Send a repo and the command you run against it, whatever that is: an eval suite, an RL rollout, a CI job, a queue of coding tasks. We build the environment once, run it a thousand times, and send back the timings, the failures and an exact price. Free, once, on your real workload.
- 1,000 runs of your own command, on your own repo
- What each one cost in wall clock, and anything that failed
- Whether a failure was your code or the environment
- An exact price for your real volume
Not ready to hand over a repo? Read the quickstart or check the numbers first.