Skip to content

SWE-bench and friends, without the broken-environment tax

Promigence runs SWE-bench-class benchmarks on environments verified before the instance starts and reproducible from a hash afterwards. Published work has found roughly a third of the hardest instances broken at the environment level, a tax paid out of every score computed on them.

Who this is for: Teams publishing or consuming SWE-bench, SWE-rebench, Terminal-Bench, SetupBench and similar results.

39 ms

warm fork on the box

177 ms

at 200 concurrent, all served

207 ms

end to end, through the CLI

Three things that are true of almost every team doing this

01

A third of the hardest instances do not set up

Breakage in public task sets is concentrated in exactly the instances that separate good agents from great ones, so scores computed across them are not wrong by a random amount.

02

Nobody can reproduce your run

A benchmark defined by a repo and a script measures the repo as much as the agent. Rebuild it six months later and you are comparing two different things.

03

The harness is fine; the fleet is the bottleneck

Most harnesses are fine. What fails at 500 instances is image pulls, concurrency limits and cold setup repeated per instance.

With Promigence underneath

  • One verified snapshot per instance family

    Built once, verified once, forked per instance. Instances that cannot build show up as a list at build time, not as mysterious agent failures.

  • A hash you can put in the paper

    snap_… identifies a bit-identical environment. Anyone re-running against that hash gets what you had.

  • Drop-in behind your existing harness

    Adapters for common eval harnesses (pre-release), plus a compatibility SDK class so existing sandbox call sites keep working.

Point an existing harness at Promigence
TypeScript
// Most existing sandbox call sites keep working through the compatibility class.
import { Sandbox } from "promigence";

const sandbox = await Sandbox.create("swebench-django");   // a snapshot you built
await sandbox.commands.run("git apply /tmp/model_patch.diff");
const result = await sandbox.commands.run("pytest -q tests/");
await sandbox.kill();               // ← the episode ends here; billing stops
Language: typescript

Full quickstart

Run your own workload on it, free

Send a repo and the command you run against it, whatever that is: an eval suite, an RL rollout, a CI job, a queue of coding tasks. We build the environment once, run it a thousand times, and send back the timings, the failures and an exact price. Free, once, on your real workload.

  • 1,000 runs of your own command, on your own repo
  • What each one cost in wall clock, and anything that failed
  • Whether a failure was your code or the environment
  • An exact price for your real volume

A person replies, usually the same day. Or write to support@promigence.ai.

Not ready to hand over a repo? Read the quickstart or check the numbers first.

Join the waitlist

Promigence is in private beta. Leave your work email and we will send an invite as places open.