Run agent evaluations on environments you can prove were valid
Promigence verifies every environment before an agent runs in it, forks a thousand from one snapshot, and reports which failures were the environment's and which were the agent's. That distinction is the difference between an eval number you can publish and one you cannot.
Who this is for: Evaluation and platform engineers at AI agent companies, benchmark producers, and eval-as-a-service teams.
warm fork on the box
at 200 concurrent, all served
forks in six hours, zero failures
Three things that are true of almost every team doing this
You cannot say how many environments were broken
Ask a team what fraction of their environments failed to install and the answer is a shrug. Those failures are recorded as agent failures, and they concentrate in the repos hardest to set up, which are the repos the benchmark is about.
Last month's run is not comparable to this month's
Environments built from a recipe re-resolve package versions every time. When the environment drifts underneath you, a regression in your agent and a change in your environment look identical.
One run per day, because a run takes hours
Cold clone, install and first build costs 60 to 100 seconds on every episode in the suite. That is what turns an eval into an overnight job.
With Promigence underneath
Every environment ran its own tests before your agent arrived
Declare a verify command with the snapshot. If it fails, the snapshot is rejected, never reaches an episode, and is not charged for.
Each run ends with a number about the environments
Verified, environment failures, task failures, each replayable from the flight recorder. Publish it next to your result.
The same hash reproduces the same environment in March
Cite the snapshot hash in your methods section. Re-run against it later and the comparison means something.
Many runs a day instead of one
Forks start warm, so an episode begins where your setup script left off, not at an empty disk.
# 1. Build a verified snapshot, rejected if the typecheck fails
promigence snapshot create --image node:22 \
--setup "git clone https://github.com/your-org/evals /work/repo && git -C /work/repo checkout <40-char sha>" \
--setup "cd /work/repo && npm ci" \
--verify "cd /work/repo && npm run typecheck" --alias evals --wait --wait-timeout 10m
# 2. Know the bill before you spend it
promigence run quote --snapshot evals --count 20 --timeout 15m
# → quote $0.83 — 20 x 900s x $0.000046/s (tier small)
# 3. Fan out, capped
promigence run --snapshot evals --count 20 --cap 1 \
--verify-each --json -- ./run_agent.sh
# 4. The number about your environments
promigence run report <run_id>
Language: bashRun your own workload on it, free
Send a repo and the command you run against it, whatever that is: an eval suite, an RL rollout, a CI job, a queue of coding tasks. We build the environment once, run it a thousand times, and send back the timings, the failures and an exact price. Free, once, on your real workload.
- 1,000 runs of your own command, on your own repo
- What each one cost in wall clock, and anything that failed
- Whether a failure was your code or the environment
- An exact price for your real volume
Not ready to hand over a repo? Read the quickstart or check the numbers first.