Skip to content

Run agent evaluations on environments you can prove were valid

Promigence verifies every environment before an agent runs in it, forks a thousand from one snapshot, and reports which failures were the environment's and which were the agent's. That distinction is the difference between an eval number you can publish and one you cannot.

Who this is for: Evaluation and platform engineers at AI agent companies, benchmark producers, and eval-as-a-service teams.

39 ms

warm fork on the box

177 ms

at 200 concurrent, all served

108,929

forks in six hours, zero failures

Three things that are true of almost every team doing this

01

You cannot say how many environments were broken

Ask a team what fraction of their environments failed to install and the answer is a shrug. Those failures are recorded as agent failures, and they concentrate in the repos hardest to set up, which are the repos the benchmark is about.

02

Last month's run is not comparable to this month's

Environments built from a recipe re-resolve package versions every time. When the environment drifts underneath you, a regression in your agent and a change in your environment look identical.

03

One run per day, because a run takes hours

Cold clone, install and first build costs 60 to 100 seconds on every episode in the suite. That is what turns an eval into an overnight job.

With Promigence underneath

  • Every environment ran its own tests before your agent arrived

    Declare a verify command with the snapshot. If it fails, the snapshot is rejected, never reaches an episode, and is not charged for.

  • Each run ends with a number about the environments

    Verified, environment failures, task failures, each replayable from the flight recorder. Publish it next to your result.

  • The same hash reproduces the same environment in March

    Cite the snapshot hash in your methods section. Re-run against it later and the comparison means something.

  • Many runs a day instead of one

    Forks start warm, so an episode begins where your setup script left off, not at an empty disk.

An eval suite, quoted before it runs
bash
# 1. Build a verified snapshot, rejected if the typecheck fails
promigence snapshot create --image node:22 \
  --setup "git clone https://github.com/your-org/evals /work/repo && git -C /work/repo checkout <40-char sha>" \
  --setup "cd /work/repo && npm ci" \
  --verify "cd /work/repo && npm run typecheck" --alias evals --wait --wait-timeout 10m

# 2. Know the bill before you spend it
promigence run quote --snapshot evals --count 20 --timeout 15m
# → quote $0.83 — 20 x 900s x $0.000046/s (tier small)

# 3. Fan out, capped
promigence run --snapshot evals --count 20 --cap 1 \
  --verify-each --json -- ./run_agent.sh

# 4. The number about your environments
promigence run report <run_id>
Language: bash

Full quickstart

Run your own workload on it, free

Send a repo and the command you run against it, whatever that is: an eval suite, an RL rollout, a CI job, a queue of coding tasks. We build the environment once, run it a thousand times, and send back the timings, the failures and an exact price. Free, once, on your real workload.

  • 1,000 runs of your own command, on your own repo
  • What each one cost in wall clock, and anything that failed
  • Whether a failure was your code or the environment
  • An exact price for your real volume

A person replies, usually the same day. Or write to support@promigence.ai.

Not ready to hand over a repo? Read the quickstart or check the numbers first.

Join the waitlist

Promigence is in private beta. Leave your work email and we will send an invite as places open.