---
title: "Agent evaluation infrastructure, verified environments at 1,000 episodes"
description: "Promigence runs agentic coding evaluations on verified, reproducible environments and reports how many were valid, so your eval number is not quietly wrong."
url: "https://www.promigence.ai/for/agent-evals"
site: "Promigence"
---

# Run agent evaluations on environments you can prove were valid

Promigence verifies every environment before an agent runs in it, forks a thousand from one snapshot, and reports which failures were the environment's and which were the agent's. That distinction is the difference between an eval number you can publish and one you cannot.

**Who this is for:** Evaluation and platform engineers at AI agent companies, benchmark producers, and eval-as-a-service teams.

## What breaks today
### You cannot say how many environments were broken
Ask a team what fraction of their environments failed to install and the answer is a shrug. Those failures are recorded as agent failures, and they concentrate in the repos hardest to set up, which are the repos the benchmark is about.

### Last month's run is not comparable to this month's
Environments built from a recipe re-resolve package versions every time. When the environment drifts underneath you, a regression in your agent and a change in your environment look identical.

### One run per day, because a run takes hours
Cold clone, install and first build costs 60 to 100 seconds on every episode in the suite. That is what turns an eval into an overnight job.

## What changes with Promigence
### Every environment ran its own tests before your agent arrived
Declare a verify command with the snapshot. If it fails, the snapshot is rejected, never reaches an episode, and is not charged for.

### Each run ends with a number about the environments
Verified, environment failures, task failures, each replayable from the flight recorder. Publish it next to your result.

### The same hash reproduces the same environment in March
Cite the snapshot hash in your methods section. Re-run against it later and the comparison means something.

### Many runs a day instead of one
Forks start warm, so an episode begins where your setup script left off, not at an empty disk.

## An eval suite, quoted before it runs
```bash
# 1. Build a verified snapshot, rejected if the typecheck fails
promigence snapshot create --image node:22 \
  --setup "git clone https://github.com/your-org/evals /work/repo && git -C /work/repo checkout <40-char sha>" \
  --setup "cd /work/repo && npm ci" \
  --verify "cd /work/repo && npm run typecheck" --alias evals --wait --wait-timeout 10m

# 2. Know the bill before you spend it
promigence run quote --snapshot evals --count 20 --timeout 15m
# → quote $0.83 — 20 x 900s x $0.000046/s (tier small)

# 3. Fan out, capped
promigence run --snapshot evals --count 20 --cap 1 \
  --verify-each --json -- ./run_agent.sh

# 4. The number about your environments
promigence run report <run_id>
```

