# Promigence, complete reference > Promigence is a sandbox platform for AI agents: production-grade, fully isolated sandboxes that start in milliseconds, stay fast at a thousand at once, and bill by the second. Generated from the site's own content modules on 2026-10-03. Canonical site: https://www.promigence.ai --- ## What Promigence is Promigence is a sandbox platform for AI agents: production-grade, fully isolated sandboxes that start in milliseconds, stay fast at a thousand at once, and bill by the second. **Category:** sandbox platform for AI agents. Promigence does not build datasets, scorers, judges or experiment dashboards; it is the execution plane those products run on. **Primitive:** `verify → snapshot → fork(N) → exec → validity report` **Who it is for:** Anyone running AI agents: coding agents, evals, RL rollouts, agentic CI, background agents, or a sandbox per user inside an app. ### Key facts - Promigence sells verified execution environments for AI agents, metered per second with the maximum cost of a run quoted and capped before it starts. - A Promigence sandbox is ready for its first command in 8 ms at p50. - Every Promigence environment is checked ready before you get it; a six-hour soak ran 108,929 of them with zero failures and zero leaks. - A Promigence snapshot is stored only if it passes its own test at build time, so a broken environment never reaches an agent. - Each Promigence environment is fully isolated from every other, and a busy environment is never killed for doing the work you asked for. - The same Promigence snapshot hash returns a bit-identical environment on any later date, which makes an eval result citable. - Every Promigence run reports environment failures separately from task failures, so a broken environment never counts against your agent. - Promigence quotes the exact maximum cost of a run before it starts and enforces a hard spend cap. - Promigence runs any coding-agent work, not only evaluations: an episode is fork, run the agent, terminate, so an eval episode, an RL rollout, a CI job and a queue of bug-fix tasks are the same shape to it. - Promigence is model-neutral and lab-neutral, which is what a cross-lab evaluation requires. --- ## The guarantees ### 1. A busy environment is never marked unhealthy _Mechanism: health is never starved by your workload_ Health checks cannot be starved by your workload, so nothing kills an environment for doing what you asked, even with every vCPU busy. ### 2. A snapshot that fails its own test is never stored _Mechanism: verified builds, rejected builds are free_ Every snapshot runs your verify command at build time. If it fails, the build is rejected and you are not charged for it. ### 3. An environment that is not ready is never handed over _Mechanism: checked ready before hand-over_ Every environment is checked ready before you get it. You never get a handle that looks healthy and is silently cold. ### 4. Resource caps are checked at build time _Mechanism: build-time cap validation_ Memory and disk ceilings are validated when the snapshot is built, so you learn your monorepo needs 12 GB during the build and not 762 seconds into a paid episode. ### 5. Every failure is attributed _Mechanism: validity report, environment versus task_ Each run splits environment failures from task failures. An agent that failed because the environment was broken is never counted as an agent that failed the task. ### 6. The bill is quoted before the run and capped _Mechanism: quote before run, hard spend cap_ `promigence run quote` returns the exact ceiling before anything starts. A run that would exceed your cap is refused, not truncated halfway. --- ## Numbers ### Headline - **207 ms**, end to end, through the CLI _(measured, 2026-09-13)_. 207 ms from the US east, 204 ms from the US west, 209 ms from Europe, all p50 as the client saw it. Measured 2026-09-13 through the path a customer uses: the promigence CLI, over the network, an 8 vCPU / 8 GiB sandbox, from clients in three regions. - **108,929**, forks in six hours, zero failures _(measured, 2026-09-14)_. 5.99 h, p50 43 ms at the first hour and 43 ms at the last, slope −0.09 ms/hour. Measured on the server itself rather than over the network. - **177 ms**, at 200 concurrent, all served _(measured, 2026-09-14)_. 200 at once, every one served on the first attempt, 0 retries. A thousand in waves of a hundred holds the same: p50 99 ms, p99 165 ms. Measured on the server itself rather than over the network. ### Measured facts from our own research - **104 ms**, cold start of an empty sandbox _(measured, 2026-09-14)_. Per tier: 104 ms small, 127 ms medium, 142 ms large. - **39 ms**, warm fork on the box _(measured, 2026-09-14)_. One at a time 39.4 ms, a hundred at once 121.6 ms, two hundred at once 177.3 ms with all 200 served. Measured on the server itself rather than over the network. - **987 ms**, first copy on a server that has never held it _(measured, 2026-09-14)_. The first copy of a 12.94 GB environment on a server that had not held it before. - **12 → 100**, how far one provider's burst moved in 12 days _(measured, 2026-09-16)_. An independent public burst test: time-to-init, 100 concurrent, 120 s timeout. Not our measurement; read from its published raw results. --- ## Pricing Metered per second at $0.0504 per vCPU-hour plus $0.0162 per GiB-hour, the market's standard list rate. Paused time is not billed and nothing accrues before an environment is handed over. The exact maximum cost of a run is quoted before it starts and a hard spend cap is enforced. | Tier | Environment | Per hour | Per second | Typical workload | |---|---|---|---|---| | Small | 2 vCPU · 4 GiB | $0.17 | $0.000046 | A SWE-bench-class episode: clone, patch, run the failing test. Ten minutes costs about $0.03. | | Medium | 4 vCPU · 16 GiB | $0.46 | $0.000128 | Monorepo typechecks and mid-size builds. | | Large | 8 vCPU · 32 GiB | $0.92 | $0.000256 | Heavy builds and memory-hungry test suites. | ### Plans - **Free**, $0 pay as you go. Start free, then pay only for the seconds you run. Usage: $50 of free credit included; then per second at list, only what you use. Includes: $50 of free credit to start; 20 sandboxes at once, sessions up to an hour; Validity report on every run; Quote and hard spend cap on every run. - **Pro**, $50 per month. For teams running agents, evals and CI every day. Usage: More usage than you pay for included; then per second past it. Includes: Everything in Free, plus $50 of free credit; More compute each month than the plan costs; Up to 100 sandboxes at once; Verified snapshot builds, failed builds free; Harness adapters: Inspect, Harbor, SWE-bench; Email support. - **Max**, $500 per month. For large eval fleets, RL and continuous agentic CI. Usage: About 10× the usage of Pro included; then below list, by commitment. Includes: Everything in Pro, plus $50 of free credit; About 10× the usage of Pro; Up to 1,000 sandboxes at once; Capacity held warm for your snapshots; Shared Slack channel. - **Custom**, Custom volume pricing, set on one call. For vendors and labs running millions of episodes. Usage: Set on the call included; then volume rates. Includes: Everything in Max, plus $50 of free credit; Committed capacity contracts and reserved capacity; Running it in your own cloud account: talk to us; Custom environment shapes; Direct line to engineering. ### Other charges - **Verified snapshot build**: $1 per build. A build that fails its own verify command is not charged. - **Snapshot storage**: Included. Included in every plan. - **Paused time**: Not billed. A paused environment accrues nothing, and nothing accrues before one is handed to you. - **Egress**: Included. Registries and git hosts are not metered separately. The trial caps data out at 20 GB a month. ### Cost of a 10-minute episode, compared | Option | Cost | What that buys | |---|---|---| | Self-hosted Docker, raw cloud compute | ~$0.014 | Before engineering time, image maintenance and burst orchestration, with unknown validity. | | Typical managed sandbox, list | ~$0.03 | The same rate we charge. The difference is the 60 to 100 s of cold setup you pay for on every episode. | | Premium managed sandbox, list | ~$0.06 | Roughly three times the typical per-vCPU rate. | | Promigence, list | ~$0.03 | Warm verified fork, validity report, reproducible by hash, quoted and capped before the run. | ### Pricing principles - **The same rate as the market, per second**, $0.0504 per vCPU-hour and $0.0162 per GiB-hour, the market's standard list rate. A price nobody can compare against is worth very little, so we did not invent a new unit. - **You see the ceiling before the run**, `promigence run quote` returns the exact maximum before a single environment starts. `--cap` refuses a run that would exceed it rather than stopping halfway and charging for the half. - **A faster run is a cheaper run**, The meter runs from the moment an environment is handed to you until it ends. Starting warm instead of cold cuts 60 to 100 seconds off every episode, and that comes straight off the bill. - **Our failures are on us**, A build that fails its own verify command is not charged. Neither is an episode the environment broke, when our side can show the failure was ours. We can only refuse to bill for that category because we measure it. --- ## Workloads ### Run agent evaluations on environments you can prove were valid Promigence verifies every environment before an agent runs in it, forks a thousand from one snapshot, and reports which failures were the environment's and which were the agent's. That distinction is the difference between an eval number you can publish and one you cannot. **Who:** Evaluation and platform engineers at AI agent companies, benchmark producers, and eval-as-a-service teams. **What breaks today** - **You cannot say how many environments were broken**, Ask a team what fraction of their environments failed to install and the answer is a shrug. Those failures are recorded as agent failures, and they concentrate in the repos hardest to set up, which are the repos the benchmark is about. - **Last month's run is not comparable to this month's**, Environments built from a recipe re-resolve package versions every time. When the environment drifts underneath you, a regression in your agent and a change in your environment look identical. - **One run per day, because a run takes hours**, Cold clone, install and first build costs 60 to 100 seconds on every episode in the suite. That is what turns an eval into an overnight job. **What changes** - **Every environment ran its own tests before your agent arrived**, Declare a verify command with the snapshot. If it fails, the snapshot is rejected, never reaches an episode, and is not charged for. - **Each run ends with a number about the environments**, Verified, environment failures, task failures, each replayable from the flight recorder. Publish it next to your result. - **The same hash reproduces the same environment in March**, Cite the snapshot hash in your methods section. Re-run against it later and the comparison means something. - **Many runs a day instead of one**, Forks start warm, so an episode begins where your setup script left off, not at an empty disk. ```bash # 1. Build a verified snapshot, rejected if the typecheck fails promigence snapshot create --image node:22 \ --setup "git clone https://github.com/your-org/evals /work/repo && git -C /work/repo checkout <40-char sha>" \ --setup "cd /work/repo && npm ci" \ --verify "cd /work/repo && npm run typecheck" --alias evals --wait --wait-timeout 10m # 2. Know the bill before you spend it promigence run quote --snapshot evals --count 20 --timeout 15m # → quote $0.83 — 20 x 900s x $0.000046/s (tier small) # 3. Fan out, capped promigence run --snapshot evals --count 20 --cap 1 \ --verify-each --json -- ./run_agent.sh # 4. The number about your environments promigence run report ``` ### Reinforcement-learning environments that cannot poison the reward Promigence runs thousands of concurrent RL environments forked from one verified snapshot, each proven to work before the episode starts. A broken environment does not just lose you an episode: the optimiser treats its failure as real and the model learns from something that never happened. **Who:** RL environment vendors, data vendors, and labs' own post-training teams running coding agents. **What breaks today** - **A broken environment is a poisoned reward**, The agent writes a correct patch, the environment failed to install a dependency, the test fails, and the optimiser updates on it. Unlike a wrong eval chart, this one is baked into weights. - **Concurrency ceilings you discover in production**, Published accounts put real fleets capping at 300 to 550 concurrent, with orchestration the bottleneck past 256. A training step that needs 2,000 does not degrade gracefully, it stalls. - **Per-second billing on millions of episodes**, At millions of episodes, cold setup is most of the bill, and you are paying it to run the same install over and over. **What changes** - **Verification is a precondition, not a hope**, No snapshot is stored that cannot pass its own test, and no fork is handed over cold. Environment failures are reported separately, so they never enter the reward signal. - **Fan-out designed for the training step, not the demo**, A thousand forks for one training step start warm, not as a thousand cold installs. - **Volume pricing on committed capacity**, Committed-use rates for sustained volume. Talk to us about running it in your own cloud account. ```python from promigence import Promigence client = Promigence() # reads PROMIGENCE_API_KEY # The most one training step can cost, before anything starts quote = client.runs.quote("task-suite", count=20, spend_cap=1.0) # 20 at once in the private beta # Fan out the step: copies of one verified snapshot, each checked first run = client.runs.start( "task-suite", count=20, spend_cap=1.0, verify_each=True, exec={"sh": f"python -m rollout --policy {policy_url}"}, ) # Environment failures apart from task failures, and what the step cost report = client.runs.report(run["run_id"]) print(report["totals"], report["cost"]) ``` ### SWE-bench and friends, without the broken-environment tax Promigence runs SWE-bench-class benchmarks on environments verified before the instance starts and reproducible from a hash afterwards. Published work has found roughly a third of the hardest instances broken at the environment level, a tax paid out of every score computed on them. **Who:** Teams publishing or consuming SWE-bench, SWE-rebench, Terminal-Bench, SetupBench and similar results. **What breaks today** - **A third of the hardest instances do not set up**, Breakage in public task sets is concentrated in exactly the instances that separate good agents from great ones, so scores computed across them are not wrong by a random amount. - **Nobody can reproduce your run**, A benchmark defined by a repo and a script measures the repo as much as the agent. Rebuild it six months later and you are comparing two different things. - **The harness is fine; the fleet is the bottleneck**, Most harnesses are fine. What fails at 500 instances is image pulls, concurrency limits and cold setup repeated per instance. **What changes** - **One verified snapshot per instance family**, Built once, verified once, forked per instance. Instances that cannot build show up as a list at build time, not as mysterious agent failures. - **A hash you can put in the paper**, `snap_…` identifies a bit-identical environment. Anyone re-running against that hash gets what you had. - **Drop-in behind your existing harness**, Adapters for common eval harnesses (pre-release), plus a compatibility SDK class so existing sandbox call sites keep working. ```typescript // Most existing sandbox call sites keep working through the compatibility class. import { Sandbox } from "promigence"; const sandbox = await Sandbox.create("swebench-django"); // a snapshot you built await sandbox.commands.run("git apply /tmp/model_patch.diff"); const result = await sandbox.commands.run("pytest -q tests/"); await sandbox.kill(); // ← the episode ends here; billing stops ``` ### Agentic CI with a known cost before every run Promigence gives every agentic CI job a warm environment forked from a verified snapshot of your repo, so an agent starts at a built tree rather than at a `git clone`. Most of an agent-heavy CI bill is repeated setup. Forking removes it, and a quote before each run keeps the monthly number predictable. **Who:** Platform teams running coding agents against their own repositories in CI. **What breaks today** - **Setup is most of the job**, A five-minute agent task behind a ninety-second install is 30% overhead, paid on every job. - **The bill scales with enthusiasm**, Agentic CI usage grows the moment it works, and per-second billing means the finance conversation happens afterwards. - **Flaky environments look like flaky agents**, A dependency that fails to resolve at 3am produces a red build nobody can attribute, and trust erodes in the agent rather than the fleet. **What changes** - **A snapshot per branch, rebuilt when the lockfile moves**, So most pushes start from a snapshot that is already built and verified. - **Quote and cap per pipeline**, Every workflow knows its ceiling before it starts, and a runaway loop is refused rather than invoiced. - **Environment failures show up as environment failures**, Red builds get a cause attached, which is the only way an agent-driven pipeline keeps anyone's trust. ```yaml - name: Agent pass on the PR run: | promigence run \ --snapshot app --secret ANTHROPIC_API_KEY \ --count 1 --cap 0.50 --timeout 10m --json \ -- sh -c "cd /work/repo && git fetch -q origin ${GITHUB_SHA} && git checkout -q ${GITHUB_SHA} && claude -p 'fix the failing tests in this diff'" env: PROMIGENCE_API_KEY: ${{ secrets.PROMIGENCE_API_KEY }} ``` ### CI runners that finish the same job 2–3.5× sooner Promigence runs your existing GitHub Actions and GitLab CI jobs in its sandboxes. One line of the workflow changes (`runs-on: promigence`, or a runner tag in GitLab), the job itself does not. On the same job at the same size, measured against two hosted runner services, it finished 2–3.5× sooner and cost 7–10× less. Coding agents working inside CI have their own page, Agentic CI. **Who:** Teams on GitHub Actions or GitLab CI whose jobs are slow, costly, or too big for the default runner. **What breaks today** - **Billed by the minute**, A 50-second job is billed as a full minute, and each step up in runner size multiplies the per-minute rate. - **The default runner runs out of memory**, On the TypeScript monorepo we measured, the default runner of both hosted services ran out of memory and never finished. - **Bigger runners wait for a machine**, A larger hosted runner can sit provisioning for minutes before the first step runs. **What changes** - **One line moves a job**, `runs-on: promigence` in GitHub Actions, a runner tag in GitLab CI. Steps, caches and secrets stay where they are. - **Per second, with a cap per job**, Each job is billed for the seconds it runs, and a cap sets the most one job may cost. - **A size per job**, Name a size in `runs-on`, and keep runners waiting when a job should start the moment it is queued. ```yaml jobs: test: runs-on: promigence # was: ubuntu-latest steps: - uses: actions/checkout@v4 - run: npm ci && npm test # then, on any machine with the CLI: # promigence runners github --repo your-org/your-repo ``` ### Coding agents that start work on a repo that is already built Promigence gives a coding agent a sandbox forked from a verified snapshot of the repository, with dependencies installed and the build done. The agent starts at a built tree instead of a git clone, so its edit-and-check loop runs at warm speed from the first step. **Who:** Teams building coding-agent products and software factories that run agents against customer repositories. **What breaks today** - **Every task starts with a clone and an install**, Before an agent writes a line, it waits for the repository to download, the dependencies to install and the first build to finish. On a real repository that is minutes, paid again on every task. - **Each check in the loop is a cold build**, An agent that edits, type-checks and edits again pays the cold price on every check. A type-check that takes 24 seconds cold sets the pace of the whole loop. - **Your users watch the spinner**, The wait is theirs. A task that takes nine minutes loses the person who asked for it. **What changes** - **Start from a built snapshot**, Build the repository once, with a verify command. Every sandbox forks from that result, so the agent's first command runs against an installed, built tree. - **Warm checks at every step**, The type-check that costs 24 seconds on a cold machine costs 1.5 seconds on a warm fork, so a twenty-step task finishes in about a minute instead of nine. - **A sandbox per task, as many as you need**, Fork one per task. Each is isolated, and one that sits idle pauses until it is needed again. ```bash # Once per repo, or when the lockfile changes promigence snapshot create --image node:22 \ --setup "git clone https://github.com/your-org/app /work/repo && git -C /work/repo checkout <40-char sha>" \ --setup "cd /work/repo && npm ci && npm run build" \ --verify "cd /work/repo && npm test" --alias app --wait --wait-timeout 10m # Per task: a warm sandbox, then your agent's commands promigence fork --snapshot app --count 1 --cap 1 promigence exec --sandbox "$SANDBOX" -- sh -c "cd /work/repo && npm run typecheck" ``` ### Background agents with a full dev environment, and no bill while they wait Promigence gives background agents a full development environment forked from a verified snapshot: your repository, its services and its build, ready when the agent starts. A long-running agent can pause while it waits on a person, a queue or a timer, and paused time is not billed. **Who:** Platform teams building internal background agents, and teams running vendor coding agents on their own pool. **What breaks today** - **A real task needs the whole dev environment**, Databases, queues and a built frontend, not just a repository. Setting that up for every task takes longer than the task. - **Agents spend most of their time waiting**, On the model, on a tool, on a reviewer. A sandbox billed for every second of that wait costs more than the work it does. - **Secrets and network access need a boundary**, An agent with production-like access must reach the hosts it should and nothing else, with credentials that are not baked into the environment. **What changes** - **One snapshot of the full environment**, Set up the services and the build once, verify them, and fork the finished environment for each agent run. - **Pause while the agent waits**, Set an idle rule and a sandbox pauses itself when nothing is happening, then wakes on the next request. Paused time is not billed. - **Allow-listed network and secrets at run time**, An outbound allow-list, with every refused host recorded, and secrets injected at run time, never stored in the snapshot your build produces. ```bash # Store a credential once; it is injected at run time, never stored in the snapshot your build produces promigence secrets set GITHUB_TOKEN --from-env GITHUB_TOKEN # Fork the environment and let it pause itself when idle promigence fork --snapshot devenv@9b0d3f1 --count 1 --cap 1 promigence sandbox autopause "$SANDBOX" --after 30s # What it may reach, and every host it was refused promigence sandbox network "$SANDBOX" ``` ### AI code review that checks every pull request on a warm, built repo Promigence gives an AI code-review agent a sandbox forked from a verified, already-built snapshot of the repository, so it can build, run the tests and check its suggestions on every pull request without paying for a fresh clone and install each time. **Who:** Companies building AI code-review products, and teams running review agents on their own repositories. **What breaks today** - **Every pull request pays for a fresh build**, To check a suggestion, the reviewer needs the code built and the tests runnable. From a cold machine that is a clone, an install and a build on every pull request. - **Suggestions nobody ran**, A review comment that was never compiled or tested is a guess, and developers learn to ignore guesses. - **Review traffic arrives in bursts**, A busy morning lands dozens of pull requests at once, and a review that queues behind the others arrives after the author has moved on. **What changes** - **Fork from the latest built snapshot**, Snapshot the default branch with its build. Each pull request forks from it and brings only its own changes. - **Run the checks before you comment**, Build and test the suggested change inside the sandbox, so a comment can say whether it passed. - **Many reviews at once**, Fork one sandbox per pull request. A burst of reviews starts together instead of waiting in a queue. ```bash # When the default branch moves promigence snapshot create --image node:22 \ --setup "git clone https://github.com/your-org/app /work/repo && git -C /work/repo checkout <40-char sha>" \ --setup "cd /work/repo && corepack enable && pnpm install && pnpm build" \ --verify "cd /work/repo && pnpm test" --alias app --wait --wait-timeout 10m # Per pull request: fork, check out the change, test before commenting promigence fork --snapshot app --count 1 --cap 1 promigence exec --sandbox "$SANDBOX" -- sh -c "cd /work/repo && git fetch origin $PR_REF && git checkout FETCH_HEAD" promigence exec --sandbox "$SANDBOX" -- sh -c "cd /work/repo && pnpm test" ``` ### Code interpreters that answer while your user is still looking Promigence runs the code your AI app writes in an isolated sandbox that is ready in milliseconds. In our benchmark a code-interpreter session took 1.15 seconds from start to result at the median, and Promigence works with the sandbox API many AI apps already use. **Who:** Teams building AI analysts, chat-with-your-data features and any AI app that runs code for its users. **What breaks today** - **Your user is watching**, A code-interpreter answer is interactive. Every second of sandbox start-up is a second your user stares at a spinner in your product. - **Model-written code needs a real boundary**, The model writes the code, and anyone can prompt the model. It has to run somewhere isolated, with network access you control. - **Switching providers should not be a rewrite**, Code written against one sandbox API should not have to change to move to a faster one. **What changes** - **Ready in milliseconds**, A sandbox is ready for its first command 8 milliseconds after it is requested. - **Isolated, with network you decide**, Each session runs in its own isolated sandbox, which can sit behind an outbound allow-list. - **A drop-in for the API you use**, Promigence works as a drop-in replacement for the sandbox API many AI apps already use, so moving over is a small code change. ```bash # A fresh sandbox from the base image promigence fork --count 1 --cap 1 # Run the model's code, isolated promigence exec --sandbox "$SANDBOX" -- python3 analysis.py ``` ### A sandbox backend for agent-builder and workflow platforms Promigence gives agent-builder and workflow platforms a sandbox backend for the code their users' agents run: isolated, ready in milliseconds, billed by the second, and compatible with the sandbox API many platforms already integrate. It fits as a second provider next to the one you have. **Who:** Agent-builder, workflow-automation and low-code AI platforms that run user-defined code or agents. **What breaks today** - **One provider is a single point of failure**, When your only sandbox provider has a bad day, every one of your customers' workflows fails with it. - **Your users' code is untrusted**, Workflows run code your platform did not write. Each run needs its own isolated sandbox and a network boundary. - **Idle workflows still cost money**, Workflows wait on models, tools and people. Paying for sandboxes through that wait turns every idle workflow into a bill. **What changes** - **A second provider behind the same API**, Promigence works as a drop-in replacement for the sandbox API many platforms already integrate, so it can take traffic next to your current provider. - **Isolated runs with network you decide**, Every run gets its own isolated sandbox, which can sit behind an outbound allow-list. - **Billed by the second, not for the wait**, Usage is billed per second, and a sandbox can pause itself when idle. Paused time is not billed. ```bash # Know the most a batch can cost before it runs promigence run quote --snapshot workflows --count 20 # Run it with a hard cap promigence run --snapshot workflows --count 20 --cap 2 -- ./run_workflow.sh ``` ### AI pentest agents that attack a fresh copy of the target every time Promigence lets an offensive-security agent fork the same prepared target as many times as it needs, so every exploit attempt starts from an identical, clean state, behind an outbound allow-list. **Who:** Teams building AI penetration-testing and offensive-security agents. **What breaks today** - **Every attempt changes the target**, An exploit that half-worked leaves the target in a state the next attempt did not expect, and results stop being comparable. - **Parallel attempts need parallel targets**, Rebuilding the target for each attempt is slow, so attempts run one after another. - **An attacking agent must not reach the wrong hosts**, An agent with open network access is a liability. It has to reach the target and nothing else. **What changes** - **Fork the prepared target**, Build the target once as a verified snapshot and fork a fresh copy per attempt. Each starts identical and isolated. - **Many attempts at once**, Forks start warm, so dozens of attempts can run side by side from the same starting point. - **An outbound allow-list**, An outbound allow-list, every refused host recorded, and a sandbox's egress can be cut at once. ```bash promigence snapshot create --image debian:12 \ --setup "apt-get update && apt-get install -y git" \ --setup "git clone https://github.com/your-org/target-app /work/app && git -C /work/app checkout <40-char sha>" \ --setup "cd /work/app && ./provision.sh" --verify "cd /work/app && ./healthcheck.sh" \ --alias target-app --wait --wait-timeout 10m promigence fork --snapshot target-app --count 20 --cap 5 # What a copy may reach, and every host it was refused promigence sandbox network "$SANDBOX" ``` --- ## Definitions ### Environment runtime (also: agent environment runtime, execution plane) An environment runtime is infrastructure that builds, verifies, reproduces and forks the environment an AI agent runs inside, treating the environment itself as the product. A sandbox provider sells a box that starts: it boots fast, it isolates the process, and what happens inside is your problem. An environment runtime sells an environment that is known good, it ran its own tests before you got it, it reproduces exactly from a hash, and a thousand copies of it are all the same. The distinction matters because cold start has converged industry-wide to a floor of tens of milliseconds. What still differs by orders of magnitude is how many of your environments were silently broken, and whether the bill is predictable. ### Verified environment A verified environment is one that ran its own declared test command at build time, so it is known to work before an agent is placed inside it. Verification is a build-time gate, not a runtime check. You declare a verify command with the snapshot, usually the same command your CI runs, and the snapshot is stored only if it exits zero. A failed verification costs nothing. Without it, a broken environment reaches an episode, the agent fails inside it, and the result is recorded as the agent failing the task. At one episode you would notice. At a thousand you do not. ### Episode An episode is one environment's complete life, fork, execute, terminate, and it is the unit a quote and a spend cap are written against. The term comes from reinforcement learning, and the shape is the same in agentic evaluation: an environment is created, an agent works in it, the result is scored, the environment is destroyed. Promigence meters by the second, like the rest of the market, but it prices a run by the episode before you start it. `promigence run quote` multiplies the number of episodes by what each can cost at most, and `--cap` refuses a run that would go past your ceiling. So you get a per-second bill with a number you agreed to first. ### Snapshot fork (also: warm fork) A snapshot fork creates a new environment from an already-warm one that has finished installing and building, so it costs milliseconds rather than seconds. The alternative is booting a machine, cloning a repo, installing dependencies and building it: 60 to 100 seconds of identical work, repeated on every episode. A fork skips all of it. The parent is already built, already warm, and the copy starts from there. A Promigence copy is ready for its first command 39 ms in, and a hundred at once still land at 122 ms. ### Warm start A warm start begins work in an environment where dependencies are installed and caches are populated, so the first command you run is the one you care about. The industry benchmarks cold start because it is easy to measure and easy to market. It is also converging, with a floor of tens of milliseconds. The warm path still differs by an order of magnitude. A typecheck that takes 24.35 s cold takes 1.55 s warm; another goes from 101.4 s to about 1.5 s. Across a twenty-cycle agent task that is 527 s against 71 s for identical work. ### Validity report (also: environment validity) A validity report states, for one run, how many environments were verified, which failed, and whether each failure belonged to the environment or the task. It answers a question most teams running evaluations cannot: how many of the environments behind this number were actually working. Without it, an environment that failed to install is indistinguishable from an agent that failed the task. In reinforcement learning the consequence is worse than a wrong chart, the training loop updates on a failure that never really happened. Promigence returns a report after every run and keeps a flight recorder so a flagged failure can be opened. ### Poisoned reward A poisoned reward is a training signal produced by a broken environment rather than by the agent's behaviour, so the model learns from an outcome that never happened. If an environment fails to install a dependency, a correct patch still fails the test, the harness records a failure, and the optimiser updates on it. Nothing in the pipeline separates 'the agent was wrong' from 'the environment was broken'. A wrong eval number is embarrassing and recoverable. A poisoned reward is baked into weights and costs a training run to undo. The defence is to verify before the episode and attribute every failure in the report. ### Burst capacity Burst capacity is how many environments a provider actually delivers when many are requested at once, which is often very different from its advertised latency for one. The gap is measurable, and it moves. In an independent test requesting 100 concurrent sandboxes, one provider went from serving 12 to serving all 100 within two weeks. Both runs are real, which is the point: one good day tells you nothing about the next one. For batch work this is the reliability number that matters. A fan-out that half-completes invalidates an evaluation rather than degrading it, because the environments that did not start are not a random sample. ### Flight recorder A flight recorder is a per-episode log of every command executed and every file diff produced, kept so a failed episode can be read rather than guessed at. Debugging a batch failure without one means re-running the episode and hoping it fails the same way. With one, the flagged episodes in a validity report are openable: exact commands, exit codes, output, and what changed on disk. Promigence keeps each failed episode for 7 days by default, so you can open it or re-run it with `promigence run replay`. --- ## CLI reference Install: `npm i -g ./promigence-.tgz` (the file sent with your invite; not on npm), then `promigence signup` ### Global flags - `--json`, one JSON document on stdout; the default whenever stdout is not a TTY, except for exec --sandbox - `--text`, human-readable text - `--api-key KEY`, precedence: --api-key > PROMIGENCE_API_KEY > ~/.config/promigence/credentials - `--version`, print the version - `--help`, help; `promigence --help` for one command ### Setup Sign up, check the install, and get an agent oriented in about 1,200 tokens. #### `promigence login` sign up or sign in (email; GitHub/Google where enabled); stores a new API key - `--provider NAME`, a sign-in method the API offers (email; github|google where enabled): skip the chooser - `--no-browser`, headless: print the URL, paste the code back - `--key-name N`, label for the key this login mints (default cli@) - `--api-url URL`, sign in to this API instead of the default (no prompt) - `--yes`, accept a non-default API from PROMIGENCE_API_URL without the prompt - `--ref CODE`, a referral code (new accounts) - `--json`, force the JSON envelope #### `promigence signup` create an account (asks for an invite code first while invite-only) - `--invite CODE`, the code (prefer the prompt; argv is visible) - `--provider NAME`, a sign-in method the API offers (email; github|google where enabled): skip the chooser - `--no-browser`, headless: print the URL, paste the code back - `--key-name N`, label for the key this sign-in mints (default cli@) - `--api-url URL`, sign up on this API instead of the default (no prompt) - `--yes`, accept a non-default API from PROMIGENCE_API_URL without the prompt - `--ref CODE`, a referral code (new accounts) - `--json`, force the JSON envelope #### `promigence logout` forget this CLI's key (revoked when `promigence login` minted it) - `--keep-key`, only forget it locally; the key keeps working - `--json`, force the JSON envelope #### `promigence auth login` store an API key (--from-env VAR | stdin; never on argv) - `--from-env VAR`, read the key from an environment variable - `--api-key K`, accepted but warns: key on argv #### `promigence keys list` your org's API keys - `--all`, include revoked keys - `--project P`, another project's - `--org ORG`, with --project: in this org - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence keys create` mint a key (secret printed once) - `--name N`, label, e.g. ci-prod - `--org ORG`, in this org (default: this key's) - `--project P`, in this project - `--service`, owners: outlives your membership (CI) - `--expires WHEN`, 30d, 1y, a date; default never - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence keys revoke KEY_ID` revoke a key - `--project P`, a key in another project - `--org ORG`, with --project: in this org - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org list` your organizations (* = default) - `--all`, every one - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org create NAME` create one you own (no trial credit) - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org switch ORG` sign in, make ORG the default, store a key for it - `--project P`, the new key's project - `--keep-old-key`, keep the replaced key working - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org members` members and roles - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org invite EMAIL` invite an email (owners) - `--role R`, owner|member (member) - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org invites` pending invites (14-day expiry) - `--mine`, sent to you - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org uninvite INVITE_ID` revoke an invite (owners) - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org role USER_ID ROLE` set owner|member (owners) - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org remove USER_ID` remove a member; their keys die (owners) - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org leave [ORG]` leave an organization - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org join INVITE_ID` accept an invite sent to you - `--decline`, decline it instead - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org audit` who did what (owners; needs --sign-in) - `--before SEQ`, older entries - `--limit N`, how many (100) - `--org ORG`, another org than this key's - `--sign-in`, sign in in your browser: the log is read as you, never with a key - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence org delete ORG` delete it; keys die (owners) - `--confirm NAME`, its name (required) - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence project list` projects (keys belong to one) - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence project create NAME` create a project - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence project rename PROJECT NAME` rename (not `default`) - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence project delete PROJECT` delete an empty one (owners) - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence project cap PROJECT USD|off` monthly USD cap, or off (owners) - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence project move PROJECT` to another org; keys die (owners) - `--to ORG`, yours, with a card - `--org ORG`, another org than this key's - `--no-browser`, sign in headless (or PROMIGENCE_ID_TOKEN) - `--yes`, sign in to PROMIGENCE_API_URL without the prompt - `--json`, force the JSON envelope #### `promigence me` plan, concurrency in use, trial credit, spend, limits - `--json`, force the JSON envelope #### `promigence billing card` add a card: prints a checkout URL (trial credits kept) - `--json`, force the JSON envelope #### `promigence doctor` check client, auth and API - `--json`, force the JSON envelope #### `promigence init` write promigence.toml - `--repo U`, git URL (default: the `origin` remote) - `--commit SHA`, commit to pin (default: HEAD) - `--setup CMD`, override the inferred setup step - `--verify CMD`, override the inferred verify command - `--tier xsmall|small|medium|large|xlarge|2xlarge`, default medium - `--force`, overwrite an existing promigence.toml - `--print`, write nothing; print what it would write - `--json`, force the JSON envelope #### `promigence examples` what forks with no build (no auth) - `--limit N`, how many to list (default 20) - `--bundle`, write each public snapshot as a .promigence file (none yet) - `--json`, force the JSON envelope #### `promigence agent-guide` ~1.2k-token guide for agents (--section commands|errors|cost) - `--section commands|errors|cost`, one section instead of the core - `--code C`, the entry for one error code - `--format md|json`, default md - `--install`, write the fence into AGENTS.md/CLAUDE.md, .claude/skills, .cursor/rules - `--create`, with --install: create AGENTS.md if neither file exists - `--check`, report what --install would change; write nothing #### `promigence mcp` MCP server on stdio: `claude mcp add promigence -- promigence mcp` (Windows: `-- cmd /c promigence mcp`) - `--read-only`, offer only the tools that read: list, read a file, fetch output, a report - `--allow-tools A,B`, offer only these tools (comma-separated names) - `--cap USD`, spend cap for every sandbox and batch it starts (else PROMIGENCE_SPEND_CAP, else the org default) - `--keep`, leave the sandboxes this session created running at shutdown #### `promigence status` regions · your concurrency in use vs your limit - `--json`, force the JSON envelope #### `promigence billing usage` this month's spend: plan, usage by tier, credits, next invoice - `--ledger`, also print this key's project's metered rows - `--limit N`, with --ledger: how many rows (default 20) - `--cursor C`, with --ledger: the next_cursor of an earlier page - `--json`, force the JSON envelope #### `promigence referrals` your referral code, its offer, your referrals - `--json`, force the JSON envelope #### `promigence referrals apply ` apply a code (owner, before the first paid invoice) - `--json`, force the JSON envelope #### `promigence pricing` the price list, live (yours with a key) - `--new-account`, what a new account gets - `--json`, force the JSON envelope #### `promigence credits` your free-credit offer: state, credit, expiry, campaign - `--json`, force the JSON envelope #### `promigence credits claim` open the claim page (accept the terms in the browser); prints its URL - `--no-open`, print the URL, open nothing - `--json`, force the JSON envelope #### `promigence billing portal` manage the card on file and past invoices: prints a portal URL - `--json`, force the JSON envelope ### Snapshots Build a verified environment once. Everything after this is a fork of it. #### `promigence snapshot create` build + verify a snapshot (--image | --dockerfile | --devcontainer) - `--repo R`, owner/name (connected GitHub) or a git URL - `--ref R`, with owner/name: branch|tag|sha - `--commit SHA`, 40-char sha, never a branch - `--image REF`, a tag (we pin it) or REF@sha256:D - `--dockerfile PATH`, a Dockerfile we build - `--context DIR|URL@SHA`, default: the Dockerfile's dir - `--devcontainer PATH`, a devcontainer.json (default: the one here) - `--no-devcontainer`, ignore the repo's devcontainer.json - `--dir PATH`, upload the files git tracks there - `--from SNAP`, owner/name|--dir: start from (base) - `--track`, owner/name: rebuild on every push - `--setup CMD`, setup step (ordered) - `--verify CMD`, must exit 0 or nothing is stored - `--start CMD`, background process kept running - `--tier xsmall|small|medium|large|xlarge|2xlarge`, build and fork size - `--alias N`, unique per org - `--grade CMD`, score the work in a clean env - `--fresh`, no 10-min replay (identical inputs still reuse their snapshot) - `--wait`, bounded by --wait-timeout (default 100s) - `--wait-timeout DUR`, bound (100s); expiry = exit 11, nothing killed - `--idempotency-key K`, a key you pick (24 h) - `--json`, force the JSON envelope #### `promigence snapshot status ` poll a build by build_id - `--wait`, bounded by --wait-timeout (default 100s) - `--wait-timeout DUR`, bound (100s); expiry = exit 11, nothing killed - `--log N`, print the last N lines of the build log - `--json`, force the JSON envelope #### `promigence snapshot list` list snapshots in your org - `--json`, force the JSON envelope #### `promigence snapshot inspect ` show a snapshot manifest - `--json`, force the JSON envelope #### `promigence snapshot verify ` re-verify a snapshot in a fork - `--json`, force the JSON envelope #### `promigence snapshot export ` write a portable .promigence bundle (commit it, cite it, mail it) - `--out FILE`, default .promigence; `-` or a pipe writes stdout - `--no-layers`, citation form: identity + provenance + verification, no chunk table - `--layers`, keep the chunk table however large (default: drop it above 100 kB) - `--json`, force the JSON envelope #### `promigence image resolve REF` resolve REF (python:3.12) to REF@sha256:… without a local container runtime - `--registry-secret NAME`, stored secret holding user:password for a private registry - `--json`, force the JSON envelope #### `promigence github connect` install the GitHub App (browser); waits for it - `--no-open`, print the link, open nothing - `--no-wait`, print the link and return - `--wait-timeout DUR`, for the install (10m); expiry = exit 11 - `--json`, force the JSON envelope #### `promigence github status` connected accounts, pending installs, tracked branches - `--json`, force the JSON envelope #### `promigence github disconnect ` stop building from a GitHub account - `--json`, force the JSON envelope #### `promigence github untrack ` stop rebuilding on push (the snapshot stays) - `--json`, force the JSON envelope #### `promigence repos` repos your GitHub connection reads (for --repo) - `--page N`, page (100 a page) - `--json`, force the JSON envelope #### `promigence snapshot delete ` remove one of your snapshots (refused while a fork of it is alive) - `--json`, force the JSON envelope #### `promigence snapshot builds` your builds, newest first (a rejected build keeps its reason) - `--limit N`, how many (default 20) - `--state S`, only this state, e.g. running | rejected - `--cursor C`, the next_cursor of an earlier page - `--json`, force the JSON envelope #### `promigence snapshot alias [NAME]` name a snapshot, or clear its name (unique per org) - `--clear`, remove the alias instead - `--json`, force the JSON envelope ### Runs Quote, cap, fork N copies, execute, and read which copies were valid. #### `promigence run -- CMD` fork N, exec CMD in each, print the validity report - `--snapshot H`, id | alias[@commit7] | ./env.promigence (default promigence/base) - `--image REF`, not with --snapshot: a tag or REF@sha256:D - `--count N` **(required)**, sandboxes; each has PROMIGENCE_INDEX (0..N-1) and PROMIGENCE_SEED - `--timeout DUR`, sandbox timeout (15m) - `--cap USD` **(required)**, spend cap; refused above it (exit 5) - `--allow-partial`, start what the cap allows; the rest not_started - `--secret NAME`, stored secret as an env var - `--verify-each`, verify forks before hand-over - `--network none|egress`, default egress - `--allow HOST`, outbound allow-list - `--optimize-for start|compute`, default start - `--keep-delta`, keep the writable delta - `--label K=V`, run label - `--wait-timeout DUR`, bound (100s); expiry = exit 11, nothing killed - `--grade CMD`, score the work in a clean env - `--holdout CMD`, second test set; gap reported - `--grade-score M`, exit|last_line_json|tap (exit); JSON: {"score":0.8,"pass":true} - `--grade-timeout DUR`, per grader (10m) - `--grade-output`, keep last 16 KiB - `--fail-on task|env|none`, what exits non-zero (task) - `--keep`, keep sandboxes after CMD - `--stream`, live output (auto on a TTY) - `--no-stream`, one envelope even on a TTY - `--max-output BYTES`, output cap (64k) - `--idempotency-key K`, a key you pick (24 h) - `--verbose`, text: fork timings and totals too #### `promigence run quote` price a run before launch - `--snapshot H`, id | alias[@commit7] | ./env.promigence (default promigence/base) - `--count N` **(required)**, number of sandboxes - `--timeout DUR`, sandbox timeout (15m) - `--tier T`, must match the snapshot's assigned tier - `--cap USD`, price it against this cap; spends nothing - `--json`, force the JSON envelope #### `promigence run report ID` validity report: env vs task failures - `--json`, force the JSON envelope - `--sandboxes all|failed`, default: summary + non-completed sandboxes - `--sandbox SID`, one sandbox - `--repro`, print the repro command for --sandbox - `--limit N`, text rows before '… N more' (50; 0 = all). --json is never paged #### `promigence run status ID` poll a run - `--wait`, bounded by --wait-timeout (default 100s) - `--wait-timeout DUR`, bound (100s); expiry = exit 11, nothing killed - `--json`, force the JSON envelope #### `promigence run cancel ID` stop queued starts - `--json`, force the JSON envelope #### `promigence run kill ID` kill every sandbox of a run - `--json`, force the JSON envelope #### `promigence run replay ID` replay sandboxes (client-triggered, never automatic) - `--sandbox SID`, only these sandboxes - `--json`, force the JSON envelope #### `promigence run open ID` reopen a failed sandbox as a live one, at the moment it failed - `--sandbox SID`, which failure (default: the only one) - `--cap USD`, what the reopened sandbox may spend - `--timeout DUR`, how long it may run (default: the run's) - `--json`, force the JSON envelope #### `promigence run captures ID` what of this run is preserved, and when it expires - `--delete SNAPSHOT_ID`, destroy one now, before it expires - `--json`, force the JSON envelope #### `promigence run list` your runs, newest first - `--limit N`, how many (default 20) - `--state S`, only this state - `--cursor C`, the next_cursor of an earlier page - `--json`, force the JSON envelope ### Sandboxes Long-lived sandboxes you drive command by command: exec, files, pause, resume, fork. #### `promigence fork` fork N sandboxes from a snapshot (--count --cap) - `--snapshot H`, snap_ | alias[@commit7] | ./env.promigence (default promigence/base) - `--count N` **(required)**, number of sandboxes - `--timeout DUR`, sandbox timeout (15m) - `--cap USD` **(required)**, spend cap; refused above it (exit 5) - `--allow-partial`, start what the cap allows; the rest not_started - `--secret NAME`, stored secret as an env var - `--verify-each`, verify forks before hand-over - `--network none|egress`, default egress - `--allow HOST`, outbound allow-list - `--optimize-for start|compute`, default start - `--keep-delta`, keep the writable delta - `--label K=V`, run label - `--wait-timeout DUR`, bound (100s); expiry = exit 11, nothing killed - `--no-wait`, return after POST /runs - `--fresh`, skip idempotent replay - `--stream`, live output (auto on a TTY) - `--no-stream`, one envelope even on a TTY - `--idempotency-key K`, a key you pick (24 h) - `--verbose`, text: fork timings and totals too - `--json`, force the JSON envelope #### `promigence exec -- CMD` run CMD in a sandbox (--sandbox SID) or a run (--run ID --all) - `--sandbox SID`, stdout is the child's own bytes (JSON only with --json or PROMIGENCE_OUTPUT=json); its exit code passes through; not found 127; Promigence's own failures 125 - `--run ID`, with --all: every ready sandbox; queued reported skipped unless --wait - `--all`, 0 if every child exits 0, else 4 - `--no-passthrough`, return 0–12 instead of the child's code - `--background`, with --sandbox: start it, print its exec_id, return now - `--stdin`, with --background: keep stdin open for `promigence exec stdin` - `--secret NAME`, stored secret as an env var - `--stream`, live output (auto on a TTY) - `--no-stream`, one envelope even on a TTY - `--max-output BYTES`, per stream (64k; 32m, the most, into a pipe or file); past it the rest is cut, and text exits 12 - `--timeout DUR`, child timeout (default 15m) - `--wait`, bounded by --wait-timeout (default 100s) - `--wait-timeout DUR`, bound (100s); expiry = exit 11, nothing killed - `--idempotency-key K`, a key you pick (24 h) #### `promigence exec output EXEC_ID` an exec's output by id (background x_…: --follow) - `--range A-B`, byte range, 0-based and B inclusive; `A-` runs to the end - `--sandbox SID`, resolve the exec inside this sandbox (required for x_…) - `--run ID`, resolve the exec inside this run - `--stderr`, the child's stderr instead of its stdout - `--follow`, background exec: stream until it exits, pass its code through - `--json`, force the JSON envelope #### `promigence exec ps` list a sandbox's background execs - `--sandbox SID` **(required)**, the sandbox - `--json`, force the JSON envelope #### `promigence exec stdin EXEC_ID` write to a background exec's stdin (needs `exec --background --stdin`) - `--sandbox SID` **(required)**, the sandbox it runs in - `--data TEXT`, write this (default: read stdin to EOF) - `--file P`, write this local file's bytes - `--eof`, close stdin after writing (implied when reading a pipe) - `--no-eof`, keep it open after piped input - `--json`, force the JSON envelope #### `promigence exec kill EXEC_ID` signal a background exec (default SIGKILL) - `--sandbox SID` **(required)**, the sandbox it runs in - `--signal SIG`, TERM | INT | HUP | KILL | a number - `--forget`, then drop its handle and buffered output - `--json`, force the JSON envelope #### `promigence files put SRC DST` copy a local file into a sandbox (--recursive: a directory) - `--sandbox SID` **(required)**, target sandbox - `--recursive`, SRC is a local directory; DST is a directory in the sandbox - `--json`, force the JSON envelope #### `promigence files get SRC DST` copy a file out of a sandbox (--recursive: a directory) - `--sandbox SID` **(required)**, source sandbox - `--recursive`, SRC is a directory in the sandbox; DST a local directory - `--exclude GLOB`, with --recursive: leave these out, e.g. .git - `--keep-links`, with --recursive: keep links leaving DST - `--json`, force the JSON envelope #### `promigence files ls [PATH]` list a directory in a sandbox (lstat per entry) - `--sandbox SID` **(required)**, the sandbox - `--recursive`, walk below PATH, not one level - `--depth N`, with --recursive: how many levels - `--limit N`, stop after N entries (truncated is reported) - `--no-hidden`, drop dotfiles (they are listed by default) - `--json`, force the JSON envelope #### `promigence files search [PATH]` find files by name glob and/or text inside them - `--sandbox SID` **(required)**, the sandbox - `--name GLOB`, name glob, e.g. '*.py' - `--contains TEXT`, literal text to find inside matching files - `--ignore-case`, case-insensitive --contains - `--limit N`, stop after N hits (truncated is reported) - `--no-hidden`, skip dotfiles and dot-directories - `--json`, force the JSON envelope #### `promigence files mkdir PATH` create a directory in a sandbox - `--sandbox SID` **(required)**, the sandbox - `--parents`, mkdir -p: intermediate directories, and no error if it exists - `--mode OCTAL`, e.g. 755 - `--json`, force the JSON envelope #### `promigence files mv SRC DST` move or rename a path inside a sandbox - `--sandbox SID` **(required)**, the sandbox - `--overwrite`, replace DST if it is already there - `--json`, force the JSON envelope #### `promigence files cp SRC DST` copy a path inside a sandbox (both ends are in the sandbox) - `--sandbox SID` **(required)**, the sandbox - `--recursive`, a directory and everything under it - `--overwrite`, replace DST if it is already there - `--json`, force the JSON envelope #### `promigence files chmod PATH` change mode and/or owner of a path in a sandbox - `--sandbox SID` **(required)**, the sandbox - `--mode OCTAL`, e.g. 755 - `--uid N`, owner - `--gid N`, group - `--recursive`, the directory and everything under it - `--json`, force the JSON envelope #### `promigence agent start claude|codex|opencode` a coding agent in a new sandbox - `--key-secret NAME[=VAR]`, stored secret with the model key (default ANTHROPIC_API_KEY; codex: OPENAI_API_KEY) - `--key-from-env VAR`, store the key from this variable first (never argv) - `--snapshot H`, start from this snapshot (default promigence/base; the agent is installed if missing) - `--image REF@sha256:D`, instead of --snapshot - `--repo URL`, clone this public repository and start the agent in it - `--prompt TEXT`, work on this task headless, in the background - `--attach`, open the agent's terminal now - `--timeout DUR`, sandbox lifetime (default 1h) - `--cap USD`, spend cap (else PROMIGENCE_SPEND_CAP, else the org default) - `--wait-timeout DUR`, bound (100s); expiry = exit 11, nothing killed - `--json`, force the JSON envelope #### `promigence sandbox kill SID` kill one sandbox - `--json`, force the JSON envelope #### `promigence sandbox network SID` what this sandbox may reach, and every host it was refused - `--cut`, cut this sandbox's egress now; it cannot be turned back on - `--json`, force the JSON envelope #### `promigence sandbox shell SID [-- CMD]` interactive terminal (raw TTY, follows window size) - `--cmd CMD`, run this instead of the login shell (sh -c) - `--cwd DIR`, start in this directory - `--env K=V`, environment variable - `--secret NAME[=VAR]`, stored secret in the terminal's environment; NAME=VAR binds it to VAR #### `promigence sandbox ssh SID [-- CMD]` log in over SSH (scp, sftp and -L work too) - `--config`, print an ssh config block instead of logging in - `--print`, print the connection details instead of logging in - `--key PATH`, use this key pair instead of making one - `--ttl S`, how long the access lasts (default 3600, max 3600) - `--idle S`, close a session after this long with no activity - `--proxy`, carry one connection on stdin/stdout (used by ssh itself) - `--list`, the keys that have access, then exit - `--revoke KEY_ID`, end one key's access (and its tunnel credential), then exit - `--json`, force the JSON envelope #### `promigence port expose PORT` publish a port as a preview URL (anyone with it can reach it) - `--sandbox SID` **(required)**, the sandbox - `--json`, force the JSON envelope #### `promigence port list` list a sandbox's preview URLs - `--sandbox SID` **(required)**, the sandbox - `--json`, force the JSON envelope #### `promigence port revoke TOKEN|PORT` take a preview URL down (a token, or every URL for a port) - `--sandbox SID` **(required)**, the sandbox - `--json`, force the JSON envelope #### `promigence secrets set NAME` store secret NAME (env/file/stdin, never argv) - `--from-env VAR`, read the value from an environment variable - `--from-file P`, read the value from a file - `--host HOST`, bind it to a host: sandboxes get a placeholder, replaced by the value only in requests sent to the address this command prints for HOST - `--header NAME`, with --host: the request header the key is sent in (default: the usual key headers) - `--no-raw`, with --host: the value itself can never be given to a sandbox - `--list`, list stored secret names (no values, ever) - `--delete`, remove NAME - `--json`, force the JSON envelope #### `promigence sandbox list` your sandboxes, newest first - `--limit N`, how many (default 50) - `--state S`, only these states, e.g. running,ready - `--label K=V`, only sandboxes whose run carries every label given - `--cursor C`, the next_cursor of an earlier page - `--json`, force the JSON envelope #### `promigence sandbox pause SID` park a sandbox: it holds no capacity and bills no seconds - `--json`, force the JSON envelope #### `promigence sandbox resume SID` bring a paused sandbox back and restart its clock - `--optimize-for start|compute`, default: as created - `--json`, force the JSON envelope #### `promigence sandbox snapshot SID` snapshot it without stopping it - `--json`, force the JSON envelope - `--alias NAME`, name it - `--move-alias`, take the alias from the snapshot that has it - `--force`, take it even though the sandbox was given a secret; the snapshot, and what starts from it, can then contain that value #### `promigence secrets list` stored secret names (never a value) - `--json`, force the JSON envelope #### `promigence secrets delete NAME` remove a stored secret - `--json`, force the JSON envelope #### `promigence sandbox info SID` one sandbox: state, tier, time left, cost so far, and its execs - `--json`, force the JSON envelope #### `promigence sandbox extend SID` give a live sandbox more time, within its run's quote and --cap (exit 5 past them) - `--timeout DUR` **(required)**, the new total timeout - `--json`, force the JSON envelope #### `promigence sandbox autopause SID` pause it once it has sat idle, or read the rule it is under - `--after DUR`, idle this long and it pauses itself (2s or more) - `--off`, never pause it on idle; only its timeout ends it - `--mode MODE`, requests (default) | activity: also no CPU, no bytes, no open connection inside - `--json`, force the JSON envelope #### `promigence sandbox probe SID` is this still the verified environment? (its recorded paths, or --path) - `--path P`, check these paths instead; PATH=blake3: where the snapshot recorded none (base, --image) - `--json`, force the JSON envelope #### `promigence files rm PATH` delete a path in a sandbox - `--sandbox SID` **(required)**, the sandbox - `--recursive`, a directory and everything under it - `--json`, force the JSON envelope ### Ops Accounts, keys, billing, organisations, connections and CI runners. #### `promigence bench burst` N sandboxes at once: hand-off percentiles + success rate - `--snapshot H` **(required)**, snap_ | alias[@commit7] | ./env.promigence - `--count N` **(required)**, number of sandboxes - `--timeout D`, per-sandbox timeout (default 120s, the public burst harness's) - `--cap USD`, spend cap for the burst (required unless PROMIGENCE_SPEND_CAP is set) - `--keep`, leave the sandboxes running (default: killed once measured) - `--json`, force the JSON envelope #### `promigence runners github` GitHub Actions jobs in sandboxes: `runs-on: promigence` (until Ctrl-C) - `--repo OWNER/REPO`, take this repository's jobs - `--org ORG`, or every repository's of an organization - `--runner-group NAME`, with --org: the runner group (Default) - `--token-env VAR`, variable holding a GitHub token (GH_TOKEN, then GITHUB_TOKEN) - `--labels A,B`, the runs-on labels (promigence); promigence-small|medium|large pick a size - `--tier small|medium|large`, size of a job that names none (small) - `--warm N`, keep N runners waiting: instant pickup, billed while idle - `--max M`, sandboxes at once (10) - `--cap USD`, most one job may cost (default: its timeout at its size) - `--timeout DUR`, longest job (default: your plan's longest sandbox) - `--idle-timeout DUR`, stop an on-demand runner no job took (90s) - `--kill-on-exit`, Ctrl-C cancels running jobs instead of waiting - `--prepare`, build the --tier runner snapshot, then exit - `--yes`, build a missing runner snapshot (the build price, per size) without asking - `--github-api URL`, GitHub Enterprise Server's API URL - `--json`, force the JSON envelope #### `promigence runners gitlab` GitLab CI jobs in sandboxes, as a runner with your tags (until Ctrl-C) - `--url URL`, your GitLab (https://gitlab.com) - `--token-env VAR`, variable holding the runner authentication token (GITLAB_RUNNER_TOKEN) - `--executor docker|shell`, jobs in containers (docker) or in the sandbox - `--default-image REF`, docker: image of a job with no image: (ubuntu:24.04) - `--pool N`, runners kept up, billed while idle (1) - `--api-token-env VAR`, read_api token: start runners only for pending jobs - `--max M`, sandboxes at once (10) - `--tier small|medium|large`, runner size (small) - `--cap USD`, most one job may cost (default: its timeout at its size) - `--timeout DUR`, longest job (default: your plan's longest sandbox) - `--idle-timeout DUR`, with --api-token-env: a runner no job took exits (90s) - `--kill-on-exit`, Ctrl-C cancels running jobs instead of waiting - `--prepare`, build the --tier runner snapshot, then exit - `--yes`, build a missing runner snapshot (the build price) without asking - `--json`, force the JSON envelope --- ## Errors and exit codes Every failure is attributed. An environment failure is ours: kept out of your result, retried once before hand-over, and not billed when our side can show it was ours. A task failure is yours: billed, because the environment worked. ### Exit codes | Exit | Meaning | Billing | |---|---|---| | 0 | Success | billed | | 1 | Internal error on our side | not billed | | 2 | Usage error, bad or missing arguments | not billed | | 3 | Environment failure, the environment was at fault, not your command | not billed when our side shows it was ours | | 4 | Task failure, your command exited non-zero | billed | | 5 | Refused: the quote exceeds your spend cap | not billed | | 6 | The snapshot is not verified | not billed | | 7 | Missing or invalid API key | not billed | | 8 | Not found | not billed | | 9 | Unreachable | not billed | | 10 | Rate limited | not billed | | 11 | A --wait expired. Nothing was killed | unchanged | | 12 | Partial, some environments did not start, or were stopped at the cap boundary | billed for what ran | | 130 | Interrupted by the client | billed for what ran | ### Error codes | Code | Category | HTTP | Exit | Billed | Meaning | |---|---|---|---|---|---| | fork_failed | env | 502 | 3 | no | The sandbox could not be started. | | liveness_timeout | env | 502 | 3 | no | The sandbox did not answer its first command within {deadline_ms} ms. | | probe_warm_failed | env | 502 | 3 | no | The sandbox failed its readiness check. | | warm_ratio_failed | env | 502 | 3 | no | Verify cpu_ms {verify_cpu_ms} exceeds max(2 x {baseline_cpu_ms}, {baseline_cpu_ms} + 250). | | verify_each_failed | env | 502 | 3 | no | Verify_each run failed in the fork (exit {exit_code}, digest {digest}). | | toolchain_mismatch | env | 502 | 3 | no | Toolchain check `{argv}` in the fork gave digest {got}, the verified snapshot recorded {expected}. | | guest_agent_lost | env | 502 | 3 | yes | The sandbox stopped responding. | | oom_within_tier | env | 502 | 3 | no | The sandbox was killed by a memory fault on our side while under the tier's {tier_mib} MiB. | | disk_full_within_tier | env | 502 | 3 | no | A storage fault on our side while the sandbox was under its {tier_disk_gib} GiB quota. | | node_lost | env | 502 | 3 | no | The host running the sandbox became unreachable. | | layer_integrity | env | 502 | 3 | no | The environment differs from the verified snapshot at {path}. | | egress_unavailable | env | 502 | 3 | no | The sandbox's network egress check failed. | | infra_timeout | env | 504 | 3 | yes | An internal deadline ({deadline_ms} ms) fired before your timeout. | | preempted | preempt | 502 | 3 | no | The sandbox was stopped because the capacity it was running on had to be given back. | | verify_did_not_cover | snapshot | 422 | 4 | yes | `{path}` is missing but was not under a verified path; the snapshot's verify did not cover it. | | grader_failed | grader | 422 | 3 | yes | The grade could not be computed ({kind}). | | nonzero_exit | task | 422 | 4 | no | Command exited {exit_code}. | | grade_not_passed | task | 422 | 4 | no | The task completed and its grade did not pass (the grader exited {exit_code}). | | task_timeout | task | 422 | 4 | yes | Command exceeded timeout_ms={timeout_ms}. | | oom_over_tier | task | 422 | 4 | no | The process was killed for exceeding the tier's {tier_mib} MiB of memory. | | disk_over_tier | task | 422 | 4 | no | The sandbox ran out of its {tier_disk_gib} GiB of disk. | | shutdown_by_task | task | 422 | 4 | yes | The sandbox shut itself down ({kind}). | | modified_by_task | task | 422 | 4 | yes | `{missing}` was removed or changed inside the sandbox; the snapshot's copy is intact. | | cap_killed | cap | 402 | 12 | yes | Sandbox killed when its accrued spend reached the spend cap ({spend_cap}). | | cap_exceeded | cap | 409 | 5 | no | Quote {max_cost} exceeds remaining spend cap {remaining}. | | aborted_by_client | other | 200 | 130 | yes | Sandbox killed by the client. | | expired | other | 200 | 12 | no | Sandbox timeout reached (counted from hand-off). | | not_started | other | 200 | 12 | no | Sandbox not started: outside the spend cap. | | superseded | other | 200 | 0 | no | Attempt 1 superseded by a successful retry. | | usage | other | 400 | 2 | no | {detail}. | | internal | other | 500 | 1 | no | Internal error. | | auth | other | 401 | 7 | no | Missing or invalid API key. | | forbidden | other | 403 | 7 | no | {detail}. | | not_found | other | 404 | 8 | no | {kind} `{id}` not found. | | unreachable | other | 503 | 9 | no | Cannot reach {url}. | | rate_limited | other | 429 | 10 | no | Rate limited by {host}; retry after {retry_after_s} s. | | timeout | other | 504 | 11 | no | Wait of {wait_timeout} expired; nothing was killed. | | not_verified | other | 409 | 6 | no | Snapshot {snapshot_id} is not verified (build {build_id}: {state}). | | pool_exhausted | env | 503 | 3 | no | No capacity free right now. | | quota_exhausted | env | 503 | 3 | no | {region} is at capacity right now. | | region_exhausted | env | 503 | 3 | no | {region} has no free capacity right now. | | builder_unavailable | env | 503 | 3 | no | New snapshots cannot be built right now. | | builder_busy | env | 503 | 3 | no | Snapshot builds are backed up; retry after {retry_after_s} s. | | overloaded | env | 503 | 3 | no | The API is at capacity; retry in {retry_after_s} s. | | tier_unavailable | env | 503 | 3 | no | The {tier} tier ({vcpus} vCPU / {mem_gib} GiB) is not available here yet. | | snapshot_unavailable | env | 409 | 3 | no | Snapshot {snapshot_id} is no longer available: the capacity that held it was retired before it was saved. | | capacity_limit | env | 409 | 3 | no | {requested} {tier} sandboxes at once is more than can run together; at most {max_at_once} at once. | | build_timeout | other | 422 | 4 | no | The build ran past its {limit_min}-minute limit. | | dockerfile_build_failed | other | 400 | 2 | yes | Your image could not be built from the Dockerfile: {detail}. | | image_unresolvable | other | 422 | 2 | yes | `{image}` could not be resolved to a digest: {detail}. | | registry_unavailable | env | 503 | 3 | no | `{image}` could not be resolved to a digest because its registry did not answer usably: {detail}. | | enoent | other | 404 | 8 | no | `{argv0}` not found in the sandbox. | | trial_credits_exhausted | cap | 402 | 5 | no | Trial credits left ({remaining}) do not cover this ({needed}). | | credit_unclaimed | other | 402 | 5 | no | This account has free credit waiting to be claimed. | | card_required | other | 402 | 5 | no | A card is required for new work since {effective_at}. | | billing_unavailable | other | 503 | 5 | no | Card payments are not open yet. | | concurrency_limit | other | 409 | 5 | no | {live} live + {requested} requested sandboxes is over your limit of {limit}. | | egress_cap_reached | other | 403 | 10 | no | Monthly data-out cap reached: {used_gb} of {cap_gb} GB. | | project_cap_reached | cap | 403 | 5 | no | Project {project} has {remaining} of its {limit} monthly cap left; this needs {needed}. | | spend_limit_reached | cap | 402 | 5 | no | Today's spend limit {limit} has {remaining} left; this needs {needed}. | | tier_over_plan | other | 403 | 2 | no | The {tier} tier ({vcpus} vCPU) is over the {limit} vCPU per-sandbox limit of your {plan} plan. | | network_host_refused | other | 403 | 9 | no | {host} is not on this sandbox's allowed list, so the connection was refused. | | network_policy_unsupported | env | 502 | 3 | yes | This sandbox asked to reach only {allow}, and that could not be applied where it was placed. | | secrets_in_snapshot | other | 409 | 2 | no | Sandbox {sandbox_id} was given a secret's value, so a snapshot of it can contain that value. | | secrets_unavailable | env | 503 | 3 | no | Secrets cannot be used right now. | | secret_binding_unsupported | env | 502 | 3 | yes | A secret bound to a host could not be set up for this sandbox. | | idempotency_key_reused | other | 409 | 2 | no | This Idempotency-Key was used for a different request. | | idempotency_key_in_use | env | 409 | 3 | no | A request with this Idempotency-Key is still running. | | repo_credentials_not_supported | other | 400 | 2 | yes | Credentials in repository URLs are not accepted ({field}). | | account_exists | other | 409 | 7 | no | This email address has an account that signs in with {sign_in_with}. | | invite_invalid | other | 403 | 7 | no | This invite code is not valid. | | invite_required | other | 403 | 7 | no | New accounts need an invite code right now. | | key_expired | other | 401 | 7 | no | This API key expired at {expired_at}. | --- ## Frequently asked questions ### Product **What is Promigence?** Promigence is a sandbox platform for AI agents: production-grade, fully isolated sandboxes that start in milliseconds, stay fast at a thousand at once, and bill by the second. It runs coding agents, background agents, evals and RL environments, each in its own sandbox built from your repository. **What is an environment runtime?** An environment runtime treats the execution environment as the product: it proves the environment works before anything runs in it, reproduces it exactly from a hash, and forks it many times over. A sandbox provider gives you a box that starts; an environment runtime gives you an environment that is known good. **What problem does Promigence solve?** Promigence solves silently broken environments. In a thousand-episode run some environments fail to install or come back cold, and that shows up as the agent failing the task, which corrupts the eval number and poisons the reward signal in RL. Promigence verifies every environment first and reports environment failures separately from task failures. **Who is Promigence for?** Promigence is for teams that run AI agents: agent companies building products on them, eval and platform engineers, benchmark producers and RL environment vendors. Any agent work fits, from one interactive sandbox to thousands of episodes at once. **Is Promigence an eval platform?** No. Promigence is the execution plane eval platforms run on, it has no datasets, scorers, judges or experiment tracking, which belong to your eval tools. It provides the verified environment the episode runs inside, and plugs in as a backend for those harnesses. **What is a validity report?** A validity report is what Promigence returns after every run: how many environments were verified, which failed, and whether each failure was the environment's or the task's. Most teams running evals today cannot produce that number, which means they cannot say how much of their result is real. **What is an episode in Promigence?** An episode is one environment's life in Promigence: fork, execute, terminate: 15 minutes by default, up to an hour in the private beta. Usage is metered per second, like the rest of the market, but the episode is the unit a quote and a spend cap are written against, so the maximum cost of a run is a multiplication you can do before you start it. **Does Promigence support long-running or interactive sandboxes?** Yes. A sandbox pauses when it goes idle, wakes on the next request, and bills only while it is awake. ### Comparisons **How is Promigence different from other sandbox providers?** Promigence is built around four things together: sandboxes that are ready in milliseconds and stay fast at a thousand at once, environments verified before an agent runs in them, a report that labels every failed run as the task's, the environment's or Promigence's, and a bill that is quoted and capped before the run starts. Runs that fail because of Promigence are not billed when our side can show the failure was ours. **Why not just run Docker on my own cloud machines?** Because a do-it-yourself setup cannot tell you how many of your environments were broken on the last run. What Promigence adds is verification, reproducibility, burst scheduling and nobody on call. ### Pricing **How much does Promigence cost?** Promigence meters per second at $0.0504 per vCPU-hour plus $0.0162 per GiB-hour, the market's standard list rate. That works out at $0.17 an hour for a small sandbox (2 vCPU, 4 GB), $0.46 for medium (4 vCPU, 16 GB) and $0.92 for large (8 vCPU, 32 GB). New accounts get $50 of free credit for 30 days: $10 straight away with no card, and $40 more once a card is added (card payments open after the private beta). **Does Promigence bill per second or per run?** Per second, at the market's standard list rate, because that is what the market meters and a price nobody can compare against is worth very little. The episode survives as the unit you are quoted in: `promigence run quote` returns the maximum a run can cost before it starts, and `--cap` refuses one that would go past your ceiling. **Can I know what a run will cost before it starts?** Yes. `promigence run quote` returns the exact maximum cost before any environment starts, and `--cap` refuses a run that would exceed your ceiling rather than stopping halfway through. Paused time is not invoiced, and nothing accrues before an environment is handed to you. **Is there a free trial?** Yes: $50 of free credit for 30 days, $10 straight away with no card and $40 more once you add one (card payments open after the private beta), with 20 concurrent sandboxes and sessions up to an hour. During the private beta, signing up needs an invite code. Separately, any team can send their environments in for 1,000 verified episodes, a validity report and an exact bill quote, free and once. ### Technical **How does Promigence work?** Promigence runs one primitive: verify, snapshot, fork, exec, report. You declare a container image, setup steps (a clone of your repository at a pinned commit among them) and a verify command. The snapshot is stored only if verification passes. You then copy it N times into warm environments and get back a validity report and a bill that never exceeds the quote you saw first. **How are Promigence environments isolated?** Each environment is fully isolated from every other, and nothing is shared with another customer. We test that boundary by attacking it rather than asserting it, and runs can start with no network at all. **How does Promigence make a fork start warm?** Promigence copies an environment that has already finished installing and building, instead of booting and installing it again. A copy takes 39 ms, and a hundred at once still land at 122 ms. **What does reproducible mean in Promigence?** The same snapshot hash returns a bit-identical environment on any later date, so an eval you ran in October can be re-run in March and the comparison means something. That is stronger than a recipe that re-resolves package versions underneath you on every build. **Does Promigence work with existing eval harnesses?** Yes. Adapters for common eval harnesses are in pre-release, and the TypeScript SDK exposes a compatibility `Sandbox` class with the familiar sandbox calls, so existing call sites run unchanged. **How are secrets handled?** You store a secret once and name it on the runs or commands that need it (`--secret NAME`). It is encrypted at rest, given only to the commands that name it, and removed from logs and outputs. It is never written into a snapshot you build, and a snapshot of a sandbox that was given a secret is refused unless you force it. **Can Promigence run inside my own cloud account?** Not as a self-serve product today. If your data cannot leave your boundary, talk to us about running it in your own cloud account. ### Reliability **What happens when an environment fails mid-run?** Promigence records the failure as the environment's in the validity report. A sandbox that fails before hand-over is retried once automatically; one that fails mid-run can be re-run with `promigence run replay`, as a new sandbox billed under the run's cap. A run with a 3% environment failure rate reports 3%, instead of quietly reporting a 3% worse agent. **How do I know the reliability numbers are real?** Every figure on the benchmarks page states its date, its sample size and the conditions it was measured under, and the free validity report runs your own workload so you can check the numbers yourself. ### Company **How do I get access to Promigence?** Promigence is in a private beta: signing up needs an invite code, and support@promigence.ai is where to ask for one. New accounts get $50 of free credit: $10 straight away with no card, and $40 more once a card is added (card payments open after the private beta). Teams who want a number on their own workload first can send their environments in for a free 1,000-episode run and a validity report. **Is Promigence tied to a particular model provider or lab?** No. Promigence is not owned by a lab or a hyperscaler and runs any vendor's agent. Neutrality is a requirement rather than a stance: a cross-lab evaluation cannot run inside one lab's harness. --- ## Contact Email support@promigence.ai. The fastest way to evaluate Promigence is the free validity report at https://www.promigence.ai/validity-report, send one repo or harness, get 1,000 verified episodes, the validity report, and an exact price for your own volume.