Reinforcement-learning environments that cannot poison the reward
Promigence runs thousands of concurrent RL environments forked from one verified snapshot, each proven to work before the episode starts. A broken environment does not just lose you an episode: the optimiser treats its failure as real and the model learns from something that never happened.
Who this is for: RL environment vendors, data vendors, and labs' own post-training teams running coding agents.
at 200 concurrent, all served
warm fork on the box
forks in six hours, zero failures
Three things that are true of almost every team doing this
A broken environment is a poisoned reward
The agent writes a correct patch, the environment failed to install a dependency, the test fails, and the optimiser updates on it. Unlike a wrong eval chart, this one is baked into weights.
Concurrency ceilings you discover in production
Published accounts put real fleets capping at 300 to 550 concurrent, with orchestration the bottleneck past 256. A training step that needs 2,000 does not degrade gracefully, it stalls.
Per-second billing on millions of episodes
At millions of episodes, cold setup is most of the bill, and you are paying it to run the same install over and over.
With Promigence underneath
Verification is a precondition, not a hope
No snapshot is stored that cannot pass its own test, and no fork is handed over cold. Environment failures are reported separately, so they never enter the reward signal.
Fan-out designed for the training step, not the demo
A thousand forks for one training step start warm, not as a thousand cold installs.
Volume pricing on committed capacity
Committed-use rates for sustained volume. Talk to us about running it in your own cloud account.
from promigence import Promigence
client = Promigence() # reads PROMIGENCE_API_KEY
# The most one training step can cost, before anything starts
quote = client.runs.quote("task-suite", count=20, spend_cap=1.0) # 20 at once in the private beta
# Fan out the step: copies of one verified snapshot, each checked first
run = client.runs.start(
"task-suite", count=20, spend_cap=1.0, verify_each=True,
exec={"sh": f"python -m rollout --policy {policy_url}"},
)
# Environment failures apart from task failures, and what the step cost
report = client.runs.report(run["run_id"])
print(report["totals"], report["cost"])
Language: pythonRun your own workload on it, free
Send a repo and the command you run against it, whatever that is: an eval suite, an RL rollout, a CI job, a queue of coding tasks. We build the environment once, run it a thousand times, and send back the timings, the failures and an exact price. Free, once, on your real workload.
- 1,000 runs of your own command, on your own repo
- What each one cost in wall clock, and anything that failed
- Whether a failure was your code or the environment
- An exact price for your real volume
Not ready to hand over a repo? Read the quickstart or check the numbers first.