---
title: "RL environment infrastructure, verified environments at scale"
description: "Promigence runs reinforcement-learning environments for coding agents: verified before every episode, forked warm by the thousand, and quoted and capped before the run starts."
url: "https://www.promigence.ai/for/rl-environments"
site: "Promigence"
---

# Reinforcement-learning environments that cannot poison the reward

Promigence runs thousands of concurrent RL environments forked from one verified snapshot, each proven to work before the episode starts. A broken environment does not just lose you an episode: the optimiser treats its failure as real and the model learns from something that never happened.

**Who this is for:** RL environment vendors, data vendors, and labs' own post-training teams running coding agents.

## What breaks today
### A broken environment is a poisoned reward
The agent writes a correct patch, the environment failed to install a dependency, the test fails, and the optimiser updates on it. Unlike a wrong eval chart, this one is baked into weights.

### Concurrency ceilings you discover in production
Published accounts put real fleets capping at 300 to 550 concurrent, with orchestration the bottleneck past 256. A training step that needs 2,000 does not degrade gracefully, it stalls.

### Per-second billing on millions of episodes
At millions of episodes, cold setup is most of the bill, and you are paying it to run the same install over and over.

## What changes with Promigence
### Verification is a precondition, not a hope
No snapshot is stored that cannot pass its own test, and no fork is handed over cold. Environment failures are reported separately, so they never enter the reward signal.

### Fan-out designed for the training step, not the demo
A thousand forks for one training step start warm, not as a thousand cold installs.

### Volume pricing on committed capacity
Committed-use rates for sustained volume. Talk to us about running it in your own cloud account.

## A training step: environments verified, fanned out, capped
```python
from promigence import Promigence

client = Promigence()   # reads PROMIGENCE_API_KEY

# The most one training step can cost, before anything starts
quote = client.runs.quote("task-suite", count=20, spend_cap=1.0)   # 20 at once in the private beta

# Fan out the step: copies of one verified snapshot, each checked first
run = client.runs.start(
    "task-suite", count=20, spend_cap=1.0, verify_each=True,
    exec={"sh": f"python -m rollout --policy {policy_url}"},
)

# Environment failures apart from task failures, and what the step cost
report = client.runs.report(run["run_id"])
print(report["totals"], report["cost"])
```

