Skip to content

Reinforcement-learning environments that cannot poison the reward

Promigence runs thousands of concurrent RL environments forked from one verified snapshot, each proven to work before the episode starts. A broken environment does not just lose you an episode: the optimiser treats its failure as real and the model learns from something that never happened.

Who this is for: RL environment vendors, data vendors, and labs' own post-training teams running coding agents.

177 ms

at 200 concurrent, all served

39 ms

warm fork on the box

108,929

forks in six hours, zero failures

Three things that are true of almost every team doing this

01

A broken environment is a poisoned reward

The agent writes a correct patch, the environment failed to install a dependency, the test fails, and the optimiser updates on it. Unlike a wrong eval chart, this one is baked into weights.

02

Concurrency ceilings you discover in production

Published accounts put real fleets capping at 300 to 550 concurrent, with orchestration the bottleneck past 256. A training step that needs 2,000 does not degrade gracefully, it stalls.

03

Per-second billing on millions of episodes

At millions of episodes, cold setup is most of the bill, and you are paying it to run the same install over and over.

With Promigence underneath

  • Verification is a precondition, not a hope

    No snapshot is stored that cannot pass its own test, and no fork is handed over cold. Environment failures are reported separately, so they never enter the reward signal.

  • Fan-out designed for the training step, not the demo

    A thousand forks for one training step start warm, not as a thousand cold installs.

  • Volume pricing on committed capacity

    Committed-use rates for sustained volume. Talk to us about running it in your own cloud account.

A training step: environments verified, fanned out, capped
Python
from promigence import Promigence

client = Promigence()   # reads PROMIGENCE_API_KEY

# The most one training step can cost, before anything starts
quote = client.runs.quote("task-suite", count=20, spend_cap=1.0)   # 20 at once in the private beta

# Fan out the step: copies of one verified snapshot, each checked first
run = client.runs.start(
    "task-suite", count=20, spend_cap=1.0, verify_each=True,
    exec={"sh": f"python -m rollout --policy {policy_url}"},
)

# Environment failures apart from task failures, and what the step cost
report = client.runs.report(run["run_id"])
print(report["totals"], report["cost"])
Language: python

Full quickstart

Run your own workload on it, free

Send a repo and the command you run against it, whatever that is: an eval suite, an RL rollout, a CI job, a queue of coding tasks. We build the environment once, run it a thousand times, and send back the timings, the failures and an exact price. Free, once, on your real workload.

  • 1,000 runs of your own command, on your own repo
  • What each one cost in wall clock, and anything that failed
  • Whether a failure was your code or the environment
  • An exact price for your real volume

A person replies, usually the same day. Or write to support@promigence.ai.

Not ready to hand over a repo? Read the quickstart or check the numbers first.

Join the waitlist

Promigence is in private beta. Leave your work email and we will send an invite as places open.