Skip to content

What is poisoned reward?

A poisoned reward is a training signal produced by a broken environment rather than by the agent's behaviour, so the model learns from an outcome that never happened.

If an environment fails to install a dependency, a correct patch still fails the test, the harness records a failure, and the optimiser updates on it. Nothing in the pipeline separates 'the agent was wrong' from 'the environment was broken'.

A wrong eval number is embarrassing and recoverable. A poisoned reward is baked into weights and costs a training run to undo. The defence is to verify before the episode and attribute every failure in the report.

This is one of the terms in the Promigence glossary.

Run your own workload on it, free

Send a repo and the command you run against it, whatever that is: an eval suite, an RL rollout, a CI job, a queue of coding tasks. We build the environment once, run it a thousand times, and send back the timings, the failures and an exact price. Free, once, on your real workload.

  • 1,000 runs of your own command, on your own repo
  • What each one cost in wall clock, and anything that failed
  • Whether a failure was your code or the environment
  • An exact price for your real volume

A person replies, usually the same day. Or write to support@promigence.ai.

Not ready to hand over a repo? Read the quickstart or check the numbers first.

Join the waitlist

Promigence is in private beta. Leave your work email and we will send an invite as places open.