Skip to content

Find out how many of your environments are silently broken

Send the repo or harness you run today. We build the verified snapshot, run 1,000 episodes, and send back the validity report and an exact per-episode price. No card, no call.

  • 1,000 verified episodes on your own repo or harness
  • How many environments were actually valid, and why the rest were not
  • Environment failures separated from task failures, per episode
  • An exact per-episode price for your real volume

Request your validity report

A person replies, usually the same day. Or write to support@promigence.ai.

Four steps, and only the first one costs you anything

  1. 01

    You send one repo or harness

    A public repo, a private one with a deploy key, or the harness you already run. We use your setup command as-is. Hand-tuning it would defeat the point.

    5 minutes of your time

  2. 02

    We build a verified snapshot

    Setup runs, verify runs, and the snapshot is stored only if it passes. If it fails, that is already a finding.

    Usually under 10 minutes

  3. 03

    We run 1,000 episodes

    A thousand forks, your command in each, at full burst. Every copy is checked ready before your command runs, and every failure is categorised.

    Minutes, not hours

  4. 04

    You get the report and the quote

    How many environments were valid, which failed and why, and your real per-episode cost.

    Same day

A number about your environments, not about us

The report is the product. It is the thing your current provider does not give you, and it is the thing that decides whether the eval number you publish is real.

Want the definition first? What a validity report is

promigence run report, example output
text
run    run_8fa21c        snapshot  snap_2b9f…08fb4737
count  1000              wall      4m 12s

environments
  verified            1000 / 1000
  environment failures   0
  task failures         37   ← yours, not ours
  silently cold          0   (each one checked ready first)

cost
  quoted              $120.00 max
  billed              $118.40
  per episode         $0.1184

failed episodes kept for replay: 7 days
Language: text

The obvious questions

Is this actually free, or is it a trial that converts?

Free, and not a trial. If the report says your environments are fine and your provider is serving you well, that is a legitimate outcome and the conversation ends there. The offer exists because this is a problem most teams cannot see until someone measures it.

What do you need from me?

A repo URL and commit, your setup commands, and the command that counts as a pass. A read-only deploy key if it is private. No call, and no credentials for your existing provider.

What happens to my code?

Cloned, snapshotted, and deleted when you ask or after 30 days, whichever is first. We do not train on it, and we will sign a mutual NDA first if you prefer.

Can I run the benchmark myself instead?

Yes, and we would rather you did: sign up and run the same workload with your free credits. The offer just saves you the setup.

Join the waitlist

Promigence is in private beta. Leave your work email and we will send an invite as places open.