An RL environment where an agent reads a client brief, researches, makes a truthful & affordable offer, and ships a real deliverable — and only gets paid for honest, complete work.
All figures baked from canonical results · real
Coding benchmarks score one thing: does it pass? A business engagement has to be right on four axes at once — and the reward refuses to pay for promises.
The offer must cover everything the client actually asked for — no more, no less.
Price has to land inside the client's budget band. Overcharge → rejected. Undercharge → money left on the table.
No claims the business can't back up. Promise SOC2 it doesn't have → the deal is policy-failed.
A great pitch with no artifact earns almost nothing. The delivery gate halves offer credit and zeros efficiency.
“You only get paid for honest, complete work” isn't a slogan — it's coded into the grader.
HUD v6 environment: 7 FastMCP tools, a hybrid grader, and GRPO training on the Tinker backend. Click any stage.
A real scenario — easy_ticket_triage, a $100 ticket-classifier for a clinic. The inputs are real; the score gauges are illustrative.
Pick one of the 30 — or paste your own scenario JSON — and run it through the real grader. You get the reward, how long it took, and exactly where it needs to improve. …
A deterministic majority anchors every score; the LLM judge (claude-haiku-4-5) only moves the quality slice. That pairing is our answer to “how do you know the judge isn't gamed?”
On-policy GRPO on the model's own rollouts. The after-eval covers all 30 scenarios — including the 24 it never trained on — so this is generalization, not memorization.
Frontier reference vs. the open model before and after RL. real · HUD eval jobs
9 business domains × 3 difficulty tiers. Add a scenario by dropping a JSON file — no code change. Click any card.
GRPO only learns when attempts in a group differ. The base model's high variance (std 0.289) is exactly the spread RL needs to push on.
Mean ± standard deviation across 30 scenarios.
Average sub-score by criterion (sample export).
A degraded shared pool doesn't slow you down — it fabricates wrong numbers. So honesty is coded in.
Any baseline below a floor set from prior clean runs is rejected — a low junk baseline can't fake a huge “improvement.”
The after-eval must have ≥90% of rollouts actually graded, or it's rejected and retried — an outage can't bias the verdict.
Rejected numbers go to quarantine.jsonl for transparency — never the canonical file.
The first three are captured live from real commands (scripts/make_receipts.py); the rest are real HUD-platform jobs. Click any screenshot to enlarge.