⬡ HUD × YC FRONTIER RL HACKATHON · AUTONOMOUS BUSINESS TRACK

SWE-bench taught models to code.
Agent Inc. teaches them to run a business.

An RL environment where an agent reads a client brief, researches, makes a truthful & affordable offer, and ships a real deliverable — and only gets paid for honest, complete work.

Explore the build
Open model · base
0.327
Qwen3.5-4B
→
After RL
0.647
16 GRPO steps
→
Improvement
+0.321
closed ~½ the gap to Claude

All figures baked from canonical results · real

The problem · why it's novel

Running a business is a harder RL task than it looks

Coding benchmarks score one thing: does it pass? A business engagement has to be right on four axes at once — and the reward refuses to pay for promises.

In scope

The offer must cover everything the client actually asked for — no more, no less.

Affordable

Price has to land inside the client's budget band. Overcharge → rejected. Undercharge → money left on the table.

Honest

No claims the business can't back up. Promise SOC2 it doesn't have → the deal is policy-failed.

Delivered

A great pitch with no artifact earns almost nothing. The delivery gate halves offer credit and zeros efficiency.

“You only get paid for honest, complete work” isn't a slogan — it's coded into the grader.

How it's built · technology

One scenario JSON in. A graded reward out.

HUD v6 environment: 7 FastMCP tools, a hybrid grader, and GRPO training on the Tinker backend. Click any stage.

The agent loop · live replay

Watch an engagement, step by step

A real scenario — easy_ticket_triage, a $100 ticket-classifier for a clinic. The inputs are real; the score gauges are illustrative.

QUEST: Northstar Clinic · ticket triage
budget $100
HONEST DEAL — OFFER ACCEPTED
The delivery gate: if the agent had sent that offer but never submitted a deliverable, pricing & policy credit are halved and efficiency drops to 0. Promises don't pay.
TOTAL REWARD
0.00
Try it yourself · interactive

Run a business scenario, scored live

Pick one of the 30 — or paste your own scenario JSON — and run it through the real grader. You get the reward, how long it took, and exactly where it needs to improve. …

▶ Pick a scenario and hit Run & score.

The reward · how the judge can't be fooled

70% deterministic Python · 30% LLM judge

A deterministic majority anchors every score; the LLM judge (claude-haiku-4-5) only moves the quality slice. That pairing is our answer to “how do you know the judge isn't gamed?”

Deterministic · key-free
0.70
completeness · pricing · efficiency · policy
LLM judge
0.30
quality (claude-haiku-4-5)
Reinforcement learning · the headline

The open model leveled up

On-policy GRPO on the model's own rollouts. The after-eval covers all 30 scenarios — including the 24 it never trained on — so this is generalization, not memorization.

baseline GRPO step (training subset) final eval · all 30
Level 1 · rookie
0.327
▼ +0.321 XP
Level 2 · operator
0.647

Where it lands

Leaderboard

Frontier reference vs. the open model before and after RL. real · HUD eval jobs

The environment · data-driven

A universe of 30 client engagements

9 business domains × 3 difficulty tiers. Add a scenario by dropping a JSON file — no code change. Click any card.

Score distributions

Variance is the fuel

GRPO only learns when attempts in a group differ. The base model's high variance (std 0.289) is exactly the spread RL needs to push on.

Reward spread per model real

Mean ± standard deviation across 30 scenarios.

Per-criterion: frontier vs. open illustrative

Average sub-score by criterion (sample export).

Result integrity · completion proof

We guard the numbers, not just the uptime

A degraded shared pool doesn't slow you down — it fabricates wrong numbers. So honesty is coded in.

Plausibility floor

Any baseline below a floor set from prior clean runs is rejected — a low junk baseline can't fake a huge “improvement.”

Graded-fraction guard

The after-eval must have ≥90% of rollouts actually graded, or it's rejected and retried — an outage can't bias the verdict.

Quarantine, don't delete

Rejected numbers go to quarantine.jsonl for transparency — never the canonical file.

Receipts — real runs, real jobs

The first three are captured live from real commands (scripts/make_receipts.py); the rest are real HUD-platform jobs. Click any screenshot to enlarge.

—