The free judge that guessed instead of saying "I don't know"
Yad's free-provider pool just grew to 10+, including a free local model that can now act as the "judge" that checks whether each browser step actually did what it was supposed to. Before trusting it, we asked a blunt question: when it doesn't have enough evidence, does it admit that, or does it guess? First answer: it guessed. Here's the test, the guess, and the one-paragraph fix, verified before and after.
Every action Yad takes gets checked by a small second opinion, the "judge": given what the step was supposed to accomplish and what evidence came back, classify the outcome as match, mismatch, or unknown. Unknown is the safe answer when there isn't enough evidence to decide either way — it's supposed to make the agent pause and ask for help rather than plow ahead on a guess. That check used to always go to a paid-tier-quality cloud model. It now can run on a small, free, local model instead, which matters because it means the check can happen far more often without burning through the free quota the actual planning needs.
A cheaper judge is only worth using if it's still an honest one. So before turning it on more broadly, we built a side-by-side test: run Yad's real benchmark tasks, and every time the judge actually gets invoked, send the exact same evidence to both the cheap local model and a strong cloud model, then log both verdicts next to each other. No synthetic test cases — only real judge calls that happened during real runs.
First run: 2 out of 7 agreed
Twenty-five tasks, but only seven produced a real judge call — most of this benchmark set is pure reading (extract this, count that), which doesn't need step verification. Of those seven, the cheap and strong model agreed on two. Not once did the cheap model say "match" while the strong model said "mismatch" — the worst possible failure never happened. But a clear pattern did: in five of the seven cases, the action had returned no extracted evidence at all, just an unchanged URL. The strong model correctly said "unknown" every time — nothing to go on, so no verdict. The cheap model, given that same nothing, still said "mismatch" four times and, once, "match." A guess dressed up as a verdict.
The fix was one paragraph
The judge's instructions already said, in prose, that unknown is the answer when evidence is "absent, unclear, or too ambiguous." A strong model reliably applies that. A small model, it turned out, did not — it needed the rule spelled out as a mechanical check rather than left to its own judgment call. The fix in packages/companion/src/judge/judge.ts adds one hard rule ahead of everything else: if there's no extracted evidence, the answer is unknown, full stop, regardless of what the URL looks like. No new infrastructure, no new model, one paragraph of prompt.
Second run, same tasks: 6 out of 7
Same benchmark set, same real judge moments, after the fix. Agreement went from 2/7 to 6/7. The one remaining disagreement leans safe, not dangerous: the cheap model said unknown where the strong model was confident enough to say match — overly cautious, never overconfident. Zero dangerous disagreements in either run.
| Run | Agreement | Dangerous (cheap=match, strong=mismatch) |
|---|---|---|
| Before fix | 2/7 (28.6%) | 0 |
| After fix | 6/7 (85.7%) | 0 |
What this does and doesn't prove
Seven comparisons is a real number, not a big one — it's what this particular task set actually produced, not a number we picked. It's enough to catch and fix a genuine overconfidence pattern, and enough to say the free local judge is no longer guessing on thin evidence in this sample. It is not enough to call the free judge equivalent to the paid one across every situation Yad runs into. The honest next step, if this becomes a broader default, is a larger run specifically on tasks that produce rich extracted evidence, not just this set's mostly-read, mostly-thin-evidence mix — that's a deliberately separate test from this one, not a reason to hold back what got fixed here.