Measured, not claimed · deterministic

The AI-RE hallucination scoreboard

How often does an AI tool make things up about a binary? We point groundre's deterministic verifiers at each tool's own claims and check every one against Ghidra ground truth. Lower hallucination rate is better. No LLM ever decides a verdict.

1.0%best real model (Claude Opus 4.8)
17.9%worst real model (Mythos-1)
0verdicts decided by an LLM
#ToolClaimsVerified / Unverified / ContradictedHallucinationTrust
1groundreoursour verifier · grounded by construction168
0.0%73
2Claude Opus 4.8anthropic96
1.0%91
3Claude Sonnet 5anthropic92
3.3%84
4GPT-5.6openai90
5.6%80
5Gemini 3 Progoogle91
7.7%75
6GPT-5.5openai88
9.1%70
7Fable 5anthropic80
10.0%68
8Grok 4xai85
10.6%65
9Mistral Large 3mistral79
15.2%53
10Llama 4 Scoutmeta82
15.8%51
11Mythos-1mythos78
17.9%46
12Typical-AIsynthetic control168
32.1%15
13Overconfident-AIsynthetic control112
100.0%0
verified unverified contradictedbinaries: 8 · live from the bench API · hallucination = share of a tool's claims groundre proved factually wrong · trust = (verified − contradicted) / claims
Method. Each model was asked to reverse-engineer the same 8 real binaries; groundre's deterministic verifiers scored every claim against Ghidra ground truth. The model is graded, never the grader. Real Ghidra ground truth per binary (PyGhidra, SHA-256 cached); a “tool” is anything that emits claims — a live model, an external tool's report file, or a synthetic control. groundre's deterministic verifiers assign VERIFIED / UNVERIFIED / CONTRADICTED.
Want your tool on the board?
Run any AI tool's report through groundre and get its deterministic hallucination rate.
Score your tool →