How often does an AI tool make things up about a binary? We point groundre's deterministic verifiers at each tool's own claims and check every one against Ghidra ground truth. Lower hallucination rate is better. No LLM ever decides a verdict.
| # | Tool | Claims | Verified / Unverified / Contradicted | Hallucination | Trust |
|---|---|---|---|---|---|
| 1 | groundreoursour verifier · grounded by construction | 168 | 0.0% | 73 | |
| 2 | Claude Opus 4.8anthropic | 96 | 1.0% | 91 | |
| 3 | Claude Sonnet 5anthropic | 92 | 3.3% | 84 | |
| 4 | GPT-5.6openai | 90 | 5.6% | 80 | |
| 5 | Gemini 3 Progoogle | 91 | 7.7% | 75 | |
| 6 | GPT-5.5openai | 88 | 9.1% | 70 | |
| 7 | Fable 5anthropic | 80 | 10.0% | 68 | |
| 8 | Grok 4xai | 85 | 10.6% | 65 | |
| 9 | Mistral Large 3mistral | 79 | 15.2% | 53 | |
| 10 | Llama 4 Scoutmeta | 82 | 15.8% | 51 | |
| 11 | Mythos-1mythos | 78 | 17.9% | 46 | |
| 12 | Typical-AIsynthetic control | 168 | 32.1% | 15 | |
| 13 | Overconfident-AIsynthetic control | 112 | 100.0% | 0 |