hub
Prime Agent benchmarks: self-reported result ledger
Prime Intellect reported ARC-AGI-3, long-context, EmulatorBench, PMPP-Hard, Factorio, and MazeBench results in its launch article. Most numbers are self-reported and lack publicly available prompts, raw runs, variance, or hardware details. This hub lists what is inspectable, what is missing, and links to methodology.
v v0.7.1reviewed 2026-08-09evidence self-reportedsources S001, S037cutoff 2026-08-09
Benchmark or case-study claim made by Prime Intellect without independent reproduction in this corpus.
Self-reported results
Reported scores remain labeled self-reported unless the public record includes enough artifacts for independent reproduction.
Available results
| Suite | Claim | Evidence class | Gap |
|---|---|---|---|
| ARC-AGI-3 | 95.5% RHAE Best@1 | Self-reported | No raw runs or variance published |
| Long-context suite | Mixed win/lose rows | Self-reported | Some competitor figures are reused |
| Methodology | N/A | Methodology page | Missing prompts and hardware on most rows |
Missing evidence
- complete prompts and model snapshots,
- raw runs and variance,
- hardware and budget details,
- independent reproductions.
Methodology and limitations
Use the methodology page before comparing results, and follow the source chain rather than quoting a headline number in isolation.