hub

Prime Agent benchmarks: self-reported result ledger

Prime Intellect reported ARC-AGI-3, long-context, EmulatorBench, PMPP-Hard, Factorio, and MazeBench results in its launch article. Most numbers are self-reported and lack publicly available prompts, raw runs, variance, or hardware details. This hub lists what is inspectable, what is missing, and links to methodology.

v v0.7.1reviewed 2026-08-09evidence self-reportedsources S001, S037cutoff 2026-08-09
Install Prime Agent

Benchmark or case-study claim made by Prime Intellect without independent reproduction in this corpus.

Self-reported results

Reported scores remain labeled self-reported unless the public record includes enough artifacts for independent reproduction.

Available results

SuiteClaimEvidence classGap
ARC-AGI-395.5% RHAE Best@1Self-reportedNo raw runs or variance published
Long-context suiteMixed win/lose rowsSelf-reportedSome competitor figures are reused
MethodologyN/AMethodology pageMissing prompts and hardware on most rows

Missing evidence

  • complete prompts and model snapshots,
  • raw runs and variance,
  • hardware and budget details,
  • independent reproductions.