Sequential language-model interfaces discard useful computation, context, and execution state. Harness faults then look like model faults, while long tasks require recovery, delegation, resource control, and continuity across trajectories.
Prime Agent combines a persistent IPython REPL for recursive context processing, a Continual Harness that preserves histories, memories, skills, prompts, and subagent specifications, direct communication among recursive subagents, daemon-backed sessions, a human Agents View, verification, recovery, and explicit resource accounting. Strategy remains model-generated rather than hard-coded.
The authors report ARC-AGI-3 RHAE Best@1 rising from 30% to 95.5%. They say the harness matches or beats native and popular alternatives on long-context coding, GPU-kernel generation, emulator construction, and autonomous nanoGPT optimization; in Factorio, refinement sustains technology progression and dedicated subagents parallelize work.
See the full analysis.
Model leaderboards increasingly confound weights with context management, tools, budgets, and retries. Prime Agent makes that confound visible and gives researchers a system to vary.
The team authored the harness and evaluation report. Best@1 across allowed attempts is not single-run reliability. Task suites are heterogeneous, exact budget parity can be difficult, and persistent state can introduce contamination or hidden human intervention. The abstract does not establish causal attribution for each component.