Skip to main content
AI

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

OpenAI said on Tuesday that enabling two settings in its Responses API tripled the score of its GPT-5.6 Sol model on the ARC-AGI-3 benchmark. The company published the findings in a research post dated July 29, 2026. With the official benchmark harness, GPT-5.6 Sol scored 13.3% on the public task set. After turning on retained reasoning and compaction, the score rose to 38.3% while output tokens dropped by a factor of six. The benchmark measures Relative Human Action Efficiency against an estimated human baseline of 48%. OpenAI said the official harness discarded private reasoning after each action and used a rolling truncation window that removed older context. The company said its models are trained to think with private reasoning messages that are retained in conversation history. Retaining that reasoning let the model spend less time thinking per action and employ coherent strategies over time. Replacing rolling truncation with compaction preserved learned information across longer runs. The company noted that GPT-5.6 Sol has solved open mathematics problems and beaten games such as Pokemon FireRed. On the ARC-AGI-3 leaderboard, no other frontier model solves any level beyond the first, while GPT-5.6 Sol with the modified harness solves all six levels.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from OpenAI and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Products
AI Products

OpenAI says it will not pursue an IPO in 2026

OpenAI CEO Sam Altman confirmed that the company will not go public in 2026, according to Fortune. Altman cited the safety situation as making an offering inadvisable at present. The decision delays a potential public listing that...

AI Products
AI Products

Atlassian rolls out code context for AI coding agents

Atlassian said its Code Context feature is in early access and gradually rolling out through an open beta. The feature indexes codebases into the Teamwork Graph and lets tools including Cursor, Claude Code and Codex retrieve code ...