# How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

_Wednesday, July 29, 2026 at 8:00 AM EDT · AI · Latest · Tier 2 — Notable_

OpenAI said on Tuesday that enabling two settings in its Responses API tripled the score of its GPT-5.6 Sol model on the ARC-AGI-3 benchmark. The company published the findings in a research post dated July 29, 2026.

With the official benchmark harness, GPT-5.6 Sol scored 13.3% on the public task set. After turning on retained reasoning and compaction, the score rose to 38.3% while output tokens dropped by a factor of six. The benchmark measures Relative Human Action Efficiency against an estimated human baseline of 48%.

OpenAI said the official harness discarded private reasoning after each action and used a rolling truncation window that removed older context. The company said its models are trained to think with private reasoning messages that are retained in conversation history. Retaining that reasoning let the model spend less time thinking per action and employ coherent strategies over time. Replacing rolling truncation with compaction preserved learned information across longer runs.

The company noted that GPT-5.6 Sol has solved open mathematics problems and beaten games such as Pokemon FireRed. On the ARC-AGI-3 leaderboard, no other frontier model solves any level beyond the first, while GPT-5.6 Sol with the modified harness solves all six levels.

## Sources

- [OpenAI](https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores)

---
Canonical: https://techandbusiness.org/newswire/QDH94Q5CU3MgPoFUMGFpvU
Retrieved: 2026-07-30T03:04:03.286Z
Publisher: Tech & Business (techandbusiness.org)
