How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Openai··Submitted by Mads Kristian Nylund
AI BenchmarkingAI InfrastructureAI Evaluation

The article highlights that enabling specific API settings, such as "retained reasoning" and "compaction," in the GPT-5.6 Sol model significantly improved its performance on the ARC-AGI-3 benchmark, increasing scores from 7.8% to 38.3% and reducing output tokens by 6x. These settings allowed the model to better retain reasoning and manage memory, addressing limitations in previous versions. The study emphasizes that evaluating AI models in isolation is insufficient and that factors like API settings and harness design play a critical role in determining performance.

Read Article

More from Openai

Related Articles