GPT-5.6 Sol’s weak ARC-AGI-3 benchmark result was largely a harness problem: retaining reasoning and enabling compaction raised its public-set score from 13.3% to 38.3%. The same changes also reduced output-token use by roughly six times. Source
Read the full article at the source.
Comments (0)
No comments yet. Be the first to comment!