GPT-5.6 Sol’s weak ARC-AGI-3 benchmark result was largely a harness problem: retaining reasoning and enabling compaction raised its public-set score from 13.3% to 38.3%. The same changes also reduced output-token use by roughly six times. Source
Læs hele artiklen hos kilden.
Kommentarer (0)
Ingen kommentarer ennå. Bli den første til å kommentere!