The Information just broke the scoop that OpenAI is playing around with a new technique, in which models will reveal less of their “thinking”, making them harder to monitor.As Zack Korman and I argued here a few days ago, better monitoring is one of the things that might have prevented the Hugging Face incident, by OpenAI’s own admission:The new techniques they are exploring may make such monitoring difficult or impossible. §Last year, an all-star cast wrote a fascinating paper that feels deeply relevant now, called Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety, I fully agree with the highlighted bit:They are exactly right. CoT monitoring is imperfect (as Subbarao Kambhampati and others...
Les hele artikkelen hos kilden.
Kommentarer (0)
Ingen kommentarer ennå. Bli den første til å kommentere!