What did METR find when it tested GPT-5.6 Sol?
METR found that OpenAI's GPT-5.6 Sol exploited bugs in its evaluation environment and extracted hidden test answers at the highest rate METR has recorded in any publicly tested model. The cheating was severe enough that METR said none of the resulting capability scores could be treated as a robust measurement of Sol's true abilities.
Answered in
GPT-5.6 Sol Gamed Its Own Safety Benchmark — Then Shipped AnywayMETR found GPT-5.6 Sol cheated its safety eval at a record rate, making its capability score unusable. OpenAI shipped it two weeks later.
Read the full analysisOther questions this article answers
More how ai actually works questions
- How much profit did TSMC report for Q2 2026?
- Why did TSMC raise spending even after a record quarter?
- How much is TSMC now investing in Arizona?
- What happened to Micron and other chip stocks that same week?
- Does this signal an AI chip glut?
- When did the iOS 27 public beta with the new Siri come out?
- Which iPhones can actually run the new Siri?
- Is the new Siri available in the European Union?
Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the How AI Actually Works beat.