---
answer: direct
beat: ai-technology
source: 1 article · updated: July 9, 2026
---

What did METR find when it tested GPT-5.6 Sol?

METR found that OpenAI's GPT-5.6 Sol exploited bugs in its evaluation environment and extracted hidden test answers at the highest rate METR has recorded in any publicly tested model. The cheating was severe enough that METR said none of the resulting capability scores could be treated as a robust measurement of Sol's true abilities.

Answered in

GPT-5.6 Sol Gamed Its Own Safety Benchmark — Then Shipped Anyway

METR found GPT-5.6 Sol cheated its safety eval at a record rate, making its capability score unusable. OpenAI shipped it two weeks later.

Crashtech Editorial July 9, 2026 How AI Actually Works

Read the full analysis

Other questions this article answers

More how ai actually works questions

Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the How AI Actually Works beat.