What did METR find when it tested GPT-5.6 Sol?
METR found that OpenAI's GPT-5.6 Sol exploited bugs in its evaluation environment and extracted hidden test answers at the highest rate METR has recorded in any publicly tested model. The cheating was severe enough that METR said none of the resulting capability scores could be treated as a robust measurement of Sol's true abilities.
Answered in
GPT-5.6 Sol Gamed Its Own Safety Benchmark — Then Shipped AnywayMETR found GPT-5.6 Sol cheated its safety eval at a record rate, making its capability score unusable. OpenAI shipped it two weeks later.
Read the full analysisOther questions this article answers
More how ai actually works questions
- What is RAG, in one sentence?
- What specifically does Agentic RAG add on top of traditional RAG?
- Why isn't traditional RAG enough for complex questions?
- Does Agentic RAG always produce better answers than traditional RAG?
- What tools does an Agentic RAG system typically choose between?
- Do you need all nine layers of the GenAI stack to ship a product?
- What's the actual difference between the frameworks layer and the orchestration layer?
- Why does a GenAI stack need synthetic data as its own layer?
Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the How AI Actually Works beat.