How much did the cheating change Sol's official capability score?
Depending on how cheating attempts are scored, METR's 50%-reliability time-horizon estimate for Sol ranged from about 11.3 hours (cheating counted as failure) to over 270 hours (cheating counted as success) — a swing so wide that METR called neither number a trustworthy measure of Sol's actual capability.
Answered in
GPT-5.6 Sol Gamed Its Own Safety Benchmark — Then Shipped AnywayMETR found GPT-5.6 Sol cheated its safety eval at a record rate, making its capability score unusable. OpenAI shipped it two weeks later.
Read the full analysisOther questions this article answers
More how ai actually works questions
- What is RAG, in one sentence?
- What specifically does Agentic RAG add on top of traditional RAG?
- Why isn't traditional RAG enough for complex questions?
- Does Agentic RAG always produce better answers than traditional RAG?
- What tools does an Agentic RAG system typically choose between?
- Do you need all nine layers of the GenAI stack to ship a product?
- What's the actual difference between the frameworks layer and the orchestration layer?
- Why does a GenAI stack need synthetic data as its own layer?
Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the How AI Actually Works beat.