How many tasks are in the Kotlin Benchmark, and where do they come from?
The first public iteration contains 105 engineering tasks sourced from active open-source Kotlin repositories. Each task requires an agent to read a real issue description, navigate the existing codebase, and produce a patch verified inside a containerized environment — resolved only if the generated solution passes the required tests.
Answered in
JetBrains Built Its Own AI Coding Benchmark Because It Doesn't Trust Anyone Else'sJetBrains released a 105-task, open Kotlin coding benchmark on July 8, 2026 — and Claude Code beat JetBrains' own Junie agent by 3.81 points at launch.
Read the full analysisOther questions this article answers
More development best practices questions
- How many outage reports did Claude and ChatGPT get on July 14, 2026?
- What did Anthropic's own status page say about the July 14 Claude outage?
- Did Claude have more outages after July 14?
- What did ChatGPT's status checker say was wrong?
- Is this outage pattern actually unusual for AI providers?
- What GitHub Actions vulnerability did the attacker exploit?
- How many npm packages were compromised, and how widely were they used?
- What did the malicious payload actually do?
Every answer on Crashtech is written by the editor of the article it comes from — never auto-summarised. Browse all answers or the Development Best Practices beat.