AI-assisted software engineering has seen the emergence of several benchmarks to measure the capabilities of LLMs. Android developers face specific challenges that aren't covered by existing benchmarks, so we created one that focuses on a north star of high quality Android development.
Model Score (%) info Average percentage of 100 test cases successfully resolved across 5 runs for each model
arrow_range Cl range (%) info Expected performance range, reflecting the results' statistical reliability (p-value < 0.05)
Avg latency (h) info Average time taken to solve 100 tasks across 5 runs
Avg cost ($) info Average cost per full benchmark run
Claude Opus 5 91.8
87.4 — 95.7 13.2 $175.4
Claude Fable 5 90.9
86.2 — 95.4 8.7 $144.3
GPT 5.6 Sol 90.8
85.6 — 95.6 14.2 $150.0
Kimi K3 90.2
85.0 — 94.6 29.6 $112.9
GPT 5.6 Luna 87.6
81.8 — 92.3 10.6 $7.2
Qwen3.8 Max 87.0
82.2 — 91.6 41.2 $196.7
GPT 5.6 Terra 86.8
81.7 — 91.2 7.6 $48.0
GPT 5.5 80.2
73.0 — 86.6 11.4 $138.3
Claude Sonnet 5 76.2
68.8 — 82.5 12.3 $99.9
Gemini 3.6 Flash 75.6
67.5 — 82.1 10.2 $135.5
Gemini 3.1 Pro Preview 74.3
66.8 — 81.5 10.5 $86.6
GPT 5.4 74.1
66.6 — 81.2 8.4 $83.4
Claude Opus 4.8 72.4
65.1 — 79.2 6.7 $88.0
GLM 5.2 72.2
65.3 — 79.1 38.9 $117.0
Gemini 3.5 Flash 71.1
63.8 — 78.5 28.3 $165.6
Kimi K2.7 Code 70.4
63.4 — 77.1 31.8 $48.1
Claude Opus 4.7 68.7
60.9 — 76.4 7.0 $96.5
Kimi K2.6 67.6
59.8 — 74.8 57.2 $49.4
Claude Sonnet 4.6 67.0
58.5 — 75.1 16.9 $127.6
Minimax M3 63.6
56.1 — 70.7 26.0 $41.7
GLM 5.1 63.2
55.3 — 70.2 17.6 $53.5
Gemini 3 Flash Preview 62.5
54.7 — 69.7 13.1 $30.1
Mimo V2.5 Pro 60.8
53.2 — 68.6 13.6 $9.2
Deepseek V4 Pro 59.5
51.4 — 67.9 9.0 $3.7
Qwen3.7 Plus 57.7
50.2 — 65.6 18.5 $18.6
Deepseek V4 Flash 54.7
46.9 — 62.8 8.9 $1.5
Qwen3.7 Max 54.2
46.9 — 62.1 14.2 $58.3
Gemini 3.5 Flash-Lite 50.6
42.3 — 58.4 5.4 $34.1
Qwen3.6 27b 45.1
37.6 — 52.8 25.8 $97.3
MiniMax M2.7 41.6
34.7 — 49.4 18.2 $14.9
Gemma 4 31b IT 37.1
29.9 — 44.5 36.3 $10.4
Qwen3.6 35b A3b 37.0
29.1 — 44.3 16.3 $17.8
Latest results as of August 11th.
View archived leaderboards and check back periodically for updates.
Track the latest AI model benchmarks, newly introduced agent architectures, and continuous performance evaluations on the platform. Stay updated with our routine methodology updates and release logs.
  • New models • Aug 11th
    Claude Opus 5
  • New models • Aug 11th
    GPT 5.6 Sol, GPT 5.6 Luna, GPT 5.6 Terra
  • New models • Aug 11th
    Kimi K3
  • New models • Aug 11th
    Qwen3.8 Max
  • New models • Aug 11th
    Gemini 3.6 Flash, Gemini 3.5 Flash Lite
  • Archived models • Aug 11th
    Gemma 4 26B A4B IT
  • New updates • Jul 8th
    Dataset available on Harbor
  • New models • Jul 8th
    Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8
  • New models • Jul 8th
    Qwen 3.7 Max, Qwen 3.7 Plus
  • New models • Jul 8th
    GLM 5.2
  • New models • Jul 8th
    Kimi K2.7 Code
  • New models • Jul 8th
    MiniMax M3
  • Archived models • Jul 8th