Android Bench
AI-assisted software engineering has seen the emergence of several benchmarks to measure the capabilities of LLMs. Android developers face specific challenges that aren't covered by existing benchmarks, so we created one that focuses on a north star of high quality Android development.
| Model | Score (%) Average percentage of 100 test cases successfully resolved across 5 runs for each model |
arrow_range
Cl range (%)
Expected performance range, reflecting the results' statistical reliability (p-value < 0.05)
|
Avg latency (h)
Average time taken to solve 100 tasks across 5 runs
|
Avg cost ($)
Average cost per full benchmark run
|
|---|---|---|---|---|
|
|
91.8 | 87.4 — 95.7 | 13.2 | $175.4 |
|
|
90.9 | 86.2 — 95.4 | 8.7 | $144.3 |
|
|
90.8 | 85.6 — 95.6 | 14.2 | $150.0 |
|
|
90.2 | 85.0 — 94.6 | 29.6 | $112.9 |
|
|
87.6 | 81.8 — 92.3 | 10.6 | $7.2 |
|
|
87.0 | 82.2 — 91.6 | 41.2 | $196.7 |
|
|
86.8 | 81.7 — 91.2 | 7.6 | $48.0 |
|
|
80.2 | 73.0 — 86.6 | 11.4 | $138.3 |
|
|
76.2 | 68.8 — 82.5 | 12.3 | $99.9 |
|
|
75.6 | 67.5 — 82.1 | 10.2 | $135.5 |
|
|
74.3 | 66.8 — 81.5 | 10.5 | $86.6 |
|
|
74.1 | 66.6 — 81.2 | 8.4 | $83.4 |
|
|
72.4 | 65.1 — 79.2 | 6.7 | $88.0 |
|
|
72.2 | 65.3 — 79.1 | 38.9 | $117.0 |
|
|
71.1 | 63.8 — 78.5 | 28.3 | $165.6 |
|
|
70.4 | 63.4 — 77.1 | 31.8 | $48.1 |
|
|
68.7 | 60.9 — 76.4 | 7.0 | $96.5 |
|
|
67.6 | 59.8 — 74.8 | 57.2 | $49.4 |
|
|
67.0 | 58.5 — 75.1 | 16.9 | $127.6 |
|
|
63.6 | 56.1 — 70.7 | 26.0 | $41.7 |
|
|
63.2 | 55.3 — 70.2 | 17.6 | $53.5 |
|
|
62.5 | 54.7 — 69.7 | 13.1 | $30.1 |
|
|
60.8 | 53.2 — 68.6 | 13.6 | $9.2 |
|
|
59.5 | 51.4 — 67.9 | 9.0 | $3.7 |
|
|
57.7 | 50.2 — 65.6 | 18.5 | $18.6 |
|
|
54.7 | 46.9 — 62.8 | 8.9 | $1.5 |
|
|
54.2 | 46.9 — 62.1 | 14.2 | $58.3 |
|
|
50.6 | 42.3 — 58.4 | 5.4 | $34.1 |
|
|
45.1 | 37.6 — 52.8 | 25.8 | $97.3 |
|
|
41.6 | 34.7 — 49.4 | 18.2 | $14.9 |
|
|
37.1 | 29.9 — 44.5 | 36.3 | $10.4 |
|
|
37.0 | 29.1 — 44.3 | 16.3 | $17.8 |
Latest results as of
August 11th.
View archived leaderboards and check back periodically for updates.
View archived leaderboards and check back periodically for updates.
Latest Updates
Track the latest AI model benchmarks, newly introduced agent architectures, and continuous performance evaluations on the platform. Stay updated with our routine methodology updates and release logs.
-
New models • Aug 11th
Claude Opus 5
-
New models • Aug 11th
GPT 5.6 Sol, GPT 5.6 Luna, GPT 5.6 Terra
-
New models • Aug 11th
Kimi K3
-
New models • Aug 11th
Qwen3.8 Max
-
New models • Aug 11th
Gemini 3.6 Flash, Gemini 3.5 Flash Lite
-
Archived models • Aug 11th
Gemma 4 26B A4B IT
-
New updates • Jul 13th
Share results on the community repository
-
New updates • Jul 13th
Contribute new tasks to our community dataset on GitHub
-
New updates • Jul 8th
Dataset available on Harbor
-
New models • Jul 8th
Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8
-
New models • Jul 8th
Qwen 3.7 Max, Qwen 3.7 Plus
-
New models • Jul 8th
GLM 5.2
-
New models • Jul 8th
Kimi K2.7 Code
-
New models • Jul 8th
MiniMax M3
-
Archived models • Jul 8th
Claude Opus 4.6, Claude Sonnet 4.5
-
Archived models • Jul 8th
GPT OSS 120B, GPT OSS 20B
-
Archived models • Jul 8th
Qwen 3.5 9B, Qwen 3.6 Max Preview
-
New updates • Jul 8th
We have migrated our benchmark framework to Harbor
-
New updates • Jul 8th
We've updated Android Bench
Learn more about Android Bench
Our methodology
Learn more about how we created a set of common Android developer tasks.
Android best practices
Many of the tasks are based on how we define high quality Android development, which is detailed in our developer documentation.
Harbor dataset
See the full dataset on Harbor dashboard.