Google has updated Android Bench, its benchmark for evaluating LLM agents on Android app development. The suite covers 100 development tasks and now includes cost and efficiency metrics, as well as open-weight models. Eight models were added: Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max. Google also adopted a new framework intended to make the benchmark easier to use and invited developers to run tests and submit feedback. Ars Technica reports that Gemini still trails other systems, but the supplied summary provides no detailed ranking or scores.
No heat snapshots are available in the last 24 hours.