Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching1 independent reports10

Google Updates Android Bench With New LLMs, but Gemini Still Lags Behind

First seen · 7/9/2026, 12:39 AMLatest activity · 7/9/2026, 12:39 AM

Google has updated Android Bench, its benchmark for evaluating LLM agents on Android app development. The suite covers 100 development tasks and now includes cost and efficiency metrics, as well as open-weight models. Eight models were added: Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max. Google also adopted a new framework intended to make the benchmark easier to use and invited developers to run tests and submit feedback. Ars Technica reports that Gemini still trails other systems, but the supplied summary provides no detailed ranking or scores.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. MediaArs Technica AI7/9, 12:39 AMRepresentative
    Google Updates Android Bench With New LLMs, but Gemini Still Lags Behind