Android Bench 2.0 focuses on long-horizon tasks, agent evaluations
AI Generated Image

Android Bench 2.0 focuses on long-horizon tasks, agent evaluations

9to5Google technology

Key Points:

  • Google has released Android Bench 2.0, an AI benchmark designed to evaluate models on complex development tasks such as adding new features, building apps from scratch, and converting cross-platform apps to Android.
  • Unlike the original version focused on incremental changes, Android Bench 2.0 uses a continuous scoring system that assesses functionality, visual fidelity, and adherence to evaluation instructions rather than a simple pass/fail metric.
  • Top-performing models include GPT-6 Astra with a 28% pass rate on complex tasks, though overall completion rates remain low, highlighting the ongoing challenge of porting cross-platform apps and handling architectural complexity.
  • AI models perform better at writing new code than refactoring or migrations, excelling in deterministic transformations like Java to Kotlin conversion but struggling with runtime validation, breaking framework changes, and unreleased libraries.
  • Android Bench also evaluates AI agents from model providers, noting that agent design influences developer outcomes, and plans to expand testing with various model-agent combinations in future updates.

Trending Business

Trending Technology

Trending Health