#Agentic Android development
-
Product NewsToday we’re releasing the first set of long-horizon tasks (LHT), which are tasks of great complexity that take an engineer multiple days or even a week to complete. We are also introducing agentic evaluation, starting with agents from corresponding model providers.
Matthew McCullough • 3 min read -
Product NewsBack in March, we introduced Android Bench—our LLM leaderboard for real-world Android development tasks. Since then, we have enhanced the benchmark based on your feedback, including evaluating open-weight models and adding cost and efficiency dimensions to the leaderboard.
Zoe Lopez-Latorre • 3 min read
Stay in the loop
Get the latest Android development insights delivered to your inbox weekly.