My Reading Library: Evaluating LLMs on Android Tasks

Wait 5 sec.

Can LLM agents actually get through a day in the life of a normal user? That question got me reading papers on Android agents and mobile benchmarks over the past few months. A few patterns kept showing up: Most benchmarks run on emulators, making real-device metrics difficult to measure. Important deployment metrics like battery, thermals, and temperature are often missing. Everyday tasks are scattered across benchmarks, languages, and apps, rather than forming a consistent, globally relevant task set. This makes it harder to evaluate whether an agent can actually work reliably on a real phone, for real users. For now, I’ve put together a library of papers on benchmarking mobile/Android agents for you all to read! Link: https://www.alphaxiv.org/shared/folder/01a070c6-29a0-77a9-a5b4-b670d5eee169   submitted by   /u/East-Muffin-6472 [link]   [comments]