A new evaluation framework published today on arXiv reveals that the benchmarks used to judge AI agents that operate smartphones have been measuring only the easiest part of the job — and that today's ...