A new evaluation framework published today on arXiv reveals that the benchmarks used to judge AI agents that operate smartphones have been measuring only the easiest part of the job — and that today's ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results