1
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
如何评估大模型在移动端处理零散个人信息的能力?这项研究给出了全新评测基准与洞见。
arXiv:2608.10692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is …