1
Beyond Perplexity: A Behavioral Evaluation Framework for Deployment-Memory Claims in LLM Test-Time Training
测试时训练宣称能记住部署数据,但困惑度真的够吗?本文提出行为评估框架,用交互式测试验证模型记忆声明,比指标更可信。
arXiv:2607.00368v1 Announce Type: new Abstract: Large language model test-time training (TTT) is often evaluated through local proxy metrics: models a…