1
Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
推理token竟分“结构”与“内容”两类,熵引导超令牌可压缩思维链,直击大模型推理成本痛点。
arXiv:2604.26355v4 Announce Type: replace Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level …