1
LLM benchmarks are answering someone else's question
揭露LLM基准测试的致命缺陷:用大模型评大模型,只会陷入循环依赖与偏见陷阱。
Article URL: https://danlevy.net/llm-evals-are-broken/ Comments URL: https://news.ycombinator.com/item?id=48572463 Points: 5 # Comments: 0
揭露LLM基准测试的致命缺陷:用大模型评大模型,只会陷入循环依赖与偏见陷阱。
Article URL: https://danlevy.net/llm-evals-are-broken/ Comments URL: https://news.ycombinator.com/item?id=48572463 Points: 5 # Comments: 0