Show HN: Muteval – mutation testing for your LLM evals
变异测试一招揪出LLM评测漏洞,Muteval让评估数据更可靠,开源免费直接上手。
Article URL: https://github.com/AshwinUgale/muteval Comments URL: https://news.ycombinator.com/item?id=49388918 Points: 1 # Comments: 0
变异测试一招揪出LLM评测漏洞,Muteval让评估数据更可靠,开源免费直接上手。
Article URL: https://github.com/AshwinUgale/muteval Comments URL: https://news.ycombinator.com/item?id=49388918 Points: 1 # Comments: 0
别只顾着问AI,这项用户视角调查告诉你提示词到底该怎么写更有效
arXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain i…
实测让 AI 编程效果提升 30% 的代码瘦身法,开源工具 CodeSlimmer 帮你喂给大模型更精准的上下文
Article URL: https://github.com/akasula09/CodeSlimmer Comments URL: https://news.ycombinator.com/item?id=49236555 Points: 2 # Comments: 0
单条提示词狂烧6.9亿token做游戏,GPT-5.6却用5美元复刻,AI游戏开发成本战一触即发。
据说GPT-5.6很适合拿来做游戏
用ChatGPT生成CAA沟通板,为无语言自闭症者定制简单 pictogramas,免费替代昂贵商业应用。
Llevo años utilizando herramientas de comunicación aumentativa con mi hijo, que tiene autismo no verbal. Hemos probado de todo: desde apps comerciales…
用提示词让AI画出引用链,一眼看穿五个来源坍缩成一个的循环报道。
Five sources report the same detail, so you write it up as confirmed. Then you look closer. Two of them cite the third. The third cites a post that ha…
别被单一提示词骗了!这项研究揭示LLM安全评估常因表面措辞变化而失真,需警惕安全机制的真实覆盖范围。
arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canoni…
开源工具专治大模型上下文漂移,让AI严格守规矩,代码块神圣不可侵犯。
Article URL: https://github.com/bonushora/surgical-dev-ops/blob/main/README_EN.md Comments URL: https://news.ycombinator.com/item?id=48979553 Points: …
OpenAI官方亲授8个核心技巧,帮你把ChatGPT用到极致,效率翻倍!
别再怪AI了,你的提问方式才是回答空洞的元凶——学会给足上下文,对话质量立竿见影。
A while back I was researching a topic I didn't know much about — the kind of casual, late-night "let me just ask the AI a few questions" session. A f…
可视化对比各家大模型API定价,跑工作流前先算清token成本,省钱又省心。
Article URL: https://neutraloverdrive.com/tools/token-calculator/ Comments URL: https://news.ycombinator.com/item?id=48817123 Points: 1 # Comments: 0
想知道AI何时一本正经地胡说八道?这个HN热帖汇集了各路高手验证大模型错误的“必杀提问”。
Just over-heard one of our seniors trying to explain to a new intern that AI can actually be wrong, no matter how much its answers sound very eloquent…
AI写作竞赛结果出炉,看人类如何用提示词驯服机器,产出不油腻的虚构故事。
Article URL: https://www.hyperstitionai.com/unslop-results Comments URL: https://news.ycombinator.com/item?id=48782890 Points: 21 # Comments: 48
Anthropic发现新模型需要更简洁提示,示例反而限制想象力,提示词走向“短-长-短”新演变。
IT之家 7 月 3 日消息,科技媒体 The Decoder 昨日(7 月 2 日)发布博文,报道称 Anthropic 表示为迎合 Claude Fable 5 模型, 进一步精简 Claude Code 系统提示词,降幅达到 80%。 Anthropic 技术人员塔里克 · 希希帕尔(Tari…
给AI写设计规范,用“该做/不该做”硬规则,让智能体精准还原界面细节
To write DESIGN.md prose agents follow, describe intent and rules, not just values. A token says a color is a hex; good prose says it is for the prima…
来围观 HN 网友如何调教大模型,获取一手实战思路,让你的 LLM 更靠谱
Comments URL: https://news.ycombinator.com/item?id=48610221 Points: 2 # Comments: 0
AI技能总是静默失效?原因在于触发器太死板。教你用高杠杆改动,让技能随叫随到。
There is a quiet failure mode with AI agent skills that almost everyone hits: you build a custom skill, it looks perfect, and it simply never runs. No…
开源实现OpenRouter的混合智能体方案,融合多模型提升回答质量,无需配置即可上手的实用工具。
OpenRouter's fusion is great. I built this for my own deep research queries. Model fusion is useful for tool calls and complex prompts that require co…
用X平台字符限制实例,揭露LLM无法仅靠提示词绕过硬约束,给开发者敲响警钟。
Thursday morning I removed five nodes from my content pipeline. By lunch I understood something about building with language models that eleven failed…
OpenAI官方教程,掌握提示词基础,让ChatGPT输出更精准有用。
Learn prompting fundamentals and how to write clear, effective prompts to get better, more useful responses from ChatGPT.