Run manual QA test cases in a real browser with an AI agent
用AI代理在真实浏览器里自动跑手动测试用例,QA效率神器。
Article URL: https://github.com/broxhq/qpilot Comments URL: https://news.ycombinator.com/item?id=49432514 Points: 1 # Comments: 0
用AI代理在真实浏览器里自动跑手动测试用例,QA效率神器。
Article URL: https://github.com/broxhq/qpilot Comments URL: https://news.ycombinator.com/item?id=49432514 Points: 1 # Comments: 0
用AI智能体在发布前自动检查网站,省心省力的QA神器,开发者必备!
AI agents that QA your website before launch Discussion | Link
医疗视觉问答也能“知道何时不懂”,置信度感知推理让AI诊断更可靠
arXiv:2608.10964v1 Announce Type: new Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produc…
把测试用例放进 Git 会踩哪些坑?QA 社区的尖锐反对意见与真实落地教训,值得一试。
I asked r/QualityAssurance a blunt question: is managing manual test cases as YAML files inside Git a good idea? The thread got lively. People were no…
澳洲航空Project Sunrise计划实现24小时直飞破纪录,2027年超远程航班即将成真。
Record flight shows viability of Qantas Airways’ Project Sunrise routes in 2027.
两大AI巨头为争夺特定基准榜首各自优化模型,GPT-5.6与Opus 5各擅胜场,竞争白热化。
Openai is optimizing for gpqa diamond and anthropic is optimizing for humanity last exam. gpt 5.6 wins on gpqa and opus 5 wins on humanity last exam C…
用自然语言描述就能测试应用的QA智能体,告别繁琐脚本编写
QA agent that tests apps the way you'd explain them Discussion | Link
一个能先读懂代码再自动点击测试的开源QA助手,让自动化测试更智能高效。
Hey guys, we built something interesting that we're using for testing / QA with our own products and it's proving to be quite helpful. MIT License - h…
新方法STEC通过证据压缩提升开放式多跳问答的深度搜索效率与准确性
arXiv:2607.10795v1 Announce Type: new Abstract: In open-domain multi-hop question answering (QA), LLM-based search agents offer a promising approach t…
LLM多智能体系统实现自主量子编程,QAgent能编写OpenQASM代码,探索AI与量子计算的融合。
arXiv:2508.20134v2 Announce Type: replace Abstract: Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optim…
深入探索Promptfoo工具,让QA工程师像测试传统软件一样高效测试LLM,附赠完整手册。
Want the full 46-page handbook? Promptfoo for QA: The Complete Engineer's Handbook (2026 Edition) by Himanshu Agarwal covers everything below in produ…
通过声学信道实现36kbps数据传输,利用OFDM和16-QAM调制,自动消除时钟漂移与相位噪声,适合离线或低功耗场景通讯
Working with Fable 5, a MacBook, and a Pixel phone, I built an acoustic transport library that is several orders of magnitude faster than existing SoT…
大规模研究揭示LLM在多语言MCQA任务中如何从推理过程量化不确定性,跨语言表现的关键发现值得关注。
arXiv:2607.06327v1 Announce Type: cross Abstract: Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing r…
AI能自动写测试,QA该何去何从?看懂这轮职业洗牌,抓住转型红利。
อวสาน QA? — อนาคตของอาชีพทดสอบซอฟต์แวร์ในยุค AI "AI เขียน test เองได้แล้ว — แล้ว QA จะอยู่ไปทำไม?" นี่คือคำถามที่ได้ยินบ่อยขึ้นทุกวันในวงการ tech — แล…
揭示语言模型自我生成QA训练的隐藏脆弱性,一篇值得关注的AI研究论文。
arXiv:2606.32002v1 Announce Type: new Abstract: Language models are increasingly taught from synthetic question--answer (QA) supervision: a model gene…
自改进QA代理处理超5亿tokens,开源实现如何重塑软件测试流程?
Hey, I am working on building agent-qa https://github.com/vostride/agent-qa which is a self-improving QA agent for software teams. Processed over 500M…
机器验证攻克量子优化十年悬案:QAOA在环上近似比精确等于(2p+1)/(2p+2),形式化证明无懈可击。
arXiv:2606.29687v1 Announce Type: cross Abstract: We report a machine-verified resolution of a problem open for over a decade in quantum optimization:…
英语和西班牙语金融因果关系抽取任务,对比三种模型家族,多语言微调方法拔得头筹
arXiv:2606.27446v1 Announce Type: new Abstract: This paper describes team HSA_CORAL's submission to the FinCausal 2026 shared task on extracting cause…
探索视频多模态大模型对人类运动推理能力的极限,最新基准测试揭示关键短板。
arXiv:2606.27999v1 Announce Type: new Abstract: Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understandin…
AI 代码越来越多,创业公司如何在快节奏中守好质量底线?这个讨论直击测试痛点:自动化再完善,真 bug 还是靠客户和人工发现。
With AI-generated code, moving fast, and hitting PMF, what are some ways startups test their changes to deliver high-quality and what are some struggl…