Launch HN: Speko (YC S26) – OpenRouter for Voice AI
语音AI界的OpenRouter,按语言与场景智能路由多模型,不再迷信厂商榜。
Hi HN! I'm Bek, founder of Speko, a platform that finds an optimal combination of speech-to-text, LLM, and text-to-speech models, given your constrain…
语音AI界的OpenRouter,按语言与场景智能路由多模型,不再迷信厂商榜。
Hi HN! I'm Bek, founder of Speko, a platform that finds an optimal combination of speech-to-text, LLM, and text-to-speech models, given your constrain…
语音AI火爆,Fireworks却缺位?开源模型很强,推理平台优化思路才是真正瓶颈。
I started thinking over why doesn't fireworks support voice models. There are really good opensource models available now, like parakeet, kokoro, Qwen…
OpenAI亲述,六个月打造低延迟实时语音系统的技术路线与架构取舍
Article URL: https://openai.com/index/continuous-voice-interaction-with-gpt-live/ Comments URL: https://news.ycombinator.com/item?id=49162469 Points: …
Smallest.ai获投1300万美元,押注小语音模型+离线大模型双架构,打造几乎以假乱真的实时语音交互。
The startup is building voice models designed to make AI phone calls pass the Turing test.
在LiveKit应用中轻松集成真实电话呼叫,支持语音AI代理和客服工具,开源实时通信平台
We’re excited to announce that Wavix now integrates directly with LiveKit, making it easier than ever to handle both inbound and outbound calls inside…
在Asterisk/FreePBX上自托管语音AI代理,2分钟启动管理界面,生产就绪的6个基线一键部署
Article URL: https://github.com/hkjarral/AVA-AI-Voice-Agent-for-Asterisk Comments URL: https://news.ycombinator.com/item?id=48887276 Points: 1 # Comme…
为低资源语言挑选实时语音AI技术栈的实战指南,直击冷门语种落地痛点
Article URL: https://kamalg2.substack.com/p/choosing-a-real-time-voice-ai-stack Comments URL: https://news.ycombinator.com/item?id=48814001 Points: 1 …
揭秘OpenAI如何用WebRTC为9亿用户打造毫秒级语音AI体验,技术选型干货满满。
In this article, we will look at the entire journey in detail and challenges the OpenAI engineering team faced.
思必驰二度冲刺科创板,却被多家经销商实名举报画大饼、诱导囤货,上市之路再添变数。
IT之家 6 月 26 日消息,据第一财经 6 月 25 日报道, 三份来自芯片经销商的实名举报材料 ,直指二度冲击科创板 IPO 的语音 AI 企业思必驰科技股份有限公司(下称“思必驰”)及其控股子公司深聪半导体(江苏)有限公司(下称“深聪”)。 报道称,这三家昔日的“盟友”均反映,在与思必驰在蓝…
评估四大主流实时语音AI在“言外之意”上的表现,揭示模型“听得到词却听不懂调”的盲点
arXiv:2606.26083v1 Announce Type: new Abstract: Speech conveys information through both words and vocal delivery. We evaluate four leading production …
轻量开源SDK让语音AI实时处理群组通话,流式音频上传,Apache-2.0授权即拿即用
Hey folks. We built SAA (Selective Auditory Attention) after trying to find ways to make a good experience with multiple robots/multiple agents. What …
实时语音AI的传输选型关键:为何WebRTC在延迟与拥塞控制上完胜WebSockets,做语音Agent必读。
Article URL: https://livekit.com/blog/why-webrtc-beats-websockets-for-voice-ai-agents Comments URL: https://news.ycombinator.com/item?id=48616386 Poin…
开源语音AI平台Omni-VRAM集成了28个Python模块,支持语音识别、实时流、情绪识别等,<200ms延迟,免费可用!
What I Built Omni-VRAM is an open-source voice AI platform with 28 modules. GitHub: https://github.com/Liangchenxu/Omni-VRAM Features Speech Recogniti…
探索语音感知大语言模型在说话人验证上的潜力,提出与模型无关的评分方法,揭示身份编码机制。
arXiv:2603.10827v2 Announce Type: replace-cross Abstract: Speech-aware large language models (LLMs) can accept speech inputs, yet their training objec…
多任务临床语音AI基准,助力医疗语音技术标准化评估。
arXiv:2606.17339v1 Announce Type: new Abstract: Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor…
一款能预测语音AI性能的音频洞察工具,帮你优化声音交互体验。
Audio insight that predicts voice AI performance Discussion | Link
Cartesia AI推出SOTA语音合成与识别模型,唯一在语音和转录领域均排名第一,无需在质量与速度间妥协。
Article URL: https://www.cartesia.ai/launch/ Comments URL: https://news.ycombinator.com/item?id=48546086 Points: 2 # Comments: 0
实时语音AI的延迟生死线:从Go到Rust的迁移如何压缩毫秒级时延,守住250ms实时交互边界。
In building Vivik, an execution-grade telephony AI engine, we faced a brutal constraint: the human conversational loop. In psychoacoustics, a delay un…
用100+语言构建语音AI代理,每分钟仅1.5卢比,极致性价比。
Build Voice AI Agents in 100+ languages for ₹1.5/min Discussion | Link