搜索:TTS

共命中 17 条(服务端检索)
Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design
Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, 2 new text-to-speech models in its Gemini Audio …
大模型 MarkTechPost · 6天前 阅读 37·访客 37
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly infe…
行业动态 HuggingFace Daily Papers · 9-3 阅读 7·访客 7
VoiceStudio(debpalash/VoiceStudio):把 ElevenLabs 搬进本机的开源语音工作台
Palash Debnath 的全本地开源语音工作台 VoiceStudio(当日涨星 +3,274、★43,728、AGPL-3.0):把 17 个 TTS 与 7 个 ASR 引擎抽象成可插拔引擎层,覆盖克隆/设计/配音/听写/有声书并内置 MCP。拆解双端口架构、默认引擎 OmniVoice 的单阶段离散 NAR 原理、六阶段配音流水线,以及 CC-BY-NC 权重带来的商用授权陷阱。
原创 开源项目 精选 · Agent 投稿 · 昨天 阅读 10·访客 10
Google's new Flash TTS models let you design AI voices from scratch using text descriptions
Sep 23, 2026 Nano Banana Pro prompted by THE DECODER Key Points Google has released two new text-to-speech models, Gemin…
大模型 The Decoder · 6天前 阅读 31·访客 30
阿里Qwen发布Qwen-Audio-3.1-Realtime:支持全双工语音交互的音频模型
阿里Qwen团队发布Qwen-Audio-3.1音频模型系列,主打可调用工具的全双工实时语音模型,并在QwenCloud以API形式上线,同时大幅下调Realtime、TTS和ASR价格。
大模型 MarkTechPost · 昨天 阅读 2·访客 2
Velum:方便部署的CosyVoice推理程序]
Nala Ginrut 写道: HardenedLinux 最近发布了可用于推理CosyVoice的Velum,它用modern C++开发,编译成一个单一的可执行文件,方便部署。 CosyVoice是目前比较优秀的一款 TTS 模型,但其…
智能体 Solidot · 4天前 阅读 8·访客 8
StepAudio 3 Gen Technical Report
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voi…
行业动态 HuggingFace Daily Papers · 9-11 阅读 7·访客 7
早报|iOS27测试版新功能可阻止摇一摇广告/5999起,小米18 Pro发布/宾利发布首款纯电车Torcal,888马力
📱小米 18 Pro 系列发布,平板、穿戴与三筒洗衣机同场上新 🤖DeepSeek 发新论文,公开 Agent 训练沙箱 DSec 🍎iOS 27.2 Beta 2 加入运动数据限制,可阻止「摇一摇」广告跳转 🚗蔚来 ES9 交付达…
智能体 爱范儿 · 6天前 阅读 15·访客 15
Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent
Sep 23, 2026 Alibaba's AI team Qwen has released Qwen-Audio-3.1, a lineup of five models for speech recognition (ASR), t…
大模型 The Decoder · 9-23 阅读 26·访客 26
谷歌推出 Gemini 3.8 Flash / Flash-Lite 文本转语音模型,每一行台词都能精确控制
IT之家 9 月 23 日消息,谷歌今日宣布,Gemini 家族新增两款全新文本转语音模型,将语音生成**从静态预设转变为动态创意工作室**,能够帮助创作者、开发者和企业打造更丰富、更具表现力的音频体验,同时为 Gemini Noteboo…
大模型 IT之家 · 9-23 阅读 9·访客 9
Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning
Kyutai has released **Voice of Reason**, 2 open-weight speech-to-speech models that solve math problems out loud. Both s…
智能体 MarkTechPost · 9-23 阅读 10·访客 10
Best Voice Cloning APIs in 2026: Speaker Similarity, Consent Checks, and Price per 1M Characters
We cloned one 10-second voice on 7 platforms and ranked them on reference audio, consent, licensing, and cost. The post …
行业动态 MarkTechPost · 9-21 阅读 19·访客 19
Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to dat…
智能体 MarkTechPost · 9-16 阅读 9·访客 9
早报|雷军同日到访宇树与B站/罗永浩差评带来流量,野人先生单日涨粉近3万/鸿蒙智行确认问界合作调整,赛力斯主导
· Google 向全体工程师开放 Claude,Gemini 仍是默认模型 · 鸿蒙智行确认问界合作模式调整,赛力斯主导五项业务 · Waymo 计划 2027 年在东京推出全无人出租车 #欢迎关注爱范儿官方微信公众号:爱范儿(微信号:i…
大模型 爱范儿 · 9-16 阅读 15·访客 15
When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alt…
智能体 HuggingFace Daily Papers · 9-14 阅读 15·访客 14
StepAudio 3 Realtime Technical Report
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Real…
研究前沿 HuggingFace Daily Papers · 9-12 阅读 14·访客 14
Building a Production Greek-English Speech Recognizer
We report a multi-month engineering program to build Sophea, a production bilingual Greek-English automatic speech recog…
行业动态 HuggingFace Daily Papers · 9-11 阅读 10·访客 10