搜索:Qwen3-Next

共命中 50 条(服务端检索)
注意力没有被取代,而是被稀释:线性注意力、Mamba 与混合架构的 2026
「线性注意力取代 Transformer」喊了六年,结局出人意料:2026 年一线开源模型的注意力层占比从 100% 压到了 10%-25%,压掉的部分由 Mamba、Gated DeltaNet、KDA 这类常数状态层接管。本文梳理线性注意力、SSM 与 RNN 三条复兴线,拆解 Qwen3-Next 与 Kimi Linear 的混合配方,并回答那个核心问题——为什么还是要留 25% 的注意力。
研究前沿 精选 · 用户投稿 · 今天 阅读 0·访客 0
Qwen3 开源:混合思考模式与多语言覆盖
阿里通义千问发布 Qwen3 系列开源模型,首创'混合思考'模式(可按需开关推理),MoE 与 Dense 全尺寸覆盖,衍生模型数登顶全球开源生态。
开源项目 精选 · Qwen · 2025-04-30 阅读 21·访客 18
Alibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-Time Interpretation Model That Cuts Average Lag to 2.3 Seconds Across 60 Languages
Qwen has released Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model built on a new Interleave archite…
大模型 MarkTechPost · 9-20 阅读 23·访客 23
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
Alibaba's Qwen3.8-Omni-Flash understands audio and video, plans tasks, calls tools, and reports about 45.7% fewer tokens…
智能体 MarkTechPost · 9-18 阅读 36·访客 36
阿里 Qoder 平台 Qwen3.8-Flash 限时免费活动延期,10 月之后继续用
IT之家 9 月 29 日消息,根据 Qoder 官方公告,阿里 Qwen3.8-Flash 限时免费活动已延期 —— 原定 2026 年 9 月 30 日 23:59:59 结束,现改为 10 月 1 日起继续免费(结束时间未定)。 该活…
大模型 IT之家 · 9-29 阅读 15·访客 15
从代码分析到授权争议:一用户用 AI 破解 IDM 引发争议
IT之家 9 月 26 日消息,据 Wccftech 报道,一则关于“大语言模型破解付费软件”的帖子今日在 Reddit 上引发关注。 用户 No_Ideal8394 声称,他利用无审查限制的 Qwen3.8-Flash-Next-Unce…
大模型 IT之家 · 9-26 阅读 24·访客 24
BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost
BottleCap AI has released ThinkingCap-Qwen3.8-27B, the second model in its ThinkingCap series. It is a fine-tune of the …
智能体 MarkTechPost · 9-25 阅读 39·访客 38
Sakana AI hires Jürgen Schmidhuber, inventor of deep learning, world models, and your next ChatGPT update
Sep 24, 2026 Silicon Valley's next big AI idea is probably already sitting in a 1991 paper by Sakana AI's new chief advi…
研究前沿 The Decoder · 9-25 阅读 26·访客 26
Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks
Qwen3.8-Omni-Flash is Qwen's first multimodal model designed for AI agents. It processes audio and video together and in…
智能体 The Decoder · 9-19 阅读 26·访客 26
PrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance
PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B. The language model occupies 5.93 GB,…
开源项目 MarkTechPost · 9-19 阅读 34·访客 34
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-…
行业动态 HuggingFace Daily Papers · 9-9 阅读 9·访客 7
Open or closed AI? Nvidia’s Nader Khalil and Sydney Sykes take on one of the decisions shaping next-gen startups at TechCrunch Disrupt 2026
Nvidia's Nader Khalil and Sydney Sykes discuss one of the decisions shaping next-gen startups on the Builders Stage at T…
行业动态 TechCrunch · 9-18 阅读 16·访客 16
MoE 的工业进化:从 1.6 万亿的 Switch Transformer 到 2.8 万亿的 Kimi K3
2025-2026 年发布的前沿模型几乎清一色是 MoE:Llama 4、gpt-oss、Qwen3、Kimi K2,直到 2026 年 2.78T 参数的 Kimi K3。本文梳理 MoE 从 1991 年分治网络、2017 年稀疏门控,到细粒度专家与无辅助损失均衡的完整演进线,算清「省算力不省显存」的工程账,也正视微调难、训练不稳与成本争议。
大模型 精选 · 用户投稿 · 今天 阅读 4·访客 4
NVIDIA 推出 PivotOPD:教多轮智能体从关键错误中恢复
NVIDIA 联合普林斯顿大学和马里兰大学提出面向多轮 LLM 智能体的在线策略蒸馏方法 PivotOPD,训练智能体避免最致命的早期错误并在发生时恢复,在 ALFWorld、WebShop 和搜索问答任务上对 Qwen3-1.7B 与 Qwen3-8B 取得 13 个基线中的最佳平均成绩。
研究前沿 MarkTechPost · 昨天 阅读 0·访客 0
TechCrunch Disrupt 2026: Blackstone’s Jas Khaira on building the next generation of AI giants
AI startups can grow at a speed that would have been difficult to imagine a generation ago. But rapid growth comes with …
行业动态 TechCrunch · 6天前 阅读 25·访客 25
How Open Science Can Help Researchers Prepare for the Next Pandemic
When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they underst…
行业动态 NVIDIA Blog · 9-24 阅读 19·访客 19
The US Navy just told us what’s on its tech wish list for the next several years
Navy CTO Justin Fanelli talks co-investing alongside VCs instead of funding early research himself, recent buys like a $…
行业动态 TechCrunch · 9-20 阅读 31·访客 31
Your startup’s next teammate might be an AI agent: Gusto, Insight Partners, and Leland explain what that changes at TechCrunch Disrupt 2026
This session will explore how early-stage companies are building teams where humans and AI agents work alongside each ot…
智能体 TechCrunch · 9-17 阅读 27·访客 27
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies lo…
智能体 HuggingFace Daily Papers · 9-17 阅读 25·访客 25
Agent Lightning v1.0:微软用 3500 行代码,把任意 Agent 接进强化学习
给已有 Agent 框架做 RL 后训练,通常要把业务逻辑重写成训练代码。微软的 Agent Lightning 换了条路:Agent 照常跑自己的循环,训练侧伪装成一个 OpenAI 风格的 API 端点,靠拦截请求-响应对收集轨迹。v1.0 技术报告里,Qwen3.5-9B 编码 Agent 在 SWE-bench Verified 上从 41.8% 提到 56.4%。本文拆解它的架构、算法与社区争议。
智能体 精选 · 原创 · 昨天 阅读 3·访客 3
Tony Fadell on why the first wave of AI gadgets failed — and what comes next
When Tony Fadell takes the stage at the inaugural MIT Future Fest, he projects a slide with three images of once-hyped A…
行业动态 TechCrunch · 2天前 阅读 3·访客 3
开源模型全景 2026:六大家版本与许可证对照,什么场景该选谁
2026 年开放权重阵营的旗手已换成 GLM-5(744B,MIT)、Qwen3.5-397B-A17B(Apache 2.0)、DeepSeek V3.2(MIT)、Kimi K2 系与 Llama 4。本文按许可证、架构与部署成本三条线做全景对照,给出自托管、微调、企业集成三类场景的选型建议。
大模型 精选 · 原创 · 2天前 阅读 33·访客 33
LeetCode 206. 反转链表:迭代、递归与「纸牌串」直觉
迭代版 prev/curr 双指针的三步摘插配「纸牌串」类比,讲透先存 next 防断链的指针纪律;递归版逐行拆解 head.next.next = head 的回头指与置空防环时机;附两版代码与边界用例本地实测、O(n)/O(1) 对 O(n)/O(n) 的复杂度对比与工程取舍。
算法题解 精选 · 原创 · 2天前 阅读 8·访客 8
Sensor-Language-Action Models
Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models ho…
行业动态 HuggingFace Daily Papers · 3天前 阅读 0·访客 0
Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
Blaming a “significant rise” in AI submissions, Google has paused its open source bug bounty program until next year. La…
开源项目 TechCrunch · 4天前 阅读 29·访客 29
苹果 MacBook Pro 外接 iPhone 17 Pro Max 运行 AI 模型,预填充性能最高提升 44%
有用户通过 USB-C 将 iPhone 17 Pro Max 外接至 24GB 内存的 M4 Pro MacBook Pro,借助自制软件将 Qwen3.8-27B 模型的运算任务在两台设备间拆分协同处理,使预填充性能最高提升 44%。
行业动态 IT之家 · 5天前 阅读 24·访客 24
亚马逊开源 Strands Decider 2B 决策模型,支持本地 CPU/GPU 部署
亚马逊 Strands Agents 团队开源了基于 Qwen3.5-2B、采用评分指针头与 rank-16 LoRA 微调的决策模型 Strands Decider 2B,权重在 Hugging Face 上,可在本地硬件运行,2B 级模型中排名第 3,本地决策时延中位数 113ms。
大模型 IT之家 · 6天前 阅读 32·访客 32
OpenAI says planned GPT-6.1 is too insecure to release
OpenAI says it has canceled plans to release its updated GPT-6.1 model next month as it continues to investigate what te…
大模型 Ars Technica · 9-29 阅读 15·访客 15
The Pentagon wants $30 million to build an AI-powered lie detector
The US government wants to spend $30.3 million over the next five years on an improved form of lie detector, according t…
行业动态 MIT Technology Review · 9-25 阅读 12·访客 12
Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI
Scalable simulation is essential for robot data generation, policy training, evaluation, and safe iteration, yet real-wo…
智能体 HuggingFace Daily Papers · 9-23 阅读 14·访客 14
AstroForge is putting AI in command of its next spacecraft
Fly to an asteroid, land on it, make no mistakes: If only it were that easy. When NASA sends spacecraft to explore the s…
行业动态 TechCrunch · 9-22 阅读 12·访客 12
OpenAI reportedly closes in on solving the Hodge conjecture, its second Millennium Prize Problem
OpenAI is reportedly tackling the next Millennium Prize Problem. After its still unconfirmed solution to the Navier-Stok…
行业动态 The Decoder · 9-18 阅读 32·访客 30
网易有道周枫:AI能力竞争,正在进入「Model + Agent + Workflow」时代,网易有道AI Open Day展示AI时代“有道解法”
9月16日,网易有道「NEXT,AGENT|有道AI Open Day」在北京举办。]
智能体 量子位 · 9-17 阅读 36·访客 35
d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
AI inference chipmaker d-Matrix today announced it will use NVLink Fusion to connect its next-generation Raptor XPUs to …
行业动态 NVIDIA Blog · 9-10 阅读 12·访客 12
Unlocking Lossless Speedups in LLMs via Discrete Diffusion
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) str…
智能体 HuggingFace Daily Papers · 9-3 阅读 8·访客 7
SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in o…
行业动态 HuggingFace Daily Papers · 9-2 阅读 14·访客 13
GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026
NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices…
行业动态 NVIDIA Blog · 8-27 阅读 10·访客 10
JetBrains 发布 Mellum2.1:面向编码智能体的 12B MoE 开源模型
JetBrains 推出面向编码智能体的开源模型 Mellum2.1,总参数 12B、每 token 激活 2.5B,主要通过在真实软件环境中的强化学习带来提升,以 Apache 2.0 协议在 Hugging Face 发布,可自托管运行。
大模型 MarkTechPost · 今天 阅读 2·访客 2
从 4K 到百万 token:RoPE 上下文扩展算法深度解析
现代大模型的百万上下文,几乎都建立在 2021 年的旋转位置编码与 2023 年的一系列外推算法之上。本文拆解位置插值、NTK-aware 缩放、YaRN 分段插值、LongRoPE 进化搜索这条演进线,核对 Llama、Qwen、DeepSeek 三类工程实践的真实做法,并对照 RULER/NoLiMa 的实测数字:宣称窗口与有效窗口的差距有多大。
大模型 精选 · 原创 · 今天 阅读 0·访客 0
MoE 深度解析:路由、负载均衡与 671B 只激活 37B 的代价
混合专家(MoE)让模型参数量与计算量解耦,但代价是引入了一个训练时最容易翻车的部件——路由器。本文沿「辅助损失 → z-loss → 免辅助损失的偏置调节」这条负载均衡演进线,拆解从 Switch Transformer 到 DeepSeek-V3 的路由设计细节,并核对 2025 年开源 MoE 收敛到 DeepSeek 范式的过程与遗留争议。
大模型 精选 · 原创 · 今天 阅读 0·访客 0
长上下文的注意力之战:滑窗、外推与原生稀疏的三线会战
2026 年,1M token 上下文已成为头部旗舰的标配,但 RULER 与 NoLiMa 评测反复证明「宣称上下文」远大于「有效上下文」。本文梳理长上下文的三条技术路线——滑动窗口、RoPE 长度外推、稀疏化与高效内核,并聚焦 2025 年的分水岭:NSA、MoBA 与 DeepSeek DSA 证明训练时原生的稀疏注意力可以不掉点,最终产品化为 API 降价一半。
研究前沿 精选 · 用户投稿 · 今天 阅读 0·访客 0
拆开 2026 年的旗舰模型:Transformer 里还剩多少 2017 年的零件
九年过去,Transformer 的骨架纹丝未动,外围部件却几乎换了一遍:归一化从 Post-LN 走到 QK-Norm,位置编码从正弦函数走到 RoPE 与 NoPE 混布,注意力从 MHA 走到 GQA 与 MLA。本文以 12 家旗舰模型的官方配置为据梳理这条演进主线,并回答一个问题——模型之间的差距,还剩多少在架构里?
大模型 精选 · 用户投稿 · 今天 阅读 0·访客 0
Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
At a glance Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness…
智能体 Microsoft Research · 昨天 阅读 6·访客 6
Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 昨天 阅读 0·访客 0
一篇读懂 MoE:大模型「大而不贵」的经济学
DeepSeek-V3 有 671B 参数,每个 token 却只动用 37B 的计算量——这不是营销话术,而是混合专家架构(MoE)把「模型容量」和「每 token 算力」拆开的结果。本文从 dense 模型的容量两难讲起,拆解路由器与稀疏激活的工作机制、负载均衡的三代方案、DeepSeekMoE 的细粒度设计,也说清 MoE 的两个代价:省 FLOPs 但不省显存,以及「专家」并不按领域分工的反直觉事实。
一叶一世界 精选 · 原创 · 昨天 阅读 1·访客 1
一篇读懂归一化:LayerNorm、RMSNorm 与 Pre-Norm 的训练稳定性账
LayerNorm 把统计量搬回单样本,RMSNorm 再省掉中心化,换来 7%~64% 的归一化提速;而归一化挂在残差内侧还是外侧,决定梯度随深度指数衰减还是多项式失衡。本文推导公式,并用可运行的 numpy 演示把这笔训练稳定性账算给你看。
一叶一世界 精选 · 原创 · 2天前 阅读 5·访客 5
Decision AI Models Explained: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide and Open-Source Competitors
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
开源项目 MarkTechPost · 6天前 阅读 12·访客 12
NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI
Local AI is becoming more useful by the token. As AI agents move from experiments into everyday development, increasingl…
智能体 NVIDIA Blog · 10-2 阅读 31·访客 31
Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference…
研究前沿 HuggingFace Daily Papers · 9-30 阅读 7·访客 7
笔记本跑7000亿参数GLM!无GPU也行? SSD当显存用火爆GitHub
田, 晏林* 2026-09-26 17:01:00 来源:量子位 GitHub现在最火热的大模型开源小蜂鸟Colibrì是个啥? 闻乐 发自 凹非寺 量子位 | 公众号 QbitAI 25GB笔记本硬跑744B GLM-5.2,32GB…
开源项目 量子位 · 9-26 阅读 35·访客 35