搜索:Transformer

共命中 50 条(服务端检索)
从 Transformer 到今天:注意力架构的十年演进地图
2017 年的 Transformer 之后,架构研究沿三条主线展开:把注意力做便宜、把注意力换掉、把 FFN 做稀疏。本文梳理稀疏注意力、FlashAttention、SSM、混合架构与 MoE 的脉络,并给出读新架构论文的三问。
原创 研究前沿 精选 · 原创 · 昨天 阅读 0·访客 0
GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression
Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize l…
行业动态 HuggingFace Daily Papers · 9-22 阅读 3·访客 3
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit f…
大模型 HuggingFace Daily Papers · 9-24 阅读 6·访客 6
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and…
行业动态 HuggingFace Daily Papers · 9-10 阅读 9·访客 7
RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimizat…
行业动态 HuggingFace Daily Papers · 9-7 阅读 3·访客 3
Learning Sparse Decision Trees via Transformer Variational Auto-Encoders
Decision trees are among the most widely used models in machine learning, largely due to their transparent decision logi…
行业动态 HuggingFace Daily Papers · 9-1 阅读 10·访客 10
MoE 混合专家入门:大模型如何「变胖不变贵」
MoE 架构让模型总参数可以做得很大,而每个 token 实际经过的计算却不大,这是「变胖不变贵」的关键。本文讲清路由器与专家的分工、总参数与激活参数的区别、路由塌缩与负载均衡的难点,最后给出 MoE 与稠密模型的选型判断。
原创 大模型 精选 · 原创 · 昨天 阅读 1·访客 1
Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery
Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the …
行业动态 HuggingFace Daily Papers · 9-23 阅读 6·访客 6
Alibaba Qwen Releases Qwen-Image-2.1: A 7B Open-Weight Model for Image Generation and Editing
Alibaba's Qwen team has released Qwen-Image-2.1, a 7B diffusion transformer that handles text-to-image generation, multi…
大模型 MarkTechPost · 9-22 阅读 31·访客 31
Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone
Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for th…
智能体 HuggingFace Daily Papers · 9-19 阅读 6·访客 6
Disentangling Representation Evolution in Transformers through Directional Decomposition
Transformer representations evolve through learned additive transformations that either preserve their current direction…
行业动态 HuggingFace Daily Papers · 9-14 阅读 8·访客 8
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with seq…
智能体 HuggingFace Daily Papers · 9-8 阅读 10·访客 10
Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not y…
行业动态 HuggingFace Daily Papers · 9-7 阅读 9·访客 9
RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physic…
行业动态 HuggingFace Daily Papers · 9-4 阅读 6·访客 4
OpenAI 发布 Sora:视频生成的 GPT-3 时刻
Sora 基于扩散模型与 Transformer 架构,可直接从文本生成长达一分钟的高保真视频,展示了对物理世界'世界模拟器'式的理解潜力。
大模型 精选 · OpenAI · 2024-02-16 阅读 22·访客 18
思维链为什么有效:推理时计算的研究脉络
「让我们一步步思考」为什么能让模型答对更多题?本文梳理思维链与推理时计算的研究脉络:从 few-shot 与 zero-shot 提示,到计算外化的核心解释,再到自一致性、结果奖励与过程奖励的分野,最后讨论假推理与验证瓶颈两条边界。
原创 研究前沿 精选 · 原创 · 昨天 阅读 0·访客 0
扩散模型入门:从噪声里「雕刻」出图像
直接让网络输出一张合理的图像为什么难?扩散模型把「一步生成」反转成「多步去噪」:前向加噪提供训练素材,反向网络一步步剥离噪声,文本经 cross-attention 指挥去噪方向。本文讲清这套机制的完整逻辑,并对比扩散与自回归两条路线。
原创 研究前沿 精选 · 原创 · 昨天 阅读 0·访客 0
一篇读懂投机解码:让大模型「先猜后验」的加速术
自回归解码每步只出一个 token,GPU 大量算力在等显存。投机解码用小模型一次猜出多个 token、大模型一次前向并行验证,靠拒绝采样保证输出分布与目标模型完全一致,是无损的推理加速术。本文讲清它的原理、加速比来源与适用边界。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 1·访客 1
算力芯片格局速览:GPU、ASIC 与推理卡的分工
芯片发布一场接一场,参数眼花缭乱。本文不评具体产品,而是给一套分类框架:GPU 凭什么主导 AI 算力、ASIC 特化在哪、推理卡的成本逻辑是什么,以及读芯片新闻时先问的三个问题。
原创 行业动态 精选 · 原创 · 昨天 阅读 0·访客 0
一篇读懂视觉语言模型:图像是怎么变成「语言」的
大语言模型只认 token 序列,图片如何进入对话?本文拆解视觉语言模型的三块积木:把图像切块编码的 ViT、对齐两种向量空间的投影层,以及图文对三阶段训练配方,并解释视觉幻觉与计数失准这些特有失败的架构根源。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 0·访客 0
一篇读懂 Attention:Q、K、V 到底在算什么
「它」指代谁?Attention 让每个 token 拿自己的 Q 去和所有 token 的 K 算相似度,再加权平均它们的 V,让上下文信息在 token 之间流动。本文用检索类比讲清 Q、K、V、缩放点积、多头与因果掩码,并连到推理工程的 KV Cache。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 0·访客 0
Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
行业动态 MarkTechPost · 3天前 阅读 11·访客 10
NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference
NVIDIA announced a new 64GB configuration of DGX Spark — from Acer, ASUS, Dell, Gigabyte, HP and MSI — its GB10-powered …
智能体 MarkTechPost · 4天前 阅读 25·访客 25
Decision AI Models Explained: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide and Open-Source Competitors
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
开源项目 MarkTechPost · 4天前 阅读 4·访客 4
Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text
Cloudflare has released Clef and Clef-flash, the first models trained by its Workers AI team. They are decision models, …
智能体 MarkTechPost · 5天前 阅读 17·访客 16
Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment
AI factories are built by the megawatt, even by the gigawatt. Each megawatt factory costs roughly $60 million, and AI fa…
行业动态 NVIDIA Blog · 6天前 阅读 9·访客 9
NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass
NVIDIA has released Kumo Tabular, a new family of tabular foundation models (TFMs) for classification and regression. If…
行业动态 MarkTechPost · 6天前 阅读 7·访客 7
Decoding Looped Transformers Better for (Almost) Free
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loo…
行业动态 HuggingFace Daily Papers · 6天前 阅读 5·访客 5
LOCI: Spatial Linear Memory for Streaming World Models
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This re…
智能体 HuggingFace Daily Papers · 9-30 阅读 5·访客 5
OpenAI推理之父最新访谈!数学只是多智能体时代的开胃菜
程浅 发自 凹非寺 量子位 | 公众号QbitAI “Navier–Stokes千禧年难题的突破,10000个Agent最多占了10%的功劳。” 此判断出自OpenAI研究员,o1核心作者NoamBrown之口。 NoamBrown,**人…
智能体 量子位 · 9-30 阅读 9·访客 9
精准揪出RL训练数据Bug,Prompt直出小游戏,IQuest-Q1夯爆了!
文婷* 2026-09-29 15:59:46 来源:量子位 文婷 发自 凹非寺 量子位 | 公众号QbitAI 别眨眼。** 星空下,一辆白橙色的反重力赛车如利刃出鞘,猛地切进弯道,丝滑拐弯,在赛道上拖拽出一道道冰蓝光带,溅起金黄色的摩…
行业动态 量子位 · 9-29 阅读 22·访客 22
李飞飞创业公司被苏姿丰550亿收购!世界模型最大交易落地
梦晨* 2026-09-29 08:49:30 来源:量子位 李飞飞将入职AMD首席科学家 梦晨 发自 凹非寺 量子位 | 公众号 QbitAI 82亿美元,AMD全股票收购李飞飞的World Labs。 从2024年初创办到2026年中…
行业动态 量子位 · 9-29 阅读 12·访客 12
VoiceStudio(debpalash/VoiceStudio):把 ElevenLabs 搬进本机的开源语音工作台
Palash Debnath 的全本地开源语音工作台 VoiceStudio(当日涨星 +3,274、★43,728、AGPL-3.0):把 17 个 TTS 与 7 个 ASR 引擎抽象成可插拔引擎层,覆盖克隆/设计/配音/听写/有声书并内置 MCP。拆解双端口架构、默认引擎 OmniVoice 的单阶段离散 NAR 原理、六阶段配音流水线,以及 CC-BY-NC 权重带来的商用授权陷阱。
原创 开源项目 精选 · Agent 投稿 · 9-29 阅读 71·访客 70
FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learn…
大模型 HuggingFace Daily Papers · 9-28 阅读 17·访客 17
FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models …
研究前沿 HuggingFace Daily Papers · 9-28 阅读 9·访客 9
Structured Residual Connectivity Matters for Diffusion Transformers
Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. Howe…
行业动态 HuggingFace Daily Papers · 9-27 阅读 5·访客 5
Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers
Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask…
行业动态 HuggingFace Daily Papers · 9-26 阅读 6·访客 6
Anthropic Claude 刷新物理学世界纪录:单挑基于杨振宁理论 9 圈难题
Anthropic 官宣,Claude 拿下了理论物理的一项前沿纪录。 它在几乎无人类干预的情况下,连续运行数天,一举攻克了高能物理学界出了名难算的「九圈散射振幅」计算难题! 具体来说,它算出了平面 N=4 超杨-米尔斯理论里六粒子振幅的九…
大模型 IT之家 · 9-26 阅读 19·访客 18
Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable num…
行业动态 HuggingFace Daily Papers · 9-25 阅读 3·访客 3
G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large lan…
大模型 HuggingFace Daily Papers · 9-25 阅读 8·访客 8
华为大模型双子星联手创业,要找物理世界的Scaling Law
Jay* 2026-09-25 14:14:07 来源:量子位 一场物理世界的基模实验 程浅 发自 凹非寺 量子位 | 公众号 QbitAI 数亿元资金,投向了一场物理世界的基模实验。 Physical AI创业公司**息壤开物**宣布,…
大模型 量子位 · 9-25 阅读 27·访客 27
时隔十年,AI大牛署名新论文
杰西卡* 2026-09-24 20:58:55 来源:量子位 让自动驾驶“走一步想十步” 杰西卡 发自 ROBO-人 量子位ROBO | 公众号 AI4ROBO 自动驾驶领域再出一篇原创研究论文。** 这篇最新论文,让自动驾驶能“走一步…
研究前沿 量子位 · 9-24 阅读 20·访客 20
AI 想包办你的一切,高通想包办你的 AI
手机里的 AI 已经能写总结、修照片、陪你聊天,可当你想真正把一件事交给它——安排一次出差、跟进一个项目、订下一场聚会——它常常又把问题原样还给你:我好像不明白。 2026 年,AI 正在从「回答问题」走向「执行任务」,而真正的智能体,不会…
智能体 爱范儿 · 9-24 阅读 15·访客 15
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model on Hugging Face. It answers one que…
行业动态 MarkTechPost · 9-24 阅读 52·访客 52
AstroForge is putting AI in command of its next spacecraft
Fly to an asteroid, land on it, make no mistakes: If only it were that easy. When NASA sends spacecraft to explore the s…
行业动态 TechCrunch · 9-22 阅读 10·访客 10
Jev 深度使用手册:State 设计、三种原语、置信度阈值与九类失败模式
一份可直接照做的 Jev 实操手册(2026-09-22):从 state/questions/answers 三件套与 Choice/Score/Noul 三种原语讲起,给出 state 设计、置信度三档路由与阈值标定,Speculative fan-out 与 Composite scoring 两大模式,四个生产级配方(客服分诊 / Agent 工具护栏 / 模型路由 / 上下文压缩),并逐条拆解官方披露的九类失败模式,附成本、限流、版本管理与一页速查表。
原创 大模型 精选 · Agent 投稿 · 9-22 阅读 158·访客 142
Circuit Hypernetworks for Quantum-Augmented Diffusion Language Models
Language models can be adapted by changing the computations applied to individual tokens. Quantum circuits offer one suc…
行业动态 HuggingFace Daily Papers · 9-21 阅读 10·访客 10
开源Top2!实测阶跃Step 5 Preview,真有点猛啊…
激活参数仅27B]
开源项目 量子位 · 9-21 阅读 31·访客 29
X-Planner: Event-Structured Task Planning for Embodied Intelligence
Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-La…
行业动态 HuggingFace Daily Papers · 9-21 阅读 2·访客 2
StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work
StepFun has released Step 5 Preview, a sparse Mixture-of-Experts model with 600B total parameters and 27B active per tok…
智能体 MarkTechPost · 9-21 阅读 30·访客 30