搜索:Tri

共命中 10 条(服务端检索)
Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts
Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-mo…
研究前沿 HuggingFace Daily Papers · 9-5 阅读 2·访客 2
FlashAttention 深度解析:一次在数学上什么都没改的注意力加速
FlashAttention 不近似、不改公式,却把注意力训练速度提高数倍、显存从 O(N²) 降到 O(N)——靠的是把「少走显存」当成一等目标:分块装进片上内存、用 online softmax 增量修正、达到 IO 复杂度理论下界。本文拆解它的算法原理与 v1→v2→v3 三代硬件适配,以及它管不了的事:KV cache 与自定义 mask。
大模型 精选 · 原创 · 今天 阅读 1·访客 1
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a tri…
大模型 HuggingFace Daily Papers · 9-29 阅读 5·访客 5
Goldman Sachs expects Big Tech to spend $1.2 trillion on AI infrastructure by 2027, dwarfing Wall Street estimates
Manuel Uth Sep 27, 2026 Goldman Sachs expects Amazon, Alphabet, Microsoft, Oracle, and Meta to spend a combined $1.2 tri…
行业动态 The Decoder · 9-27 阅读 7·访客 7
Diffusion Policy:机器人动作生成为什么弃用回归、改用扩散模型
模仿学习的老问题是「多峰动作分布」——两种都对的做法被回归平均成一种错的。Diffusion Policy 用条件去噪扩散直接建模动作分布,在 12 个任务、4 个基准上平均成功率提升 46.9%,此后 DP3、DPPO 相继跟进,π0 的流匹配与 GR00T 的扩散 Transformer 把它推成了 VLA 时代的标配动作头。
研究前沿 精选 · 原创 · 2天前 阅读 6·访客 6
论文精读:Mamba——线性时间序列建模的选择性状态空间
Mamba 把 SSM 参数改成输入的函数,让固定大小的状态学会按内容取舍;再靠并行扫描与 kernel 融合把状态装进 SRAM,线性复杂度落地——3B 匹敌两倍大的 Transformer,5 倍推理吞吐,线性扩展到百万长度。文末梳理截至 2026-10 它与 Attention 的分工现状。
研究前沿 精选 · 原创 · 2天前 阅读 6·访客 6
一篇读懂 FlashAttention:注意力为什么能又快又省显存
标准注意力的瓶颈不在算力而在显存读写:O(N²) 的注意力矩阵要在 HBM 里反复进出。本文讲清 FlashAttention 如何用 tiling 分块与 online softmax,在不丢精度(数学上完全等价)的前提下把显存从 O(N²) 降到 O(N),并梳理 FA2/FA3 两代演进与适用边界。
一叶一世界 精选 · 原创 · 2天前 阅读 4·访客 4
Inside NVIDIA’s IsaacTeleop: From Hand and Controller Tracking to Robot Actions with the Graph-Based Retargeting Engine
!(https://www.gstatic.com/images/branding/googleg/1x/googleg_standard_color_128dp.png)Add as a preferredsource on Google…
智能体 MarkTechPost · 5天前 阅读 28·访客 27
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may…
行业动态 HuggingFace Daily Papers · 9-10 阅读 11·访客 10
UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
All-in-One medical image restoration (MedIR) aims to address diverse tasks across modalities and degradation types using…
行业动态 HuggingFace Daily Papers · 9-10 阅读 14·访客 12