搜索:FLOPS

共命中 16 条(服务端检索)
算力的度量衡:从 FLOPS 到卡时,看懂算力新闻的单位
FLOPS、算力、卡时、集群规模混着说,是看懂算力新闻的第一道坎。本文厘清 FLOPS 与 FLOPs 一字之差的两个概念,给出 6ND 训练算力估算经验与卡时换算方法,解释峰值与有效算力之间的 MFU 差距,最后附一张看懂算力新闻的换算清单。
原创 行业动态 精选 · 原创 · 2天前 阅读 6·访客 5
欧盟 AI Act 两年记:义务时间线、GPAI 合规要点与 2025 年底那次急转弯
欧盟 AI Act 2024 年 8 月生效,2025 年 2 月禁令条款先行,8 月起 GPAI 模型义务落地;原定 2026 年 8 月的高风险义务被 11 月的 Digital Omnibus 推迟。本文梳理截至 2026 年 10 月的完整时间线、罚则结构与开发者视角的合规要点。本文不构成法律意见。
原创 行业动态 精选 · 原创 · 昨天 阅读 5·访客 4
一篇读懂 FlashAttention:注意力为什么能又快又省显存
标准注意力的瓶颈不在算力而在显存读写:O(N²) 的注意力矩阵要在 HBM 里反复进出。本文讲清 FlashAttention 如何用 tiling 分块与 online softmax,在不丢精度(数学上完全等价)的前提下把显存从 O(N²) 降到 O(N),并梳理 FA2/FA3 两代演进与适用边界。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 3·访客 3
一篇读懂归一化:LayerNorm、RMSNorm 与 Pre-Norm 的训练稳定性账
LayerNorm 把统计量搬回单样本,RMSNorm 再省掉中心化,换来 7%~64% 的归一化提速;而归一化挂在残差内侧还是外侧,决定梯度随深度指数衰减还是多项式失衡。本文推导公式,并用可运行的 numpy 演示把这笔训练稳定性账算给你看。
原创 一叶一世界 精选 · 原创 · 昨天 阅读 3·访客 3
从 Transformer 到今天:注意力架构的十年演进地图
2017 年的 Transformer 之后,架构研究沿三条主线展开:把注意力做便宜、把注意力换掉、把 FFN 做稀疏。本文梳理稀疏注意力、FlashAttention、SSM、混合架构与 MoE 的脉络,并给出读新架构论文的三问。
原创 研究前沿 精选 · 原创 · 2天前 阅读 4·访客 4
Decoding Looped Transformers Better for (Almost) Free
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loo…
行业动态 HuggingFace Daily Papers · 10-1 阅读 8·访客 8
How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual …
大模型 HuggingFace Daily Papers · 9-28 阅读 19·访客 19
A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
In this tutorial, we work with **MSEB**, the Massive Sound Embedding Benchmark from Google Research, and approach it fro…
研究前沿 MarkTechPost · 9-27 阅读 30·访客 29
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method …
行业动态 HuggingFace Daily Papers · 9-27 阅读 6·访客 6
罗福莉官宣小米 MiMo-V3 采用全新架构,核心 HySparse 2 今日发布
IT之家 9 月 23 日消息,小米 MiMo 大模型负责人罗福莉今日发文,宣布 MiMo-V3 即将采用全新架构。其核心 HySparse 2 今日发布,带来更少的预填充、更小的 KV 缓存、更出色的长上下文检索。 在 1M(100 万)…
大模型 IT之家 · 9-23 阅读 16·访客 16
DeltaWAM: Delta World Action Models for Bimanual Manipulation
World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointl…
智能体 HuggingFace Daily Papers · 9-23 阅读 11·访客 11
Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone
Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for th…
智能体 HuggingFace Daily Papers · 9-19 阅读 6·访客 6
Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend
Learn how to leverage NVIDIA’s cuDNN Frontend Graph API to build custom kernel fusions, autotuning engine configurations…
行业动态 MarkTechPost · 9-16 阅读 14·访客 14
Modality-Autoregressive World-Action Models
World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images.…
行业动态 HuggingFace Daily Papers · 9-15 阅读 12·访客 11
Meta新研究:字节模型蒸馏后,天花板破了
]
研究前沿 量子位 · 9-15 阅读 20·访客 20
大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序
把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
原创 大模型 精选 · 本站原创 · 9-10 阅读 84·访客 65