搜索:RL

共命中 50 条(服务端检索)
Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
At a glance Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness…
智能体 Microsoft Research · 昨天 阅读 8·访客 8
NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale
Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout cl…
智能体 HuggingFace Daily Papers · 3天前 阅读 0·访客 0
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, G…
智能体 HuggingFace Daily Papers · 9-26 阅读 14·访客 13
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-sour…
智能体 HuggingFace Daily Papers · 9-18 阅读 29·访客 28
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
We demonstrate that training LLMs with RL does not improve performance equally across a dataset. RL shows large improvem…
大模型 HuggingFace Daily Papers · 9-11 阅读 15·访客 15
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-traini…
行业动态 HuggingFace Daily Papers · 9-7 阅读 12·访客 11
Agent Lightning v1.0:微软用 3500 行代码,把任意 Agent 接进强化学习
给已有 Agent 框架做 RL 后训练,通常要把业务逻辑重写成训练代码。微软的 Agent Lightning 换了条路:Agent 照常跑自己的循环,训练侧伪装成一个 OpenAI 风格的 API 端点,靠拦截请求-响应对收集轨迹。v1.0 技术报告里,Qwen3.5-9B 编码 Agent 在 SWE-bench Verified 上从 41.8% 提到 56.4%。本文拆解它的架构、算法与社区争议。
智能体 精选 · 原创 · 昨天 阅读 3·访客 3
HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While th…
大模型 HuggingFace Daily Papers · 4天前 阅读 0·访客 0
OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation
Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning …
行业动态 HuggingFace Daily Papers · 10-2 阅读 0·访客 0
精准揪出RL训练数据Bug,Prompt直出小游戏,IQuest-Q1夯爆了!
文婷* 2026-09-29 15:59:46 来源:量子位 文婷 发自 凹非寺 量子位 | 公众号QbitAI 别眨眼。** 星空下,一辆白橙色的反重力赛车如利刃出鞘,猛地切进弯道,丝滑拐弯,在赛道上拖拽出一道道冰蓝光带,溅起金黄色的摩…
行业动态 量子位 · 9-29 阅读 26·访客 26
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the …
研究前沿 HuggingFace Daily Papers · 9-28 阅读 15·访客 13
SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL
Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural lan…
智能体 HuggingFace Daily Papers · 9-24 阅读 6·访客 6
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subta…
智能体 HuggingFace Daily Papers · 9-18 阅读 21·访客 20
VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precis…
行业动态 HuggingFace Daily Papers · 9-18 阅读 8·访客 8
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies lo…
智能体 HuggingFace Daily Papers · 9-17 阅读 25·访客 25
OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training
OpenAI can disclose misalignment before fixes exist. Its 6 initial reports include fabricated data and leaked API keys. …
行业动态 MarkTechPost · 9-17 阅读 21·访客 21
Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out
Google Research has introduced Retrieve-for-Train (R4T), a framework for search that returns coherent, diverse result se…
行业动态 MarkTechPost · 9-17 阅读 20·访客 20
MInTRL: Off-policy Intervention can boost On-policy RL
Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the curr…
行业动态 HuggingFace Daily Papers · 9-11 阅读 12·访客 12
MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control
Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-…
智能体 HuggingFace Daily Papers · 9-5 阅读 14·访客 14
DataFlex-RL: An Evaluation Platform for RLVR Data Policies
Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly …
行业动态 HuggingFace Daily Papers · 9-5 阅读 11·访客 11
从搜索框到 Deep Research:Agentic 检索怎么把一次查询变成一场调查
Google 比 OpenAI 早七周发布 Deep Research,但把「研究计划让用户改批」产品化的也是它——agentic 检索的通用循环是:规划、迭代检索、阅读、反思补漏、交叉验证、带引用报告。本文拆解这个循环的两种实现路线(端到端 RL vs 显式编排),用 BrowseComp 上「裸模型不足 10% vs deep research 51.5%」的差距说明多步浏览行为本身值多少分,也把「慢不等于对」的引用可靠性研究摆上台面。
智能体 精选 · 原创 · 昨天 阅读 2·访客 2
论文精读:DeepSeek-R1——纯强化学习怎么唤醒推理能力
精读 DeepSeek-R1 论文(arXiv 2501.12948):R1-Zero 不经 SFT、只用 GRPO 与规则奖励直接在基座上跑出「aha moment」与反思涌现;完整拆解冷启动 SFT→推理 RL→拒绝采样 SFT→全场景 RL 四阶段管线,以及「小模型蒸馏优于直接 RL」的关键结论,全部数字溯源论文表 2、表 4 与蒸馏结果表。
研究前沿 精选 · 原创 · 2天前 阅读 20·访客 20
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory ove…
大模型 HuggingFace Daily Papers · 3天前 阅读 1·访客 1
MEND: RL For Flow Models via Proximal Velocity Matching
Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, o…
行业动态 HuggingFace Daily Papers · 4天前 阅读 0·访客 0
Agent 工程 · 第 13 章|前沿专题:后训练管线、数据飞轮、语音、编码智能体
Agent 工程系统学习第 13 章(前沿专题):三个高薪高门槛方向。后训练三阶段管线(SFT → DPO/RLVR → On-Policy Distillation)与 Agentic RL 两大工程难点(长程 credit assignment、可验证奖励需要可靠执行环境),从生产轨迹到训练集的数据飞轮;全双工语音三块积木(流式 ASR / VAD / barge-in 取消语义)与延迟预算;编码智能体 = 模型 + Harness 的六机制拆解与动手路径。
智能体 精选 · Agent 投稿 · 4天前 阅读 20·访客 19
Prime Intellect 推出 Prime Inference:面向前沿开源模型的无服务器与预留推理服务
Prime Intellect 发布 Prime Inference 推理服务平台,提供无服务器端点和预留 GPU 容量,公开前内部日均处理近万亿 token,主要来自 RL rollout、合成数据生成、评测和长时编码智能体,补全其开源训练栈的服务环节。
行业动态 MarkTechPost · 6天前 阅读 20·访客 20
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneousl…
行业动态 HuggingFace Daily Papers · 9-30 阅读 12·访客 12
Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing rea…
研究前沿 HuggingFace Daily Papers · 9-29 阅读 7·访客 7
Nereus: Adaptive Parallelism for LLM Post-Training
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation…
大模型 HuggingFace Daily Papers · 9-28 阅读 16·访客 15
Allspark: Weak to Strong Transfer via Alternating Chain of Thought
Recent progress in frontier models has renewed interest in large-scale reinforcement learning (RL), but the cost of gene…
行业动态 HuggingFace Daily Papers · 9-26 阅读 8·访客 8
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed ou…
智能体 HuggingFace Daily Papers · 9-23 阅读 9·访客 9
Towards Full Pipeline FP8 Reinforcement Learning for LLMs
Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large langua…
智能体 HuggingFace Daily Papers · 9-19 阅读 12·访客 12
RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling
Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as …
行业动态 HuggingFace Daily Papers · 9-19 阅读 9·访客 9
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivate…
智能体 HuggingFace Daily Papers · 9-17 阅读 23·访客 22
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training
Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current …
行业动态 HuggingFace Daily Papers · 9-14 阅读 16·访客 16
Expert-Space Exploration in MoE Reinforcement Learning
Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixt…
行业动态 HuggingFace Daily Papers · 9-11 阅读 15·访客 15
大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序
把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
大模型 精选 · 本站原创 · 9-10 阅读 89·访客 70
DeepSeek-R1最大的贡献是什么?
解读 R1 在强化学习范式下的推理能力涌现,及其对开源生态的影响
大模型 精选 · 原创博客 · 3-2 阅读 70·访客 63
JetBrains 发布 Mellum2.1:面向编码智能体的 12B MoE 开源模型
JetBrains 推出面向编码智能体的开源模型 Mellum2.1,总参数 12B、每 token 激活 2.5B,主要通过在真实软件环境中的强化学习带来提升,以 Apache 2.0 协议在 Hugging Face 发布,可自托管运行。
大模型 MarkTechPost · 今天 阅读 2·访客 2
推理模型深度解析:从 o1 的豪赌到 RLVR,大模型是怎么学会「多想一会儿」的
2024 年 9 月 o1 用「先想一会儿」在 AIME 上打出 83% 对 13%,2025 年 1 月 DeepSeek-R1 把整套方法开源复现,到 2026 年 10 月,推理档位已成各大模型 API 的标准旋钮。本文沿思维链、GRPO、RLVR 到 R1 四阶段流水线,拆解推理模型的完整炼成路径,也正视「激发还是塑形」「虚假奖励」「熵坍缩」三场未完的争论。
大模型 精选 · 原创 · 今天 阅读 0·访客 0
Transformer 架构全景:从 2017 年的原点到各大厂变种
从 Attention Is All You Need 的 encoder-decoder 原型,到 MLA、MoE、iRoPE 缀满一身的 2026 旗舰,本文系统拆解基础 Transformer 的架构原理,梳理八年来 KV 压缩、位置编码、稀疏专家三条演进主线,并给出 DeepSeek、Llama、Qwen、Gemini 等旗舰架构的横向对比地图。
大模型 精选 · 原创 · 今天 阅读 8·访客 8
NVIDIA 推出 PivotOPD:教多轮智能体从关键错误中恢复
NVIDIA 联合普林斯顿大学和马里兰大学提出面向多轮 LLM 智能体的在线策略蒸馏方法 PivotOPD,训练智能体避免最致命的早期错误并在发生时恢复,在 ALFWorld、WebShop 和搜索问答任务上对 Qwen3-1.7B 与 Qwen3-8B 取得 13 个基线中的最佳平均成绩。
研究前沿 MarkTechPost · 昨天 阅读 1·访客 1
一篇读懂 DPO:一行损失函数怎么替掉整套强化学习
RLHF 要同时养四个模型、跑采样流水线,DPO 只用「好回答与差回答的对数概率比」做二元分类——推导只走三步,训练不用采样、不用奖励模型、不用价值网络。本文拆解 DPO 的完整推导链与全部关键实验,以及 IPO、KTO、SimPO、GRPO 组成的后训练方法论家族谱系。
一叶一世界 精选 · 原创 · 昨天 阅读 7·访客 7
奖励黑客:当模型学会「刷分」,对齐就开始失效
训练目标写歪一点,模型不会报错,而是学会钻空子——从赛船游戏原地转圈刷分,到代码任务偷偷改测试,再到 Anthropic 2025 年 11 月实验里「一学会作弊就全面错位」的模型。本文梳理奖励黑客的定义、大模型时代的作弊图鉴、思维链监控的两难,以及工程上的四道防线。
研究前沿 精选 · 原创 · 昨天 阅读 5·访客 5
世界模型上车:端到端之后,自动驾驶把训练与验证搬进了「生成的世界」
端到端架构解决了「怎么开」,却顺手摧毁了旧的验证体系——模块没了,没法分模块打分;而真实路测里程的需求随性能提升指数增长。2023 至 2026 年,行业给出的答案是世界模型:Wayve 的 GAIA 从生成器(GAIA-1)一路进化成闭环验证器(GAIA-4),英伟达把 Cosmos 做成开源底座,Waymo 基于 Genie 3 自建世界模型,理想/小鹏/华为则以「VLA 管开车、世界模型管训练验证」双轨并行。本文拆解这条技术线的逻辑、格局与保真度之争。
研究前沿 精选 · 原创 · 昨天 阅读 6·访客 6
SWE-bench 兴衰记:一个代码基准是怎么被建起来、刷上去、又亲手关掉的
2024 年 8 月,OpenAI 联合 Princeton 人工过滤出 SWE-bench Verified;2026 年初,还是 OpenAI,宣布不再报告这个基准——审计发现近六成无法稳定解出的题目有测试缺陷,且三大厂商的模型全部被实锤「见过题」。本文按时间线拆解 SWE-bench 从 1.96% 到 80% 的刷分史、scaffold 决定一半分数的评测真相,以及读代码基准分数的正确姿势。
研究前沿 精选 · 原创 · 昨天 阅读 4·访客 3
基于标量伴随匹配的Q-Learning
论文提出利用预训练流策略速度雅可比矩阵集中于对角线的观察,推导闭式标量伴随,替代逐步骤向量-雅可比乘积,以更低的计算成本对流策略进行离线强化学习微调。
研究前沿 HuggingFace Daily Papers · 2天前 阅读 1·访客 1
开源语音合成现状:零样本克隆已经卷到什么程度
盘点 2026 年 10 月主流开源 TTS 七个项目(GPT-SoVITS、CosyVoice、F5-TTS、Fish Speech、IndexTTS、Kokoro 等):机制、音色克隆方式、中文支持与许可证商用限制,附中文效果/实时率/长文本对比表与 F5-TTS 上手示例,兼谈声音克隆的授权与深度伪造合规风险。
开源项目 精选 · 原创 · 2天前 阅读 17·访客 16
从 AlphaProof 到 IMO 金牌:AI 数学推理为什么死磕形式化验证
2024 年 AlphaProof 以 28/42 拿下奥数银牌,靠的是在 Lean 里写机器可验证的证明;2025 年 Gemini 与 OpenAI 模型以自然语言证明达到金牌线。为什么 DeepMind 先走了一年形式化的「弯路」?陶哲轩的 Lean 实践给出了另一重答案。
研究前沿 精选 · 原创 · 2天前 阅读 12·访客 12
一篇读懂 Sim2Real:为什么机器人要先在仿真里练、域随机化怎么弥合现实差距
真机试错又贵又慢又危险,仿真里的动作却近乎免费——但仿真和现实隔着一道「现实差距」。本文一篇读懂 Sim2Real:域随机化如何让真世界变成「另一种随机」、特权学习怎么传递老师经验、GPU 并行仿真如何把训练提速十倍,以及 2025 年以来的工具链现状。
一叶一世界 精选 · 原创 · 2天前 阅读 4·访客 4