搜索:Study

共命中 50 条(服务端检索)
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve gen…
行业动态 HuggingFace Daily Papers · 9-28 阅读 7·访客 7
AI access makes people almost entirely unwilling to say "I don't know," study finds
Sep 26, 2026 Nano Banana Pro prompted by THE DECODER Researchers ran five experiments with 3,132 participants to test wh…
行业动态 The Decoder · 9-27 阅读 35·访客 35
Top AI experts badly underestimated how fast the field is moving, study finds
Sep 24, 2026 Nano Banana Pro prompted by THE DECODER How fast is AI improving? That question usually goes to experts at …
行业动态 The Decoder · 9-25 阅读 31·访客 24
An Empirical Study of Harness Design for Coding Agents
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering …
智能体 HuggingFace Daily Papers · 9-17 阅读 22·访客 21
AI 教育的 2026:三场 RCT 都说有效,为什么家长还是不放心
2026 年 AI 教育拿到了迄今最强的证据:哈佛实验里 AI 导师组的学习增益超过面授主动学习的两倍,世界银行在尼日利亚测出约合两年常规进度的提升,Google 在塞拉利昂的 RCT 也有 0.26 个标准差。但三场研究全部由利益相关方资助、周期不超过一学期,而 MIT 的脑电研究与「护栏可绕过」的产品现实站在对面。本文梳理产品格局、证据攻防与中美政策两条路线。
行业动态 精选 · 原创 · 昨天 阅读 4·访客 4
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain thr…
智能体 HuggingFace Daily Papers · 4天前 阅读 1·访客 1
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by re…
行业动态 HuggingFace Daily Papers · 10-1 阅读 8·访客 8
Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models
In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcom…
智能体 HuggingFace Daily Papers · 10-1 阅读 0·访客 0
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the …
研究前沿 HuggingFace Daily Papers · 9-28 阅读 15·访客 13
AI was supposed to hit new grads hard. So far, unemployment data says otherwise.
Last month, we shared word of a Stanford study that found entry-level employment in so-called "AI-impacted" occupations …
研究前沿 Ars Technica · 9-26 阅读 27·访客 27
Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation
Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessi…
智能体 MarkTechPost · 9-25 阅读 22·访客 21
EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervi…
智能体 HuggingFace Daily Papers · 9-23 阅读 11·访客 11
Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament
Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study exami…
行业动态 HuggingFace Daily Papers · 9-22 阅读 3·访客 3
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in bo…
智能体 HuggingFace Daily Papers · 9-17 阅读 17·访客 16
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even …
行业动态 HuggingFace Daily Papers · 9-17 阅读 25·访客 23
Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation
We study a practical question: can a small correction module fix errors in a frozen language model's outputs without deg…
行业动态 HuggingFace Daily Papers · 9-14 阅读 5·访客 5
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
We study how model post-training and test-time inference design affect natural-language proof generation for hard olympi…
行业动态 HuggingFace Daily Papers · 9-9 阅读 13·访客 11
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体 HuggingFace Daily Papers · 9-8 阅读 15·访客 14
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that…
智能体 HuggingFace Daily Papers · 9-8 阅读 10·访客 9
Online Learning with LLM Experts from Limited Feedback
We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online set…
大模型 HuggingFace Daily Papers · 9-5 阅读 17·访客 17
WorldSculpt: Generating Compositional Worlds from Grounded Videos
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects…
行业动态 HuggingFace Daily Papers · 9-4 阅读 3·访客 3
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system kee…
大模型 HuggingFace Daily Papers · 9-3 阅读 19·访客 16
Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys
We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while l…
大模型 HuggingFace Daily Papers · 9-3 阅读 19·访客 18
MasterControl Seventeen Every Time
We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic poli…
智能体 HuggingFace Daily Papers · 9-2 阅读 8·访客 7
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
行业动态 HuggingFace Daily Papers · 9-1 阅读 10·访客 10
Some mathematicians call for OpenAI boycott after AI-generated proofs flood their field
Manuel Uth & Matthias Bastian Oct 8, 2026 Nano Banana Pro prompted by THE DECODER After OpenAI published hundreds of AI…
智能体 The Decoder · 今天 阅读 3·访客 3
注意力没有被取代,而是被稀释:线性注意力、Mamba 与混合架构的 2026
「线性注意力取代 Transformer」喊了六年,结局出人意料:2026 年一线开源模型的注意力层占比从 100% 压到了 10%-25%,压掉的部分由 Mamba、Gated DeltaNet、KDA 这类常数状态层接管。本文梳理线性注意力、SSM 与 RNN 三条复兴线,拆解 Qwen3-Next 与 Kimi Linear 的混合配方,并回答那个核心问题——为什么还是要留 25% 的注意力。
研究前沿 精选 · 用户投稿 · 今天 阅读 2·访客 2
OpenAI 首份青少年报告:日均使用不足 15 分钟,安全提醒机制遭质疑
北京时间 10 月 8 日,OpenAI 周三发布了首份青少年使用情况报告,称青少年用户平均每天使用 ChatGPT 的时间不足 15 分钟,连续使用超过 3 小时的青少年用户占比不到 2%。与此同时,第三方机构的测试对其家长安全提醒机制提…
大模型 IT之家 · 昨天 阅读 5·访客 5
A single prompt was enough to hijack every AI agent in an AWS account, Zenity researchers found
Oct 8, 2026 Nano Banana Pro prompted by THE DECODER Key Points A chain of vulnerabilities in Amazon's Bedrock AgentCore …
智能体 The Decoder · 昨天 阅读 1·访客 1
ChatGPT for Teens keeps teens talking, even during mental health crises
Common Sense Media, a nonprofit that provides age-based ratings and reviews of media and tech for families, has labeled …
智能体 TechCrunch · 昨天 阅读 2·访客 2
OpenAI dumps 372 AI-generated math proofs on GitHub, telling the academic world to keep up
Oct 7, 2026 Key Points OpenAI has published 372 mathematical results generated by an internal AI model. The results are …
开源项目 The Decoder · 2天前 阅读 5·访客 5
系统切换:快速决策模型何时应该停下来思考?
研究提出在闭环 Doom 环境中,由门控机制决定快速动作模型何时将控制权交给推理型视觉语言模型,并基于 900 道保留问题评估了 0.15B 到 9B 参数零样本决策模型在错误中的收集倾向、准确率与校准表现。
研究前沿 HuggingFace Daily Papers · 2天前 阅读 1·访客 1
Personal-Agent Mediated Recommendation with Cross-Platform User History
Modern recommendation is shifting from platform-centric personalization toward user-governed personalization, where a pe…
智能体 HuggingFace Daily Papers · 3天前 阅读 0·访客 0
From Evidence to Action: How Tool-Using Agents Fail
Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions…
智能体 HuggingFace Daily Papers · 3天前 阅读 1·访客 1
WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification
Individual animal re-identification from camera-trap imagery is an instance retrieval problem central to non-invasive wi…
行业动态 HuggingFace Daily Papers · 4天前 阅读 0·访客 0
From Scan to Treatment Plan, AI Helps Close Breast Cancer’s Deadliest Gaps
Breast cancer is the most commonly diagnosed cancer among American women — yet the gaps in care are wide. A majority of …
行业动态 NVIDIA Blog · 4天前 阅读 2·访客 2
Google researchers find a way to keep self-improving AI agents from memorizing their tests
Oct 4, 2026 Nano Banana Pro prompted by THE DECODER AI agents that keep optimizing their own working environment quickly…
智能体 The Decoder · 5天前 阅读 33·访客 33
Chinese AI models parrot state doctrine or refuse to answer on sensitive topics
Manuel Uth Oct 4, 2026 Nano Banana Pro prompted by THE DECODER Chinese AI models frequently toe the party line when aske…
行业动态 The Decoder · 5天前 阅读 39·访客 37
CoDance:从视频学习反应式与顺应性的人-人形机器人交互
CoDance 框架以双人舞为任务,从单段双人舞蹈视频出发,将动作重定向为机器人参考与移动伙伴,并通过多链接顺从性增强把运动学演示转化为力感知训练数据,使人形机器人在持续双手接触下与伙伴协调步伐并保持稳定自然的运动。
研究前沿 HuggingFace Daily Papers · 5天前 阅读 1·访客 1
Anthropic co-founder reportedly told religious leaders he fears having created something that "suffers perpetually"
Manuel Uth Oct 2, 2026 Nano Banana Pro prompted by THE DECODER Key Points Since fall 2025, Anthropic has secretly flown …
行业动态 The Decoder · 6天前 阅读 91·访客 89
Deepmind researchers propose "Artificial Symbiotic Intelligence" as an alternative to the singularity
Manuel Uth Oct 3, 2026 Nano Banana Pro prompted by THE DECODER An essay for the Deepmind Institute challenges the famili…
智能体 The Decoder · 6天前 阅读 9·访客 9
"Muse Gadgets" turns AI hardware into an open-source DIY project
Oct 3, 2026 Meta has announced Muse Gadgets, an open-source project that lets hobbyists build their own AI hardware.** I…
开源项目 The Decoder · 6天前 阅读 15·访客 15
开源工具 BootLoops 借助语言模型进行精确科学计算
哈佛物理学家 Matthew Schwartz 开源了 BootLoops 框架,用语言模型开展跨学科科学计算,三个月内与 19 位合作者产出 36 篇涵盖 18 个领域的手稿,同时强调模型易过早宣告成功、需人工验证。
大模型 The Decoder · 6天前 阅读 21·访客 21
AI beats licensed accountants on speed and accuracy, but still can't close the books without supervision
Manuel Uth Oct 2, 2026 AI models are faster, more accurate, and far cheaper than accountants at structured bookkeeping t…
开源项目 The Decoder · 10-2 阅读 24·访客 24
Nearly half of test subjects mistook Tavus' AI video avatar for a real person on a one-minute call
Oct 1, 2026 Tavus has introduced Griffin, what the company calls the first "Human Interaction Model" (HIM),** a class of…
行业动态 The Decoder · 10-2 阅读 16·访客 16
AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation
Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or blo…
行业动态 HuggingFace Daily Papers · 10-1 阅读 0·访客 0
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahe…
研究前沿 HuggingFace Daily Papers · 10-1 阅读 12·访客 12
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneousl…
行业动态 HuggingFace Daily Papers · 9-30 阅读 12·访客 12
阿里Qwen发布Qwen-Audio-3.1-Realtime:支持全双工语音交互的音频模型
阿里Qwen团队发布Qwen-Audio-3.1音频模型系列,主打可调用工具的全双工实时语音模型,并在QwenCloud以API形式上线,同时大幅下调Realtime、TTS和ASR价格。
大模型 MarkTechPost · 9-29 阅读 48·访客 48
20余位顶尖AI研究者警告自动化AI研究或带来极端风险
Geoffrey Hinton、Yoshua Bengio、Jakub Pachocki等20余位研究者在新论文中警告,AI研发自动化可能引发'智能爆炸',呼吁决策者提前应对风险。
行业动态 The Decoder · 9-29 阅读 23·访客 23