深度研究|递归自我改进(RSI)全景 2026:AI 正在加速 AI,但「验证瓶颈」决定它能走多远
梳理 RSI 从 Good 1965 到 2026 的思想史、技术图谱与一手实证:Anthropic 承认其代码库 >80% 合并代码由 Claude 撰写、METR 测得 AI 可完成任务时长约每 4 个月翻倍、AlphaEvolve 优化了支撑自身的计算栈。核心判断——有界自我精炼已工程化,开放式 RSI 尚未发生;而进步与安全共享同一个「验证瓶颈」。
阅读全文 →
研究前沿
共 68 条Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal deco…
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but t…
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers prov…
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that…
TempCloze: Can Video-LLMs Identify the Missing Middle?
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from …
诺贝尔物理学奖授予神经网络先驱 Hopfield 与 Hinton
2024 年诺贝尔物理学奖授予 John Hopfield 与 Geoffrey Hinton,表彰其利用人工神经网络实现机器学习的基础性发现;化学奖同时颁给蛋白质计算先驱。
OpenAI o1 问世:推理时计算开启新范式
o1 系列通过强化学习训练模型'先思考再回答',在数学、代码与科学推理上大幅跃升,开创了推理时扩展(Test-time Compute)的新 Scaling 维度。
AlphaFold 3 发布:AI 破解生命分子相互作用
Google DeepMind 与 Isomorphic Labs 发布 AlphaFold 3,首次将预测范围从蛋白质结构扩展到 DNA、RNA、配体等全部生命分子及其相互作用。
← 上一页
第 6 / 6 页 · 共 68 条