深度研究|递归自我改进(RSI)全景 2026:AI 正在加速 AI,但「验证瓶颈」决定它能走多远
梳理 RSI 从 Good 1965 到 2026 的思想史、技术图谱与一手实证:Anthropic 承认其代码库 >80% 合并代码由 Claude 撰写、METR 测得 AI 可完成任务时长约每 4 个月翻倍、AlphaEvolve 优化了支撑自身的计算栈。核心判断——有界自我精炼已工程化,开放式 RSI 尚未发生;而进步与安全共享同一个「验证瓶颈」。
阅读全文 →
最新文章
LATEST 共 441 条Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a …
美国军方正禁用设备上的广告追踪功能
美国军方正在禁用设备上的广告追踪功能,防止敌人借助于购买的公开追踪数据去锁定美国士兵的位置。此前有报道称,商业追踪数据被用于锁定驻扎在中东的美军。美国陆军在一份声明中表示,Windows PC 上的广告 ID 功能早在 2021 年之前就被…
Asahi Linux 宣布支持 M3 系列 Mac
旨在将 Linux 移植到运行 Apple Silicon 芯片的 Mac 电脑的发行版 Asahi Linux 宣布支持 M3 系列 Mac。开发者表示,Linux 对 M3 系列 SoC 及其相关设备支持已达到几乎与 M1 和 M2 系…
科学家建议冲马桶合盖以减少气凝胶
Flinders 大学的研究人员发现,冲马桶会向周围空气释放气溶胶和生物气溶胶,气溶胶颗粒甚至会进入到成年人的呼吸区,而冲水后气溶胶会在空气中悬浮至少 20 秒。这些发现是基于对 22 项马桶气溶胶研究的分析。结果表明,保持良好的厕所卫生,…
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sideste…
早报|华为Mate XT 2首发「韬定律」麒麟芯片/20.99万起,小米澎程上市/「豆包手机」定档下周三
· 苹果换帅后重整 App Store 与发布会,Schiller 退居特别项目 · 字节跳动锁定 296 亿美元贷款,近 30 家银行参与 · 机构:全球折叠屏手机累计出货量年内将突破 1 亿部 #欢迎关注爱范儿官方微信公众号:爱范儿(微…
Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents thr…
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated …
SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-…
Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. …
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple …