深度研究|递归自我改进(RSI)全景 2026:AI 正在加速 AI,但「验证瓶颈」决定它能走多远
梳理 RSI 从 Good 1965 到 2026 的思想史、技术图谱与一手实证:Anthropic 承认其代码库 >80% 合并代码由 Claude 撰写、METR 测得 AI 可完成任务时长约每 4 个月翻倍、AlphaEvolve 优化了支撑自身的计算栈。核心判断——有界自我精炼已工程化,开放式 RSI 尚未发生;而进步与安全共享同一个「验证瓶颈」。
阅读全文 →
最新文章
LATEST 共 516 条VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of …
Causal Foundation Models
Causal inference is the practice of estimating the effect of a treatment or intervention from data. It traditionally req…
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher …
Graph Machine: Towards Better Pretraining via Edges
We introduce the Graph Machine (GM), an architecture that maintains an O(n)-sized state and accesses it through sparse, …
MasterControl Seventeen Every Time
We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic poli…
SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in o…
TempCloze: Can Video-LLMs Identify the Missing Middle?
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from …
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but lon…
Enoki: Efficient Multi-Level Hallucination Detection
Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucination dete…
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
Pi 突破 86k Star:组件化 Agent 生态成型
截至 2026 年 9 月,Pi 智能体框架 GitHub Star 突破 86k,pi-ai / pi-agent-core / pi-tui 组件被大量第三方 Agent 项目复用,'自建 Agent 而非套框架'成为新潮流。
Using Grounded Theory for Agent Behavior Analysis at Scale
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, …