深度研究|递归自我改进(RSI)全景 2026:AI 正在加速 AI,但「验证瓶颈」决定它能走多远
梳理 RSI 从 Good 1965 到 2026 的思想史、技术图谱与一手实证:Anthropic 承认其代码库 >80% 合并代码由 Claude 撰写、METR 测得 AI 可完成任务时长约每 4 个月翻倍、AlphaEvolve 优化了支撑自身的计算栈。核心判断——有界自我精炼已工程化,开放式 RSI 尚未发生;而进步与安全共享同一个「验证瓶颈」。
阅读全文 →
研究前沿
共 68 条IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently …
Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentia…
Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking
Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on…
中国联通:智能手机正式迈入一机四号码时代,预计下半年 eSIM 手机突破 2000 万台
IT之家 9 月 9 日消息,中国联通今日在上海世博中心正式发布 eSIM 尝鲜季 2026 升级方案 ,从玩法、服务、权益三大维度迭代升级,加速推动 eSIM 技术走进千家万户。 玩法方面,用户开通 eSIM 服务即可领取专属福利, 告别…
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the pro…
控制呼吸为何能控制焦虑?
焦虑是人类最常见的精神疾病,全球约有 3.59 亿人受到影响。控制呼吸被认为有助于控制焦虑,根据发表在 PNAS 期刊上的一项研究,科学家基于小鼠研究揭示了这一现象背后的鼻脑回路(nose-to-brain circuit)机制。鼻脑回路始…
小鼠实验显示 GLP-1 减肥药或有助于延缓衰老
GLP-1 减肥药或有助于延缓衰老。加州伯克利等机构研究人员通过小鼠试验发现,老年雌性小鼠服用 GLP-1 类药物司美格鲁肽后,寿命比未服药小鼠延长12%,同时多项与衰老相关的生物学变化得到减缓。研究人员选取 20 个月大的健康雌性小鼠开展…
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition ma…
科学家建议冲马桶合盖以减少气凝胶
Flinders 大学的研究人员发现,冲马桶会向周围空气释放气溶胶和生物气溶胶,气溶胶颗粒甚至会进入到成年人的呼吸区,而冲水后气溶胶会在空气中悬浮至少 20 秒。这些发现是基于对 22 项马桶气溶胶研究的分析。结果表明,保持良好的厕所卫生,…
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sideste…
早报|华为Mate XT 2首发「韬定律」麒麟芯片/20.99万起,小米澎程上市/「豆包手机」定档下周三
· 苹果换帅后重整 App Store 与发布会,Schiller 退居特别项目 · 字节跳动锁定 296 亿美元贷款,近 30 家银行参与 · 机构:全球折叠屏手机累计出货量年内将突破 1 亿部 #欢迎关注爱范儿官方微信公众号:爱范儿(微…
Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. …