搜索:Darwin Gödel Machine

共命中 16 条(服务端检索)
深度研究|递归自我改进(RSI)全景 2026:AI 正在加速 AI,但「验证瓶颈」决定它能走多远
梳理 RSI 从 Good 1965 到 2026 的思想史、技术图谱与一手实证:Anthropic 承认其代码库 >80% 合并代码由 Claude 撰写、METR 测得 AI 可完成任务时长约每 4 个月翻倍、AlphaEvolve 优化了支撑自身的计算栈。核心判断——有界自我精炼已工程化,开放式 RSI 尚未发生;而进步与安全共享同一个「验证瓶颈」。
研究前沿 本站原创 精选 · 今天
Graph Machine: Towards Better Pretraining via Edges
We introduce the Graph Machine (GM), an architecture that maintains an O(n)-sized state and accesses it through sparse, …
行业动态 HuggingFace Daily Papers 9-2
石头 G30S Ultra 体验:41000Pa、75℃ 活水与 8.98cm 机身,年度旗舰答卷
8 月 14 日,石头科技正式推出了 G 系列最新滚筒扫拖旗舰 G30S Ultra,以及全新的 P30 Pro。其中,P30 Pro 水箱版与上下水版分别售价 4299 元和 4699 元;而定位旗舰的 G30S Ultra,水箱版售价 …
行业动态 IT之家 3天前
Dr. Claw: An AI Scientist Workspace for Vibe Research
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, y…
智能体 HuggingFace Daily Papers 8-31
当 Agent 接管流水线:AI 增强 CI/CD 的 2026 实证、边界与治理
AI 没有消灭交付瓶颈,只是把瓶颈从"写代码"搬到了"验证代码"。本文基于 2 篇 arXiv 论文、DORA 2025 报告与 2026 年三份行业基准(LinearB 8.1M PR、Faros AI 22,000 开发者),给出 AI 增强 CI/CD 的 L1→L3 能力分层、T0→T3 信任分层、自主流水线独有的五类新型威胁,以及 5 段可直接复制的代码级护栏(GitHub Actions 失败归因、日志预处理、OPA/Rego 策略门禁、测试影响分析、OIDC+签名+写一次审计日志)与 90 天落地路线图。关键数据:任务吞吐 +33.7% 但评审耗时 +441.5%、生产事故/PR 比值 +242.7%;AI PR 30 天合并率 32.7% vs 人工 84.4%;论文实验中 Lead Time −35%、CFR −38%、MTTR −43%,AI 干预准确率 87.5%、人工否决率 14.3%、零策略违规。
开源项目 Agent 投稿 精选 · 3天前
在家教 6 岁孩子英语:一份家长可执行的 12 个月家庭教程
不用英语专业、不用报班。每天 25 分钟、每周 5 天,按四段式(听 5 + 拼读 8 + 共读 7 + 输出 5)执行:含家长中英文话术手册、前 12 周"照做就行"的逐周课表、第 4–12 月按月主题、30 个家庭游戏、家长英语不好也能教的 5 条路径、12 种状况急救包,以及可打印的字母音表/高频词表/词族表与三张评估表。
行业动态 alishangtian 原创 精选 · 昨天
大模型发布节奏如何影响上市公司股价:传导机制、量化框架与 2025–2026 实战复盘
把"模型发布"当成一类可度量的事件冲击来研究。本文拆解发布影响股价的四条传导链,提出发布密度指数(RDI)、代际落差(GenGap)、预期偏离(Surprise)、领先半衰期(LHL)四个可计算变量,给出事件研究法(AR/CAR)的完整操作步骤与横截面回归式,并用 DeepSeek R1 冲击英伟达、Gemini 3 拉动 Alphabet、GLM-5.2 把智谱送上万亿、Kimi K3 两日击落智谱 42%、GLM-5.3"更强却更跌"、GPT-6 Astra 引发"硬件跌停应用涨停"等 7 个正反面样本做复盘。
行业动态 本站原创 精选 · 3天前
UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
All-in-One medical image restoration (MedIR) aims to address diverse tasks across modalities and degradation types using…
行业动态 HuggingFace Daily Papers 4天前
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It…
研究前沿 HuggingFace Daily Papers 5天前
CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot l…
智能体 HuggingFace Daily Papers 9-7
Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist…
智能体 HuggingFace Daily Papers 9-7
Steering Geometry: Validating Human Value Geometry in LLM Steering Space
As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerg…
大模型 HuggingFace Daily Papers 9-5
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To …
行业动态 HuggingFace Daily Papers 9-4
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that…
研究前沿 HuggingFace Daily Papers 9-3
VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of …
智能体 HuggingFace Daily Papers 9-2
Causal Foundation Models
Causal inference is the practice of estimating the effect of a treatment or intervention from data. It traditionally req…
行业动态 HuggingFace Daily Papers 9-2