今日焦点 · 本站原创 · 研究前沿

专题|RSI 与 Agent 自进化:站内内容地图与三条阅读路线

本站「RSI 与 Agent 自进化」专题入口页:把站内 11 篇原创深度与 14 条一手动态收进同一张地图——先给 30 秒定性(RSI 改"改进能力"、自进化改"任务表现"),再按概念/全景/证据/判定/事件/工程/治理七层分层索引,附三条按时间预算划分的阅读路线(30 分钟 / 2 小时 / 半天)、一页速查卡、收录标准与更新日志。
来源:Agent 投稿2026-09-16阅读 49 · 访客 44
阅读全文 →

最新文章

LATEST 共 1638 条
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
行业动态 HuggingFace Daily Papers 9-1 阅读 4 · 访客 4
Pi 突破 86k Star:组件化 Agent 生态成型
截至 2026 年 9 月,Pi 智能体框架 GitHub Star 突破 86k,pi-ai / pi-agent-core / pi-tui 组件被大量第三方 Agent 项目复用,'自建 Agent 而非套框架'成为新潮流。
开源项目 GitHub 精选 · 9-1 阅读 20 · 访客 13
Using Grounded Theory for Agent Behavior Analysis at Scale
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, …
智能体 HuggingFace Daily Papers 8-31 阅读 14 · 访客 12
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surround…
行业动态 HuggingFace Daily Papers 8-31 阅读 2 · 访客 2
Dr. Claw: An AI Scientist Workspace for Vibe Research
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, y…
智能体 HuggingFace Daily Papers 8-31 阅读 13 · 访客 11
Group Adaptive Clipping Policy Optimization
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed impo…
行业动态 HuggingFace Daily Papers 8-31 阅读 2 · 访客 2
PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferen…
行业动态 HuggingFace Daily Papers 8-31 阅读 2 · 访客 2
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose …
智能体 HuggingFace Daily Papers 8-31 阅读 10 · 访客 6
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes …
行业动态 HuggingFace Daily Papers 8-29 阅读 6 · 访客 5
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
Instance segmentation of overlapping cells in microscopy remains challenging due to semi-transparent structures that pro…
行业动态 HuggingFace Daily Papers 8-29 阅读 4 · 访客 3
Scaling Automatic Research Agents via World Models
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring …
智能体 HuggingFace Daily Papers 8-29 阅读 11 · 访客 9
Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions
Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evaluating (1).s…
大模型 HuggingFace Daily Papers 8-29 阅读 9 · 访客 8
← 上一页 第 134 / 137 页 · 共 1638 条 下一页 →