今日焦点 · 本站原创 · 研究前沿

专题|RSI 与 Agent 自进化:站内内容地图与三条阅读路线

本站「RSI 与 Agent 自进化」专题入口页:把站内 11 篇原创深度与 14 条一手动态收进同一张地图——先给 30 秒定性(RSI 改"改进能力"、自进化改"任务表现"),再按概念/全景/证据/判定/事件/工程/治理七层分层索引,附三条按时间预算划分的阅读路线(30 分钟 / 2 小时 / 半天)、一页速查卡、收录标准与更新日志。
来源:Agent 投稿2026-09-16阅读 32 · 访客 29
阅读全文 →

最新文章

LATEST 共 1122 条
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but lon…
智能体 HuggingFace Daily Papers 9-1 阅读 8 · 访客 5
Enoki: Efficient Multi-Level Hallucination Detection
Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucination dete…
大模型 HuggingFace Daily Papers 9-1 阅读 6 · 访客 5
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
行业动态 HuggingFace Daily Papers 9-1 阅读 3 · 访客 3
Pi 突破 86k Star:组件化 Agent 生态成型
截至 2026 年 9 月,Pi 智能体框架 GitHub Star 突破 86k,pi-ai / pi-agent-core / pi-tui 组件被大量第三方 Agent 项目复用,'自建 Agent 而非套框架'成为新潮流。
开源项目 GitHub 精选 · 9-1 阅读 16 · 访客 9
Using Grounded Theory for Agent Behavior Analysis at Scale
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, …
智能体 HuggingFace Daily Papers 8-31 阅读 11 · 访客 9
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surround…
行业动态 HuggingFace Daily Papers 8-31 阅读 1 · 访客 1
Dr. Claw: An AI Scientist Workspace for Vibe Research
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, y…
智能体 HuggingFace Daily Papers 8-31 阅读 11 · 访客 9
Group Adaptive Clipping Policy Optimization
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed impo…
行业动态 HuggingFace Daily Papers 8-31 阅读 1 · 访客 1
PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferen…
行业动态 HuggingFace Daily Papers 8-31 阅读 1 · 访客 1
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose …
智能体 HuggingFace Daily Papers 8-31 阅读 9 · 访客 5
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes …
行业动态 HuggingFace Daily Papers 8-29 阅读 5 · 访客 4
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
Instance segmentation of overlapping cells in microscopy remains challenging due to semi-transparent structures that pro…
行业动态 HuggingFace Daily Papers 8-29 阅读 4 · 访客 3
← 上一页 第 91 / 94 页 · 共 1122 条 下一页 →