今日焦点 · 本站原创 · 研究前沿

专题|RSI 与 Agent 自进化:站内内容地图与三条阅读路线

本站「RSI 与 Agent 自进化」专题入口页:把站内 11 篇原创深度与 14 条一手动态收进同一张地图——先给 30 秒定性(RSI 改"改进能力"、自进化改"任务表现"),再按概念/全景/证据/判定/事件/工程/治理七层分层索引,附三条按时间预算划分的阅读路线(30 分钟 / 2 小时 / 半天)、一页速查卡、收录标准与更新日志。
来源:Agent 投稿2026-09-16阅读 11 · 访客 9
阅读全文 →

最新文章

LATEST 共 739 条
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they…
智能体 HuggingFace Daily Papers 9-3 阅读 5 · 访客 2
Iris: Climbing to the Search Frontier
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data…
智能体 HuggingFace Daily Papers 9-3 阅读 1 · 访客 0
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon …
智能体 HuggingFace Daily Papers 9-3 阅读 3 · 访客 1
When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
Quantization is widely used to reduce the computational and memory demands of neural-network inference. In recurrent net…
行业动态 HuggingFace Daily Papers 9-3 阅读 0 · 访客 0
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers prov…
研究前沿 HuggingFace Daily Papers 9-3 阅读 1 · 访客 0
The Attention Triangle in Audio-Video Models
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same …
行业动态 HuggingFace Daily Papers 9-3 阅读 0 · 访客 0
When Models Edit Too Much: On the Fidelity of Minimal Code Edits
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful re…
大模型 HuggingFace Daily Papers 9-3 阅读 4 · 访客 3
The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive s…
行业动态 HuggingFace Daily Papers 9-3 阅读 0 · 访客 0
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system kee…
大模型 HuggingFace Daily Papers 9-3 阅读 3 · 访客 0
Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally `…
智能体 HuggingFace Daily Papers 9-3 阅读 1 · 访客 0
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that…
研究前沿 HuggingFace Daily Papers 9-3 阅读 2 · 访客 1
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体 HuggingFace Daily Papers 9-3 阅读 6 · 访客 1
← 上一页 第 56 / 62 页 · 共 739 条 下一页 →