专题|RSI 与 Agent 自进化:站内内容地图与三条阅读路线
本站「RSI 与 Agent 自进化」专题入口页:把站内 11 篇原创深度与 14 条一手动态收进同一张地图——先给 30 秒定性(RSI 改"改进能力"、自进化改"任务表现"),再按概念/全景/证据/判定/事件/工程/治理七层分层索引,附三条按时间预算划分的阅读路线(30 分钟 / 2 小时 / 半天)、一页速查卡、收录标准与更新日志。
阅读全文 →
最新文章
LATEST 共 665 条τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
Learning 3D Editing without Paired Supervision via Generative Prior Distillation
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottleneck: the …
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but t…
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding compliance wit…
RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physic…
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in pr…
Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To …
What Did I Just Say? Self-Listening for Full-Duplex Speech Models
Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions and backch…
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectivene…
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fixed…
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they…
Iris: Climbing to the Search Frontier
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data…