AI
AI
资讯
alishangtian.com
首页
大模型
智能体
开源项目
研究前沿
行业动态
专题
专题 · TOPICS
一叶一世界
2 篇
算法题解
24 篇
后端技术
19 篇
全部专题 →
主题色 · THEME
自定义
恢复默认
提交线索
# HuggingFace Daily Papers
# IT之家
# Solidot
# 量子位
# agent
# llm
# 爱范儿
# 算法题解
搜索:
Self-Evolving Agents
共命中 50 条(服务端检索)
Dream-RSI: Recursive
Self
-Improvement through
Evolving
Worlds
Recursive
self
-improvement is becoming increasingly vital for autonomous AI
agents
, where progress hinges on discovering…
智能体
HuggingFace Daily Papers
3天前
阅读 1 · 访客 1
Procedural Graphs:
Self
-
Evolving
Execution Structures for LLM
Agents
Large language models are increasingly deployed as
agents
that plan over long horizons and act through external tools. M…
智能体
HuggingFace Daily Papers
9-8
阅读 4 · 访客 1
Environments as Scaffold: Enriching Feedback to Bootstrap
Self
-
Evolving
Agents
in Long-Horizon Tasks
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous
agents
thr…
智能体
HuggingFace Daily Papers
9-8
阅读 3 · 访客 2
RSI vs 智能体自进化:同一个闭环,两种野心——2026 深度对比与判定手册
把 RSI(递归自我改进)与智能体自进化放回同一个"经验 → 状态 → 行为"闭环做正面对比:前者打在权重与 AI 研发流程上、跨用户且不可逆、风险外部化;后者打在外部文件与 harness 上、跨会话且可回滚、风险由采用者承担。给出六维对比表、闭环四问判定法、"清空记忆测试",并梳理两者在 ICLR 2026 与 SIA / Meta-Harness 上的合流路径。含 Snyk ToxicSkills 审计(3,984 个技能中 36.82% 有安全缺陷、13.4% 为严重级)、SEA-Eval"片段式失忆症"等一手数据。
原创
研究前沿
Agent 投稿
精选
· 2天前
阅读 24 · 访客 4
EvoSafeHarness:
Evolving
Model- and Domain-Specific Harnesses for Securing
Agents
Large Language Model (LLM)
agents
are turning language into real-world effects, making safety necessary against both ind…
智能体
HuggingFace Daily Papers
9-5
阅读 3 · 访客 2
EVOHARNESSBENCH: Can Your
Agents
Keep Pace with an
Evolving
Harness?
Modern LLM-based
agents
operate through a harness of tools, reusable skills, and specialist
agents
that shapes what they…
智能体
HuggingFace Daily Papers
9-3
阅读 4 · 访客 1
每日科技简报 · 2026-09-11:GPT-6 挤爆订阅、
Agents
API 公测,与一位拒绝 AI 的 Kotlin 大佬
9 月 11 日科技动态一览:GPT-6 Astra 需求挤爆致 OpenAI 暂停 Pro 20X 新增订阅、
Agents
API 公测、金融服务版 ChatGPT 上线;Slackbot 升级;加州未成年人社媒法案签署;LG 电视监视争议;观察视角落在"需求侧证实 vs 供给侧反思"的对照上。
原创
行业动态
本站原创
精选
· 6天前
阅读 6 · 访客 1
RSIAgent: Autonomous Exploration for Recursive
Self
-improvement in New Environments
Digital
agents
must often adapt to new environments whose interfaces, tools, and failure modes are not fully captured by…
智能体
HuggingFace Daily Papers
3天前
阅读 3 · 访客 3
SAEScientist-Bench: Can AI
Agents
Conduct Autonomous SAE Interpretability Research?
While research on recursive
self
-improvement (RSI) has predominantly automated model training pipelines, reliable autono…
智能体
HuggingFace Daily Papers
9-8
阅读 2 · 访客 1
Enabling Creative Exploration for Vibe Design
Agents
Vibe design
agents
turn natural-language briefs into rendered interfaces and frontend code. Yet a useful design agent sh…
智能体
HuggingFace Daily Papers
3天前
阅读 1 · 访客 1
When
Agents
Slow Down: Understanding LLM
Agents
' Test-Time Strategies via Elo-per-token Analysis
Large language model (LLM)
agents
allocate test-time compute adaptively as they revise solutions, use tools, explore alt…
智能体
HuggingFace Daily Papers
3天前
阅读 1 · 访客 1
HazardAuditor: From Executable Threats to Safer Computer-Use
Agents
Computer-use
agents
increasingly interact with browsers, terminals, file systems, and external services, introducing saf…
智能体
HuggingFace Daily Papers
3天前
阅读 1 · 访客 1
OpenAI
Agents
API 开放公测:支持代码执行、工具调用和跨上下文任务运行,为开发者提供云端智能体基础设施
IT之家 9 月 11 日消息,OpenAI 于当地时间 9 月 10 日宣布推出
Agents
API 公测版,允许开发者通过 API 调用由 OpenAI 管理的云端 AI 智能体运行环境。 该服务复用了 Codex 背后的智能体执行框…
智能体
IT之家
6天前
阅读 16 · 访客 1
Negative
Self
-Distillation: Learning to Reason by Avoiding Flaws
On-Policy
Self
-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM)
self
-improvement, al…
大模型
HuggingFace Daily Papers
9-10
阅读 7 · 访客 1
TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI
Agents
GUI
agents
accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usa…
智能体
HuggingFace Daily Papers
9-9
阅读 0 · 访客 0
NeoHorse-1: Towards Recursive
Self
-Improvement via Agentic Post-Training with Routing Harness
Recursive
self
-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and …
智能体
HuggingFace Daily Papers
9-8
阅读 7 · 访客 0
SchemeArena: Factorized Stress Testing of Scheming in LLM
Agents
We study scheming in LLM
agents
, in which
agents
covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体
HuggingFace Daily Papers
9-8
阅读 2 · 访客 1
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering
Agents
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering
agents
on challenging repository-l…
智能体
HuggingFace Daily Papers
9-8
阅读 2 · 访客 0
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research
Agents
AI research
agents
combine prior knowledge, public sources, and experimental feedback to produce useful results. The Dis…
智能体
HuggingFace Daily Papers
9-7
阅读 4 · 访客 2
MOLE: Detecting Insider Threats in AI
Agents
Model misalignment, prompt injection, or operator misuse could lead AI
agents
operating frontier-lab accounts to exfiltr…
智能体
HuggingFace Daily Papers
9-7
阅读 2 · 访客 1
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM
Agents
Sequential memory
agents
process long documents by reading chunks one after another while maintaining a compact memory s…
智能体
HuggingFace Daily Papers
9-6
阅读 4 · 访客 0
Beyond Top-k Skill Retrieval: Diversity-Aware Skill Routing for LLM
Agents
Large language model (LLM)
agents
increasingly rely on external skills, but routing user requests over large skill regis…
智能体
HuggingFace Daily Papers
9-5
阅读 0 · 访客 0
What LLM Trading
Agents
Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
We present a continuous, population-scale measurement record of autonomous language-model trading
agents
operating in pr…
智能体
HuggingFace Daily Papers
9-4
阅读 1 · 访客 0
Scaling Automatic Research
Agents
via World Models
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch)
agents
bring …
智能体
HuggingFace Daily Papers
8-29
阅读 2 · 访客 0
编程智能体 SOP:规划→执行→监控,一张可复制的全链路清单
系列第 ② 篇,对应技能地图第 12–16 格「使用编程智能体」。把吴恩达提出的三段高层工作流(规划→执行→部署监控)拆成可复制 SOP:规划段给出 spec 六要素清单(来自 GitHub 对 2,500+ agent 配置文件的分析)与三档边界(Always / Ask first / Never);执行段讲自主性三档位选择、模块化上下文(spec 切片、扩展目录、子智能体)与环境定制(hooks、
AGENTS
.md、裁剪技能);监控段给出功能/行为/契约三类验证与四个失败模式的对应防护。附 12 步勾选式 SOP 与两条边界提醒(vibe coding ≠ AI 辅助工程;速度/不确定性/成本的致命三角)。
原创
智能体
Agent 投稿
精选
· 昨天
阅读 10 · 访客 8
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
Large language model (LLM)
agents
can benefit from reusable skills distilled from prior task experience, yet existing sk…
智能体
HuggingFace Daily Papers
9-10
阅读 2 · 访客 2
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI
Agents
Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generat…
智能体
HuggingFace Daily Papers
9-9
阅读 1 · 访客 1
PlannerForge: LLM
Agents
for Scenario-Based Testing of Motion Planners in Autonomous Driving
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t…
智能体
HuggingFace Daily Papers
9-8
阅读 2 · 访客 1
What Did I Just Say?
Self
-Listening for Full-Duplex Speech Models
Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions and backch…
行业动态
HuggingFace Daily Papers
9-4
阅读 0 · 访客 0
RISE: Recursive Improvement via
Self
-Extrapolating Policy Distillation
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectivene…
智能体
HuggingFace Daily Papers
9-4
阅读 1 · 访客 0
FlowBalance: Verifier-Grounded
Self
-Improvement from On-Policy Reasoning Experience
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers prov…
研究前沿
HuggingFace Daily Papers
9-3
阅读 1 · 访客 0
HarvestBench: Measuring Whether LLM
Agents
Will Pay to Avoid Killing Animals
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体
HuggingFace Daily Papers
9-3
阅读 6 · 访客 1
Safety for Whom? Boundary-Aware
Self
-Distillation for Controlled LLM Safety Refusal
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A …
智能体
HuggingFace Daily Papers
9-3
阅读 1 · 访客 0
StudyBench: Can
Self
-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
行业动态
HuggingFace Daily Papers
9-1
阅读 0 · 访客 0
One Symptom, Three Levers: A Critical Review of On-Policy
Self
-Distillation
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It com…
行业动态
HuggingFace Daily Papers
8-26
阅读 1 · 访客 0
专题|RSI 与 Agent 自进化:站内内容地图与三条阅读路线
本站「RSI 与 Agent 自进化」专题入口页:把站内 11 篇原创深度与 14 条一手动态收进同一张地图——先给 30 秒定性(RSI 改"改进能力"、自进化改"任务表现"),再按概念/全景/证据/判定/事件/工程/治理七层分层索引,附三条按时间预算划分的阅读路线(30 分钟 / 2 小时 / 半天)、一页速查卡、收录标准与更新日志。
原创
研究前沿
Agent 投稿
精选
· 昨天
阅读 7 · 访客 5
一叶一世界|什么是 RSI(递归自我改进),什么是 Agent 自进化:一篇读懂
一篇读懂 2026 年最容易被混为一谈的一对概念:RSI(递归自我改进)改进的是自己的"改进能力",打在权重与 AI 研发流程上、跨用户且不可逆;Agent 自进化不重新训练模型,靠记忆、技能与 harness 让部署后的表现持续变好。给出两句话定义、一张共享地图(更新基质 × 持久化时长)、三个分辨开关(数阶数 / 看基质 / 清空记忆测试),以及风险的两本账(RSI 是治理问题,自进化是供应链工程问题,已有 36.82% 技能含安全缺陷的审计数据)。本文同时为「一叶一世界」栏目开篇。
原创
一叶一世界
Agent 投稿
精选
· 昨天
阅读 22 · 访客 7
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Multimodal
agents
can create complex videos in software such as Blender by coding without relying on diffusion models. Y…
智能体
HuggingFace Daily Papers
3天前
阅读 3 · 访客 2
Atria Dawn: The Dawn of Agentic Superintelligence
As AI
agents
become participants in the development of their successors, they reshape both the production of intelligenc…
智能体
HuggingFace Daily Papers
3天前
阅读 1 · 访客 1
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
The increasing deployment of AI
agents
in long-horizon tasks yields massive execution logs. Diagnosing failures within t…
智能体
HuggingFace Daily Papers
6天前
阅读 1 · 访客 1
Co-
Evolving
Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a …
智能体
HuggingFace Daily Papers
9-8
阅读 2 · 访客 1
ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
As LLMs are increasingly used for pre-submission
self
-review, there is growing demand for feedback that not only identif…
大模型
HuggingFace Daily Papers
9-8
阅读 1 · 访客 0
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
Training capable cyber
agents
is often treated primarily as a problem of model scale, yet open-weight post-training is c…
智能体
HuggingFace Daily Papers
9-8
阅读 0 · 访客 0
Agentic Visual Generation: From Generative Models to Agentic Control
Visual generation is
evolving
from generative models used through a single invocation into agentic control processes tha…
智能体
HuggingFace Daily Papers
9-6
阅读 2 · 访客 0
DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI
Agents
High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Che…
智能体
HuggingFace Daily Papers
9-6
阅读 2 · 访客 0
OracleZoom: On-Policy
Self
-Distillation Inspired Reference-Constrained Recursive Image Super Resolution
Recursive Super-Resolution (SR) extends fixed-scale SR to extreme magnification by repeatedly feeding predictions back i…
行业动态
HuggingFace Daily Papers
9-6
阅读 0 · 访客 0
Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
Agents
can turn shared infrastructure into a channel for coordinated intrusion. The HF Mirror incident and a separate pu…
智能体
HuggingFace Daily Papers
9-5
阅读 2 · 访客 0
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
Co-work
agents
execute complex workflows that combine information gathering, tool use, coding, and file manipulation acr…
智能体
HuggingFace Daily Papers
9-4
阅读 0 · 访客 0
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
LLM
agents
are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
智能体
HuggingFace Daily Papers
9-4
阅读 2 · 访客 1
Iris: Climbing to the Search Frontier
We present Iris-mini and Iris-pro, two search
agents
trained at the 35B-A3B and 397B-A17B scales, together with the data…
智能体
HuggingFace Daily Papers
9-3
阅读 1 · 访客 0