Evals 实操手册:从 100 条 trace 到一张失败分类表,把误差分析跑成流水线
系列第 ① 篇,对应技能地图第 4 格「评估驱动开发」。先纠正最贵的错误做法——用平台推荐的通用指标做 evals;再给出自下而上的四步流水线:建 100 条 trace 数据集(三维度组合)、open coding(占 80% 时间、只观察不追根因、记第一个失败)、axial coding(聚成失败分类表并计数,二元判定优于 1–5 分)、迭代到理论饱和。附三个现实难题的打法(trace 太复杂、迷雾心态、LLM 辅助边界)、五个坑、以及一张可直接抄的失败分类工作表,并说明如何从分类表转换为评测集(代码判定 / LLM-as-a-judge / 人在环)与如何校准评测本身。
阅读全文 →
全部文章
ARCHIVE 共 128 条高通全球副总裁夏权:个人 AI 时代,每个人都拥有自己的智能体
IT之家 9 月 10 日消息,据新浪科技今日消息,高通公司全球副总裁夏权在 2026 服贸会表示,随着 AI 能力不断提升,人机交互正在发生新的变化。 过去人类需要依靠应用完成工作 , 未来人们将越来越多地与智能体协同工作 。 夏权指出,…
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing sk…
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat
Multi-Agent Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous sy…
Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that pre…
Memory as Plans: World-Action Modeling with Memory-Grounded Planning
Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inhe…
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks ar…
全球首个3D原生城市世界模型ABot-Earth 0.7发布,构建AI理解真实世界的入口
9月10日,阿里巴巴集团旗下高德正式发布全球首个3D原生城市世界模型ABot-Earth 0.7。]
Show-Harness: Just a VLM Agent Can Play Robots
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence i…
拿铁熊猫发布高性能 x86 计算模块 Mu Ultra,基于英特尔 "Lunar Lake" 处理器
IT之家 9 月 9 日消息,智位机器人 (DFRobot) 旗下拿铁熊猫 (LattePanda) 团队近日发布高性能 x86 计算模块(核心板)产品 Mu Ultra。其基于英特尔酷睿 Ultra 200V "Lunar Lake" 处…
TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents
GUI agents accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usa…
Studying Without a Syllabus: Task-Agnostic Environment Preprocessing
Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusabl…
ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of em…