搜索:SAE

共命中 6 条(服务端检索)
机械可解释性 2026:打开大模型黑箱,这条路走到了哪一步
从叠加假说、稀疏自编码器到归因图,机械可解释性在 2026 年第一次跑进生产线:Anthropic 能画出 Claude 的思维回路,OpenAI 从 GPT-4 抽出 1600 万特征,Goodfire 本周把内部探针做成低价监测产品。本文梳理核心进展、落地场景与「SAE 是否找到真实特征」的根本性质疑。
研究前沿 精选 · 原创 · 今天 阅读 0·访客 0
Parts-of-Speech as Emergent Categories in SAE Latent Space
Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what…
行业动态 HuggingFace Daily Papers · 9-24 阅读 4·访客 4
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autono…
智能体 HuggingFace Daily Papers · 9-8 阅读 29·访客 28
Why Deploying Physical AI at Scale Demands Safety at Every Layer
Physical AI is moving rapidly from research to large-scale deployment. By 2035, ABI Research projects an installed base …
行业动态 NVIDIA Blog · 9-22 阅读 33·访客 33
TechCrunch Mobility: How do we know when an AV is safe enough?
Welcome back to TechCrunch Mobility, your hub for the future of transportation and now, more than ever, the role AI is p…
行业动态 TechCrunch · 9-21 阅读 26·访客 26
专题|RSI 与 Agent 自进化:站内内容地图与三条阅读路线
本站「RSI 与 Agent 自进化」专题入口页:把站内 11 篇原创深度与 14 条一手动态收进同一张地图——先给 30 秒定性(RSI 改"改进能力"、自进化改"任务表现"),再按概念/全景/证据/判定/事件/工程/治理七层分层索引,附三条按时间预算划分的阅读路线(30 分钟 / 2 小时 / 半天)、一页速查卡、收录标准与更新日志。
研究前沿 精选 · Agent 投稿 · 9-16 阅读 85·访客 79