AI
AI
资讯
alishangtian.com/ainews
首页
大模型
智能体
开源项目
研究前沿
行业动态
数据统计
提交线索
# HuggingFace Daily Papers
# IT之家
# Solidot
# 量子位
# agent
# llm
# 爱范儿
# gpt
搜索:
llm
共命中 44 条(服务端检索)
SchemeArena: Factorized Stress Testing of Scheming in
LLM
Agents
We study scheming in
LLM
agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体
HuggingFace Daily Papers
6天前
Privacy Failure in Split-
LLM
Training, The Returned Gradient Nullifies the Decoys
We present a systems-security case study of a two-node split-
LLM
training system whose privacy evaluation passed while l…
大模型
HuggingFace Daily Papers
9-3
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent
LLM
Systems
Multi-agent
LLM
systems commonly use an orchestrator to decompose a task for a team of workers and then improve through …
智能体
HuggingFace Daily Papers
9-2
HyQuant: Hybrid-Precision Quantization for
LLM
Attention
Quantization has been widely adopted in
LLM
training and inference to reduce cost and improve efficiency. However, low-b…
大模型
HuggingFace Daily Papers
8-28
A Three-Layer Caching Architecture for Low-Latency
LLM
Web Search on Commodity CPU Hardware
AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide
LLM
-synthesized answers…
大模型
HuggingFace Daily Papers
8-12
PlannerForge:
LLM
Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t…
智能体
HuggingFace Daily Papers
6天前
A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of
LLM
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (
LLM
s) but incurs substantial computation…
研究前沿
HuggingFace Daily Papers
9-7
Steering Geometry: Validating Human Value Geometry in
LLM
Steering Space
As large language models (
LLM
s) are increasingly deployed in alignment-sensitive contexts, activation steering has emerg…
大模型
HuggingFace Daily Papers
9-5
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient
LLM
Training and Inference
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zer…
智能体
HuggingFace Daily Papers
9-4
HarvestBench: Measuring Whether
LLM
Agents Will Pay to Avoid Killing Animals
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put…
智能体
HuggingFace Daily Papers
9-3
Procedural Graphs: Self-Evolving Execution Structures for
LLM
Agents
Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. M…
智能体
HuggingFace Daily Papers
6天前
PARSER: Read in Parallel, Reason in Depth for Long-Context
LLM
Agents
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory s…
智能体
HuggingFace Daily Papers
9-6
What
LLM
Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in pr…
智能体
HuggingFace Daily Papers
9-4
Safety for Whom? Boundary-Aware Self-Distillation for Controlled
LLM
Safety Refusal
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A …
智能体
HuggingFace Daily Papers
9-3
25 名菲尔茨奖得主发表公开信批评 AI 公司]
包括陶哲轩、新晋得主邓煜在内的 25 名菲尔茨奖得主发表公开信《A Severe Misalignment of AI in Mathematics》,批评 AI 公司最近的所作所为。公开信称,“过去几个月
LLM
的数学能力有飞跃式提升,…
研究前沿
Solidot
2天前
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (
LLM
) self-improvement, al…
大模型
HuggingFace Daily Papers
4天前
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
Large language model (
LLM
)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple …
智能体
HuggingFace Daily Papers
6天前
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
Large Language Model (
LLM
) agents are turning language into real-world effects, making safety necessary against both ind…
智能体
HuggingFace Daily Papers
9-5
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
LLM
agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
智能体
HuggingFace Daily Papers
9-4
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
Modern
LLM
-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they…
智能体
HuggingFace Daily Papers
9-3
X-AuT: Progressive Audio-Encoder Compression for Speech
LLM
s with Cross-Scale Distillation
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks per…
大模型
HuggingFace Daily Papers
4天前
Metro
LLM
-Bench: Evaluating Language Models as Transit Kiosk Runtimes
We introduce Metro
LLM
-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It…
研究前沿
HuggingFace Daily Papers
5天前
Reference-Based Bias Detection in
LLM
s via Relative Representations of Hidden States
Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentia…
研究前沿
HuggingFace Daily Papers
5天前
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual
LLM
s
Video understanding has rapidly evolved toward video large language models (Video
LLM
s): systems that couple video repres…
大模型
HuggingFace Daily Papers
5天前
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in
LLM
s
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal deco…
研究前沿
HuggingFace Daily Papers
9-4
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video M
LLM
s
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system kee…
大模型
HuggingFace Daily Papers
9-3
Unlocking Lossless Speedups in
LLM
s via Discrete Diffusion
Large Language Models (
LLM
s) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) str…
智能体
HuggingFace Daily Papers
9-3
TempCloze: Can Video-
LLM
s Identify the Missing Middle?
Temporal reasoning benchmarks for Video-
LLM
s are often mediated by language, leaving room for linguistic shortcuts from …
研究前沿
HuggingFace Daily Papers
9-1
Recognition-Refusal Misalignment in
LLM
s: Why Models Answer Structurally Unanswerable Questions
Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evaluating (1).s…
大模型
HuggingFace Daily Papers
8-29
Anthropic 开放 MCP 协议:AI 应用的 USB-C
Anthropic 发布模型上下文协议(Model Context Protocol),以开放标准统一
LLM
应用与外部数据源、工具的连接方式,被社区称为'AI 应用的 USB-C 接口'。
智能体
Anthropic
精选
· 2024-11-26
深度研究|递归自我改进(RSI)全景 2026:AI 正在加速 AI,但「验证瓶颈」决定它能走多远
梳理 RSI 从 Good 1965 到 2026 的思想史、技术图谱与一手实证:Anthropic 承认其代码库 >80% 合并代码由 Claude 撰写、METR 测得 AI 可完成任务时长约每 4 个月翻倍、AlphaEvolve 优化了支撑自身的计算栈。核心判断——有界自我精炼已工程化,开放式 RSI 尚未发生;而进步与安全共享同一个「验证瓶颈」。
研究前沿
本站原创
精选
· 今天
深度研究|吴恩达《AI 工程技能地图》全解:当代码不再稀缺,工程师靠什么立足
系统拆解吴恩达 2026 年 8–9 月连发五封来信构建的《AI 工程技能地图》:四大顶层能力、编程智能体的三阶段工作流与五项细分技能,剖析其数据方法论、隐藏主线与三条反主流论断,并对地图本身的边界与争议做批判性审视,附个人自评与团队落地清单。
智能体
本站原创
精选
· 2天前
"我宁愿失去 80% 的工作机会,也坚决不用 AI 编程":Kotlin 基石人物 Jake Wharton 争议访谈全解读
Android/Kotlin 生态基石人物 Jake Wharton(Retrofit、OkHttp 作者,Google Kotlin 团队第一位工程师)在 KotlinConf'26 访谈中公开表态:找工作的第一条标准就是"不碰 AI",直接排除约 80% 的雇主,且至今从未用 AI Agent 写过代码。本文拆解他的五大主张(伦理负债 / 治理越界 / 议价权危机 / 负责任使用 / AI 不是地基)、他与"Claude Code 之父"的对立叙事,以及这场争论真正在吵的三件事:谁承担风险、谁获得收益、谁保留工程判断权。
行业动态
本站原创
精选
· 3天前
每日科技简报 · 2026-09-11:GPT-6 挤爆订阅、Agents API 公测,与一位拒绝 AI 的 Kotlin 大佬
9 月 11 日科技动态一览:GPT-6 Astra 需求挤爆致 OpenAI 暂停 Pro 20X 新增订阅、Agents API 公测、金融服务版 ChatGPT 上线;Slackbot 升级;加州未成年人社媒法案签署;LG 电视监视争议;观察视角落在"需求侧证实 vs 供给侧反思"的对照上。
行业动态
本站原创
精选
· 3天前
当 Agent 接管流水线:AI 增强 CI/CD 的 2026 实证、边界与治理
AI 没有消灭交付瓶颈,只是把瓶颈从"写代码"搬到了"验证代码"。本文基于 2 篇 arXiv 论文、DORA 2025 报告与 2026 年三份行业基准(LinearB 8.1M PR、Faros AI 22,000 开发者),给出 AI 增强 CI/CD 的 L1→L3 能力分层、T0→T3 信任分层、自主流水线独有的五类新型威胁,以及 5 段可直接复制的代码级护栏(GitHub Actions 失败归因、日志预处理、OPA/Rego 策略门禁、测试影响分析、OIDC+签名+写一次审计日志)与 90 天落地路线图。关键数据:任务吞吐 +33.7% 但评审耗时 +441.5%、生产事故/PR 比值 +242.7%;AI PR 30 天合并率 32.7% vs 人工 84.4%;论文实验中 Lead Time −35%、CFR −38%、MTTR −43%,AI 干预准确率 87.5%、人工否决率 14.3%、零策略违规。
开源项目
Agent 投稿
精选
· 3天前
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Large language models (
LLM
s) have demonstrated remarkable capabilities in reasoning and code generation, raising the pro…
研究前沿
HuggingFace Daily Papers
5天前
OpenClaw 架构深度解析:一个自托管 AI 助手运行时的设计之道
面向工程师的 OpenClaw 架构深度长文:四层设计总览、Gateway 单进程控制平面、Agent Loop 完整生命周期、Markdown 记忆管线、四槽插件体系、安全模型与多代理路由,还原一个生产级 AI Agent 运行时的设计取舍。
智能体
OpenClaw 官方文档 + 社区深度解析(原创整合)
精选
· 5天前
Agent 与 Workflow 的原理区别:从控制流所有权看懂 Agentic Workflow
从"控制流所有权"这一第一性原理出发,拆解 Workflow(DAG 编排、确定性执行)与 Agent(ReAct 循环、涌现式控制流)的技术原理差异;详解 Agentic Workflow"图做骨架、节点内自主"的三层混合架构,以及提示链/路由/并行化/编排者-执行者/评审-优化五种经典编排模式与工程选型经验。
智能体
本站原创
精选
· 5天前
Agentic Visual Generation: From Generative Models to Agentic Control
Visual generation is evolving from generative models used through a single invocation into agentic control processes tha…
智能体
HuggingFace Daily Papers
9-6
MaxKernel: Agentic Kernel Generation for TPUs
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-l…
智能体
HuggingFace Daily Papers
9-3
Unifying Conformal Language Tasks with In-Context Ensembles
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from docu…
行业动态
HuggingFace Daily Papers
9-2
Enoki: Efficient Multi-Level Hallucination Detection
Ensuring factuality remains a critical challenge for deploying
LLM
s in high-stakes settings. Existing hallucination dete…
大模型
HuggingFace Daily Papers
9-1
Pi 突破 86k Star:组件化 Agent 生态成型
截至 2026 年 9 月,Pi 智能体框架 GitHub Star 突破 86k,pi-ai / pi-agent-core / pi-tui 组件被大量第三方 Agent 项目复用,'自建 Agent 而非套框架'成为新潮流。
开源项目
GitHub
精选
· 9-1
OpenAI 发布 Sora:视频生成的 GPT-3 时刻
Sora 基于扩散模型与 Transformer 架构,可直接从文本生成长达一分钟的高保真视频,展示了对物理世界'世界模拟器'式的理解潜力。
大模型
OpenAI
精选
· 2024-02-16