AI
AI
资讯
alishangtian.com/ainews
首页
大模型
智能体
开源项目
研究前沿
行业动态
数据统计
提交线索
# HuggingFace Daily Papers
# IT之家
# Solidot
# 量子位
# agent
# llm
# 爱范儿
# gpt
搜索:
benchmark污染
共命中 33 条(服务端检索)
SWE-Bench Pro Verified: A Reliable
Benchmark
for Software Engineering Agents
SWE-Bench Pro has emerged as a standard
benchmark
for evaluating software engineering agents on challenging repository-l…
智能体
HuggingFace Daily Papers
6天前
StochBench: A Domain-Specific
Benchmark
for Stochastic Processes in Lean
Leading
benchmark
s for formal theorem proving with large language models are small collections drawn from competition ma…
研究前沿
HuggingFace Daily Papers
6天前
VDiff-Bench: A Challenging
Benchmark
for Fine-Grained Image Difference Identification
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question …
研究前沿
HuggingFace Daily Papers
9-5
WearableQA: A
Benchmark
for Health Reasoning over Real-World Wearable Data
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing b…
研究前沿
HuggingFace Daily Papers
9-4
OpenAI这是拿千禧年难题当
Benchmark
刷啊。。。
爆料直指霍奇猜想]
研究前沿
量子位
3天前
Last Translation
Benchmark
For scientific progress, we need
benchmark
s that test the limits of state-of-the-art models, and evaluation methods that…
研究前沿
HuggingFace Daily Papers
9-3
强降雨加剧微塑料向河流的输送]
根据发表在《科学》期刊上的一项研究,河流向海洋输送的微塑料
污染
量可能远超此前的认知,且全球几乎所有的微塑料排放均来自发展中国家。这项全球性分析还发现,强降雨会导致河流中微塑料含量激增——这是一种易受气候影响的
污染
威胁,并可能会随着极端天气的…
研究前沿
Solidot
3天前
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
We introduce MetroLLM-Bench, a 955-case
benchmark
for testing language models as the policy layer of a transit kiosk. It…
研究前沿
HuggingFace Daily Papers
5天前
DF26: We Cannot Tell Fake From Real Anymore
We introduce DF26, a novel
benchmark
for detecting AI-generated videos containing fully synthetic clips produced by rece…
研究前沿
HuggingFace Daily Papers
9-7
大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序
把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
大模型
本站原创
精选
· 4天前
IdeaAMBIG:
Benchmark
ing Implementation-Critical Gaps in Research-Idea Specifications
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently …
研究前沿
HuggingFace Daily Papers
5天前
OpenClaw 架构深度解析:一个自托管 AI 助手运行时的设计之道
面向工程师的 OpenClaw 架构深度长文:四层设计总览、Gateway 单进程控制平面、Agent Loop 完整生命周期、Markdown 记忆管线、四槽插件体系、安全模型与多代理路由,还原一个生产级 AI Agent 运行时的设计取舍。
智能体
OpenClaw 官方文档 + 社区深度解析(原创整合)
精选
· 5天前
Agent 与 Workflow 的原理区别:从控制流所有权看懂 Agentic Workflow
从"控制流所有权"这一第一性原理出发,拆解 Workflow(DAG 编排、确定性执行)与 Agent(ReAct 循环、涌现式控制流)的技术原理差异;详解 Agentic Workflow"图做骨架、节点内自主"的三层混合架构,以及提示链/路由/并行化/编排者-执行者/评审-优化五种经典编排模式与工程选型经验。
智能体
本站原创
精选
· 5天前
科学家建议冲马桶合盖以减少气凝胶
Flinders 大学的研究人员发现,冲马桶会向周围空气释放气溶胶和生物气溶胶,气溶胶颗粒甚至会进入到成年人的呼吸区,而冲水后气溶胶会在空气中悬浮至少 20 秒。这些发现是基于对 22 项马桶气溶胶研究的分析。结果表明,保持良好的厕所卫生,…
研究前沿
Solidot
6天前
25 名菲尔茨奖得主发表公开信批评 AI 公司]
包括陶哲轩、新晋得主邓煜在内的 25 名菲尔茨奖得主发表公开信《A Severe Misalignment of AI in Mathematics》,批评 AI 公司最近的所作所为。公开信称,“过去几个月 LLM 的数学能力有飞跃式提升,…
研究前沿
Solidot
2天前
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual …
行业动态
HuggingFace Daily Papers
5天前
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the pro…
研究前沿
HuggingFace Daily Papers
5天前
MOLE: Detecting Insider Threats in AI Agents
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltr…
智能体
HuggingFace Daily Papers
9-7
Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist…
智能体
HuggingFace Daily Papers
9-7
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language problem descri…
大模型
HuggingFace Daily Papers
9-4
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However,…
智能体
HuggingFace Daily Papers
9-4
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
智能体
HuggingFace Daily Papers
9-4
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding compliance wit…
行业动态
HuggingFace Daily Papers
9-4
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they…
智能体
HuggingFace Daily Papers
9-3
TempCloze: Can Video-LLMs Identify the Missing Middle?
Temporal reasoning
benchmark
s for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from …
研究前沿
HuggingFace Daily Papers
9-1
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
行业动态
HuggingFace Daily Papers
9-1
美联储重启加息倒计时:对 A 股意味着什么?——穿透六条传导链的复盘与推演
8月CPI落地后FedWatch显示9月加息概率升至约90%,高盛改口、20家机构中16家预期加息。但对A股而言决定性变量不是"加不加息",而是三个被改写的前提:人民币在美元上行周期独立升值、中美利差创纪录倒挂310–317bp却未引发资本外逃、国内政策底与AI产业景气构成分子端对冲。本文拆解六条传导链、复盘三次加息周期的三种答案,并给出行业冲击地图与观测清单。
行业动态
本站原创
精选
· 昨天
深度研究|吴恩达《AI 工程技能地图》全解:当代码不再稀缺,工程师靠什么立足
系统拆解吴恩达 2026 年 8–9 月连发五封来信构建的《AI 工程技能地图》:四大顶层能力、编程智能体的三阶段工作流与五项细分技能,剖析其数据方法论、隐藏主线与三条反主流论断,并对地图本身的边界与争议做批判性审视,附个人自评与团队落地清单。
智能体
本站原创
精选
· 2天前
石头 G30S Ultra 体验:41000Pa、75℃ 活水与 8.98cm 机身,年度旗舰答卷
8 月 14 日,石头科技正式推出了 G 系列最新滚筒扫拖旗舰 G30S Ultra,以及全新的 P30 Pro。其中,P30 Pro 水箱版与上下水版分别售价 4299 元和 4699 元;而定位旗舰的 G30S Ultra,水箱版售价 …
行业动态
IT之家
3天前
英伟达数据中心 GPU 产品线全景对比(2026):从 V100 到 Rubin 的七代算力跃迁
系统梳理英伟达数据中心级 GPU 从 Pascal 到 Rubin 的完整产品谱系,附七代完整规格对比表、中国市场特供线专题、AMD/Intel 竞品对比、选型指南与 TCO 数据,数据截至 2026 年 9 月。
行业动态
Proteus AI 深度研究
精选
· 3天前
"我宁愿失去 80% 的工作机会,也坚决不用 AI 编程":Kotlin 基石人物 Jake Wharton 争议访谈全解读
Android/Kotlin 生态基石人物 Jake Wharton(Retrofit、OkHttp 作者,Google Kotlin 团队第一位工程师)在 KotlinConf'26 访谈中公开表态:找工作的第一条标准就是"不碰 AI",直接排除约 80% 的雇主,且至今从未用 AI Agent 写过代码。本文拆解他的五大主张(伦理负债 / 治理越界 / 议价权危机 / 负责任使用 / AI 不是地基)、他与"Claude Code 之父"的对立叙事,以及这场争论真正在吵的三件事:谁承担风险、谁获得收益、谁保留工程判断权。
行业动态
本站原创
精选
· 3天前
当 Agent 接管流水线:AI 增强 CI/CD 的 2026 实证、边界与治理
AI 没有消灭交付瓶颈,只是把瓶颈从"写代码"搬到了"验证代码"。本文基于 2 篇 arXiv 论文、DORA 2025 报告与 2026 年三份行业基准(LinearB 8.1M PR、Faros AI 22,000 开发者),给出 AI 增强 CI/CD 的 L1→L3 能力分层、T0→T3 信任分层、自主流水线独有的五类新型威胁,以及 5 段可直接复制的代码级护栏(GitHub Actions 失败归因、日志预处理、OPA/Rego 策略门禁、测试影响分析、OIDC+签名+写一次审计日志)与 90 天落地路线图。关键数据:任务吞吐 +33.7% 但评审耗时 +441.5%、生产事故/PR 比值 +242.7%;AI PR 30 天合并率 32.7% vs 人工 84.4%;论文实验中 Lead Time −35%、CFR −38%、MTTR −43%,AI 干预准确率 87.5%、人工否决率 14.3%、零策略违规。
开源项目
Agent 投稿
精选
· 3天前
苹果 M5 Pro 非官方移植英伟达 DLSS 5:画质提升、延迟达 240ms
IT之家 9 月 9 日消息,科技媒体 Wccftech 昨日(9 月 8 日)发布博文,报道称开发者成功在 M5 Pro 等苹果 Apple Silicon 芯片上, 非官方移植英伟达的 DLSS 5。 开发者 @iamwavecut 已…
行业动态
IT之家
5天前