搜索:benchmark

共命中 22 条(服务端检索)
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-l…
智能体 HuggingFace Daily Papers 6天前
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition ma…
研究前沿 HuggingFace Daily Papers 6天前
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question …
研究前沿 HuggingFace Daily Papers 9-5
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing b…
研究前沿 HuggingFace Daily Papers 9-4
OpenAI这是拿千禧年难题当Benchmark刷啊。。。
爆料直指霍奇猜想]
研究前沿 量子位 3天前
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that…
研究前沿 HuggingFace Daily Papers 9-3
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It…
研究前沿 HuggingFace Daily Papers 5天前
DF26: We Cannot Tell Fake From Real Anymore
We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by rece…
研究前沿 HuggingFace Daily Papers 9-7
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently …
研究前沿 HuggingFace Daily Papers 5天前
25 名菲尔茨奖得主发表公开信批评 AI 公司]
包括陶哲轩、新晋得主邓煜在内的 25 名菲尔茨奖得主发表公开信《A Severe Misalignment of AI in Mathematics》,批评 AI 公司最近的所作所为。公开信称,“过去几个月 LLM 的数学能力有飞跃式提升,…
研究前沿 Solidot 2天前
大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序
把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
大模型 本站原创 精选 · 4天前
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual …
行业动态 HuggingFace Daily Papers 5天前
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the pro…
研究前沿 HuggingFace Daily Papers 5天前
MOLE: Detecting Insider Threats in AI Agents
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltr…
智能体 HuggingFace Daily Papers 9-7
Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist…
智能体 HuggingFace Daily Papers 9-7
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language problem descri…
大模型 HuggingFace Daily Papers 9-4
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However,…
智能体 HuggingFace Daily Papers 9-4
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and opera…
智能体 HuggingFace Daily Papers 9-4
Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding compliance wit…
行业动态 HuggingFace Daily Papers 9-4
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they…
智能体 HuggingFace Daily Papers 9-3
TempCloze: Can Video-LLMs Identify the Missing Middle?
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from …
研究前沿 HuggingFace Daily Papers 9-1
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
行业动态 HuggingFace Daily Papers 9-1