搜索:HuggingFace Daily Papers

共命中 50 条(服务端检索)
IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently …
研究前沿 HuggingFace Daily Papers 5天前
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing b…
研究前沿 HuggingFace Daily Papers 9-4
Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs
Code world models represent worlds as executable programs, but this representation alone does not determine how to const…
行业动态 HuggingFace Daily Papers 4天前
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and…
行业动态 HuggingFace Daily Papers 4天前
Mi-Ripple: Restoring Images Degraded by Iterative AI Editing
Iterative reference-conditioned image editing can introduce grid-like and granular textures, commonly described as digit…
行业动态 HuggingFace Daily Papers 4天前
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat
Multi-Agent Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous sy…
智能体 HuggingFace Daily Papers 4天前
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks per…
大模型 HuggingFace Daily Papers 4天前
Beyond Solver Verdicts: Generative Reward Models for Autoformalization
Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally bli…
研究前沿 HuggingFace Daily Papers 4天前
Memory as Plans: World-Action Modeling with Memory-Grounded Planning
Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inhe…
智能体 HuggingFace Daily Papers 4天前
World in World: Explore the World with World Models
Autoregressive video world models enable interactive, long-horizon exploration, but flexible control remains challenging…
行业动态 HuggingFace Daily Papers 4天前
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, al…
大模型 HuggingFace Daily Papers 4天前
UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
All-in-One medical image restoration (MedIR) aims to address diverse tasks across modalities and degradation types using…
行业动态 HuggingFace Daily Papers 4天前
Generative Late-Interaction Embeddings For Visual Document Retrieval
Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. …
行业动态 HuggingFace Daily Papers 4天前
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks ar…
智能体 HuggingFace Daily Papers 4天前
SenseNova-U1.5: Towards Native Unified Visual Intelligence
We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visua…
行业动态 HuggingFace Daily Papers 4天前
Show-Harness: Just a VLM Agent Can Play Robots
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence i…
智能体 HuggingFace Daily Papers 5天前
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It…
研究前沿 HuggingFace Daily Papers 5天前
Programmable World Model
Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanis…
行业动态 HuggingFace Daily Papers 5天前
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-…
行业动态 HuggingFace Daily Papers 5天前
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain ov…
研究前沿 HuggingFace Daily Papers 5天前
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
We study how model post-training and test-time inference design affect natural-language proof generation for hard olympi…
行业动态 HuggingFace Daily Papers 5天前
The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine…
智能体 HuggingFace Daily Papers 5天前
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual …
行业动态 HuggingFace Daily Papers 5天前
Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentia…
研究前沿 HuggingFace Daily Papers 5天前
Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking
Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on…
研究前沿 HuggingFace Daily Papers 5天前
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video repres…
大模型 HuggingFace Daily Papers 5天前
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the pro…
研究前沿 HuggingFace Daily Papers 5天前
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. M…
智能体 HuggingFace Daily Papers 6天前
Omni Interaction Agent Technical Report
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic cap…
智能体 HuggingFace Daily Papers 6天前
Miles v0.1: Production-Level Post-Training
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design …
行业动态 HuggingFace Daily Papers 6天前
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and …
智能体 HuggingFace Daily Papers 6天前
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in sc…
智能体 HuggingFace Daily Papers 6天前
NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting
The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetime, yet ful…
行业动态 HuggingFace Daily Papers 6天前
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition ma…
研究前沿 HuggingFace Daily Papers 6天前
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a …
智能体 HuggingFace Daily Papers 6天前
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sideste…
研究前沿 HuggingFace Daily Papers 6天前
Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents thr…
智能体 HuggingFace Daily Papers 6天前
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated …
行业动态 HuggingFace Daily Papers 6天前
SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-…
智能体 HuggingFace Daily Papers 6天前
Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. …
研究前沿 HuggingFace Daily Papers 6天前
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how schemi…
智能体 HuggingFace Daily Papers 6天前
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple …
智能体 HuggingFace Daily Papers 6天前
ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
As LLMs are increasingly used for pre-submission self-review, there is growing demand for feedback that not only identif…
大模型 HuggingFace Daily Papers 6天前
PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used t…
智能体 HuggingFace Daily Papers 6天前
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-l…
智能体 HuggingFace Daily Papers 6天前
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods dist…
行业动态 HuggingFace Daily Papers 6天前
The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements
We consider the problem of support recovery for sparse binary signals from noisy linear measurements. For sparse Gaussia…
行业动态 HuggingFace Daily Papers 6天前
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that…
智能体 HuggingFace Daily Papers 6天前
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility ass…
行业动态 HuggingFace Daily Papers 6天前
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autono…
智能体 HuggingFace Daily Papers 6天前