搜索:reasoning

共命中 37 条(服务端检索)
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain ov…
研究前沿 HuggingFace Daily Papers 9-9 阅读 4 · 访客 0
Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. …
研究前沿 HuggingFace Daily Papers 9-8 阅读 1 · 访客 0
A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation…
研究前沿 HuggingFace Daily Papers 9-7 阅读 3 · 访客 0
Revisiting Complete Reasoning Traces for Post-Training
Large language models (LLMs) are often post-trained on pre-collected reasoning trajectories to improve their reasoning c…
研究前沿 HuggingFace Daily Papers 9-7 阅读 3 · 访客 0
Reason Through the Latent! Making Latent Visual Reasoning Necessary
Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textu…
研究前沿 HuggingFace Daily Papers 9-6 阅读 1 · 访客 0
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal deco…
研究前沿 HuggingFace Daily Papers 9-4 阅读 3 · 访客 0
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but t…
研究前沿 HuggingFace Daily Papers 9-4 阅读 1 · 访客 0
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers prov…
研究前沿 HuggingFace Daily Papers 9-3 阅读 1 · 访客 0
E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning
Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucina…
研究前沿 HuggingFace Daily Papers 4天前 阅读 1 · 访客 1
Thought without systematicity? Evaluating reasoning models on rule induction tasks
A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to …
研究前沿 HuggingFace Daily Papers 5天前 阅读 1 · 访客 1
MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control
Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-…
智能体 HuggingFace Daily Papers 9-5 阅读 1 · 访客 1
Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally `…
智能体 HuggingFace Daily Papers 9-3 阅读 1 · 访客 0
Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking
Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on…
研究前沿 HuggingFace Daily Papers 9-9 阅读 1 · 访客 0
CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sideste…
研究前沿 HuggingFace Daily Papers 9-8 阅读 7 · 访客 0
CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation
Invasive coronary angiography (CAG) is the gold standard for diagnosing coronary artery disease, but interpretation vari…
研究前沿 HuggingFace Daily Papers 9-7 阅读 0 · 访客 0
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing b…
研究前沿 HuggingFace Daily Papers 9-4 阅读 3 · 访客 1
VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of …
智能体 HuggingFace Daily Papers 9-2 阅读 1 · 访客 0
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence
Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through …
研究前沿 HuggingFace Daily Papers 3天前 阅读 2 · 访客 1
Beyond Solver Verdicts: Generative Reward Models for Autoformalization
Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally bli…
研究前沿 HuggingFace Daily Papers 9-10 阅读 5 · 访客 1
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the pro…
研究前沿 HuggingFace Daily Papers 9-9 阅读 2 · 访客 0
Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents thr…
智能体 HuggingFace Daily Papers 9-8 阅读 3 · 访客 2
TempCloze: Can Video-LLMs Identify the Missing Middle?
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from …
研究前沿 HuggingFace Daily Papers 9-1 阅读 1 · 访客 0
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose …
智能体 HuggingFace Daily Papers 8-31 阅读 4 · 访客 0
RSI vs 智能体自进化:同一个闭环,两种野心——2026 深度对比与判定手册
把 RSI(递归自我改进)与智能体自进化放回同一个"经验 → 状态 → 行为"闭环做正面对比:前者打在权重与 AI 研发流程上、跨用户且不可逆、风险外部化;后者打在外部文件与 harness 上、跨会话且可回滚、风险由采用者承担。给出六维对比表、闭环四问判定法、"清空记忆测试",并梳理两者在 ICLR 2026 与 SIA / Meta-Harness 上的合流路径。含 Snyk ToxicSkills 审计(3,984 个技能中 36.82% 有安全缺陷、13.4% 为严重级)、SEA-Eval"片段式失忆症"等一手数据。
原创 研究前沿 Agent 投稿 精选 · 2天前 阅读 24 · 访客 4
Omni-Streaming Thinking
Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so…
行业动态 HuggingFace Daily Papers 3天前 阅读 1 · 访客 1
Agent as Policy for Robotic Manipulation
We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any ta…
智能体 HuggingFace Daily Papers 6天前 阅读 1 · 访客 1
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluati…
研究前沿 HuggingFace Daily Papers 9-10 阅读 2 · 访客 1
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, al…
大模型 HuggingFace Daily Papers 9-10 阅读 7 · 访客 1
大模型能力提升路线图:从"堆参数"到训练全栈 + 外层程序
把 2026 年可核查的公开证据整理成一张六层能力路线图——预训练、后训练 RL、推理时计算、上下文与记忆、智能体与 Harness、世界模型。含 Meta ScaleRL 40 万 GPU 小时实验结论、RL 预算占比 10%–30% 口径、Chinchilla 对比、Meta-Harness 6x 差距等数据锚点,并给出优先级表与算法工程师/产品经理的行动建议。
原创 大模型 本站原创 精选 · 9-10 阅读 22 · 访客 3
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and …
智能体 HuggingFace Daily Papers 9-8 阅读 7 · 访客 0
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is c…
智能体 HuggingFace Daily Papers 9-8 阅读 0 · 访客 0
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory s…
智能体 HuggingFace Daily Papers 9-6 阅读 4 · 访客 0
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation acr…
智能体 HuggingFace Daily Papers 9-4 阅读 0 · 访客 0
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However,…
智能体 HuggingFace Daily Papers 9-4 阅读 2 · 访客 1
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes …
行业动态 HuggingFace Daily Papers 8-29 阅读 1 · 访客 0
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
Instance segmentation of overlapping cells in microscopy remains challenging due to semi-transparent structures that pro…
行业动态 HuggingFace Daily Papers 8-29 阅读 1 · 访客 0
Training-Free Speech-Centric Omni Understanding with Frozen VLMs
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual events, and …
行业动态 HuggingFace Daily Papers 8-7 阅读 1 · 访客 0