Anthropic CEO 长文《We Must Pace the Frontier》:首次呼吁为前沿 AI 能力进步"限速",三步走方案详解
Dario Amodei 发文主张放慢 AI 能力提升速度,提出"嵌入式评估员—民主国家协调—全球协调"三步走方案,第一步 Anthropic 已单方面承诺执行;Altman 与马斯克罕见附和,业内同时出现"监管俘获"质疑。
阅读全文 →
行业动态
共 218 条SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in o…
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. W…
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surround…
Group Adaptive Clipping Policy Optimization
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed impo…
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes …
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
Instance segmentation of overlapping cells in microscopy remains challenging due to semi-transparent structures that pro…
To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoor scenes, …
Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in s…
One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It com…
Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing
Change data synthesis provides a cost-effective solution for expanding training data and improving the performance of ch…
Training-Free Speech-Centric Omni Understanding with Frozen VLMs
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual events, and …
Google I/O:Gemini 2.5 与全模态搜索重塑
Google 在 I/O 大会全面转向 AI First:Gemini 2.5 Pro/Flash 系列落地百亿级用户产品,AI Mode 重构搜索交互,Veo/Imagen 视频图像生成商用化。