搜索:video-use

共命中 50 条(服务端检索)
video-use(browser-use/video-use):让编码智能体「读」懂时间轴来剪片子
Browser Use 开源的会话式视频剪辑技能包 video-use(当日涨星 +745、★26,450、MIT):不让模型看像素,只让它读 12KB 词级转写文本 + 按需抽查画面,配合 12 条硬规则与 EDL 驱动的 ffmpeg 确定性渲染管线。拆解两条读入通道、六步流水线与自评回路,附五个落地场景与成本口径。
原创 开源项目 Agent 投稿 精选 · 今天 阅读 7 · 访客 6
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
Alibaba's Qwen3.8-Omni-Flash understands audio and video, plans tasks, calls tools, and reports about 45.7% fewer tokens…
智能体 MarkTechPost 6天前 阅读 7 · 访客 7
YouTube adds new creator tools like video A/B testing, dynamic thumbnails, and live dubbing
In a bid to help creators reach more people, YouTube on Wednesday unveiled a slew of new tools, many of which use AI to …
行业动态 TechCrunch 昨天 阅读 1 · 访客 1
VideoGen-Agent: Reinforcing Video Generation Agents
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, th…
智能体 HuggingFace Daily Papers 3天前 阅读 3 · 访客 3
Runway wants to turn AI video generation into a live stream you control in real time
Runway wants to stream AI video as users prompt it, rather than make them wait for finished clips. The approach builds o…
行业动态 The Decoder 4天前 阅读 7 · 访客 6
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development throug…
智能体 HuggingFace Daily Papers 6天前 阅读 5 · 访客 5
PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipu…
行业动态 HuggingFace Daily Papers 9-15 阅读 9 · 访客 9
Streaming Video Editing with Easy Adaptation
In this paper, we propose SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion…
研究前沿 HuggingFace Daily Papers 3天前 阅读 3 · 访客 3
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations ov…
行业动态 HuggingFace Daily Papers 3天前 阅读 2 · 访客 2
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-wor…
行业动态 HuggingFace Daily Papers 4天前 阅读 1 · 访客 1
AI training built on fair use looks shaky when the companies' own people call it "astonishing theft"
Internal emails and sworn testimony undercut OpenAI and Microsoft's fair use defense. A Microsoft director described the…
行业动态 The Decoder 6天前 阅读 9 · 访客 9
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise…
研究前沿 HuggingFace Daily Papers 6天前 阅读 4 · 访客 4
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major…
智能体 HuggingFace Daily Papers 9-17 阅读 8 · 访客 7
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (D…
行业动态 MarkTechPost 9-17 阅读 5 · 访客 5
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limite…
智能体 HuggingFace Daily Papers 9-14 阅读 5 · 访客 5
LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guar…
智能体 HuggingFace Daily Papers 9-14 阅读 14 · 访客 14
HazardAuditor: From Executable Threats to Safer Computer-Use Agents
Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing saf…
智能体 HuggingFace Daily Papers 9-14 阅读 7 · 访客 7
AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity obser…
行业动态 HuggingFace Daily Papers 9-13 阅读 11 · 访客 11
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video repres…
大模型 HuggingFace Daily Papers 9-9 阅读 11 · 访客 9
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods dist…
行业动态 HuggingFace Daily Papers 9-8 阅读 4 · 访客 3
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility ass…
行业动态 HuggingFace Daily Papers 9-8 阅读 5 · 访客 4
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding …
研究前沿 HuggingFace Daily Papers 9-7 阅读 10 · 访客 8
VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes
Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them…
行业动态 HuggingFace Daily Papers 9-6 阅读 8 · 访客 6
The Attention Triangle in Audio-Video Models
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same …
行业动态 HuggingFace Daily Papers 9-3 阅读 3 · 访客 3
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system kee…
大模型 HuggingFace Daily Papers 9-3 阅读 12 · 访客 9
One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing …
行业动态 HuggingFace Daily Papers 9-3 阅读 7 · 访客 6
ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, aut…
行业动态 HuggingFace Daily Papers 9-2 阅读 5 · 访客 5
TempCloze: Can Video-LLMs Identify the Missing Middle?
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from …
研究前沿 HuggingFace Daily Papers 9-1 阅读 7 · 访客 6
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose …
智能体 HuggingFace Daily Papers 8-31 阅读 10 · 访客 6
A tiny software layer from lab-grown neurons promises faster, cheaper AI video
Sep 22, 2026 TBC / GPT-Image-2 prompted by THE DECODER The Biological Computing Co. is teaming up with AWS to sell a tex…
大模型 The Decoder 2天前 阅读 5 · 访客 5
Even Americans who use AI every day are worried about it
People across the world are increasingly turning to AI for everyday tasks, from looking up recipes and doing their homew…
行业动态 TechCrunch 今天 阅读 4 · 访客 4
YouTube releases new AI features for creators within its Studio app
At its Made On YouTube annual event, YouTube announced new tools for YouTube Studio, the app creators use to manage thei…
行业动态 TechCrunch 昨天 阅读 1 · 访客 1
YouTube’s conversational video editing tool lets creators make edits in natural language
YouTube announced a bunch of new creator tools at Wednesday’s Made On YouTube event, including a conversational editing …
智能体 TechCrunch 昨天 阅读 0 · 访客 0
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In O…
研究前沿 HuggingFace Daily Papers 6天前 阅读 6 · 访客 5
PrismML hopes its tiny LLM will change how we all use AI
If AI lab PrismML isn't on your radar yet, it should be.]
大模型 TechCrunch 6天前 阅读 29 · 访客 29
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Y…
智能体 HuggingFace Daily Papers 9-14 阅读 11 · 访客 10
Omni-Streaming Thinking
Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so…
行业动态 HuggingFace Daily Papers 9-14 阅读 5 · 访客 5
Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video
Manufacturing floors, warehouses and production lines rarely stay fixed — tasks change, layouts shift and new products a…
智能体 NVIDIA Blog 9-11 阅读 5 · 访客 5
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing…
行业动态 HuggingFace Daily Papers 9-10 阅读 7 · 访客 7
Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model
We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in the
行业动态 HuggingFace Daily Papers 9-10 阅读 4 · 访客 3
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across sh…
行业动态 HuggingFace Daily Papers 9-6 阅读 3 · 访客 3
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction …
行业动态 HuggingFace Daily Papers 3天前 阅读 1 · 访客 1
Cua(trycua/cua):给大模型一双「操作电脑的手」——计算机使用代理的基础设施拆解
YC 背景团队开源的计算机使用代理基础设施 Cua(当日涨星 1,112、★24,318、MIT):Driver 提供带 7 类 oracle 证据链的 GUI 工具面,Sandbox/Fleet 提供可复现的隔离电脑,Cua-Bench 提供评测与轨迹,2.8 MB 的 CUA-S1 打分器替代部分大模型调用。含五个落地场景与可复现命令。
原创 开源项目 Agent 投稿 精选 · 4天前 阅读 49 · 访客 46
Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks
Qwen3.8-Omni-Flash is Qwen's first multimodal model designed for AI agents. It processes audio and video together and in…
智能体 The Decoder 5天前 阅读 16 · 访客 16
AI hallucination of Chinese nuclear components almost led to US military attack
But the military's overall use of AI seems to be accelerating.]
行业动态 Ars Technica 5天前 阅读 9 · 访客 9
Petlibro’s new AI-powered feeder is a game changer for multi-cat homes
Petlibro's new Granary 2 smart feeders use a built-in scale and (on pricier models) an AI camera to track exactly how mu…
行业动态 TechCrunch 5天前 阅读 7 · 访客 7
85% 的日本游戏开发者在工作中使用生成式 AI]
日本计算机娱乐协会(Computer Entertainment Supplier's Association )的年度行业报告《Video Game Industry Report》显示,85.8% 的日本游戏开发者在工作中使用生成式 A…
行业动态 Solidot 6天前 阅读 14 · 访客 14
HuRo: Robotizing Human Videos for Scalable VLA Pretraining
Human video datasets offer an abundant and diverse source of interaction data that can complement expensive real-robot d…
智能体 HuggingFace Daily Papers 6天前 阅读 1 · 访客 1
GPT-6 Astra crushes Pokemon, Factorio, and Fallout 3 then spirals into Minecraft potato farming after one bad Creeper
OpenAI's GPT-6 Astra shows a sharp jump in video games. Pokemon FireRed in 18 hours instead of 96, plus completions in F…
大模型 The Decoder 9-17 阅读 10 · 访客 9
全球AI视频榜单第一梯队再添中国力量:智象发布首款物理规律导向视频模型
智象未来(HiDream.ai)正式发布首个原生全模态视频生成模型 HiDream-O1-Video-1.0]
大模型 量子位 9-15 阅读 13 · 访客 13