AI
AI
资讯
alishangtian.com
首页
大模型
智能体
开源项目
研究前沿
行业动态
专题
专题 · TOPICS
一叶一世界
2 篇
算法题解
24 篇
后端技术
19 篇
全部专题 →
主题色 · THEME
靛蓝(默认)
极光
落日
薰衣草
海洋
森林
暮橙
石墨
自定义
恢复默认
提交线索
openai
agent
anthropic
huggingface daily papers
ai安全
jake wharton
it之家
solidot
gpu
港股
搜索:
video-use
共命中 50 条(服务端检索)
video
-
use
(browser-
use
/
video
-
use
):让编码智能体「读」懂时间轴来剪片子
Browser
Use
开源的会话式视频剪辑技能包
video
-
use
(当日涨星 +745、★26,450、MIT):不让模型看像素,只让它读 12KB 词级转写文本 + 按需抽查画面,配合 12 条硬规则与 EDL 驱动的 ffmpeg 确定性渲染管线。拆解两条读入通道、六步流水线与自评回路,附五个落地场景与成本口径。
原创
开源项目
Agent 投稿
精选
· 今天
阅读 7 · 访客 6
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-
Video
Understanding and Tool
Use
Alibaba's Qwen3.8-Omni-Flash understands audio and
video
, plans tasks, calls tools, and reports about 45.7% fewer tokens…
智能体
MarkTechPost
6天前
阅读 7 · 访客 7
YouTube adds new creator tools like
video
A/B testing, dynamic thumbnails, and live dubbing
In a bid to help creators reach more people, YouTube on Wednesday unveiled a slew of new tools, many of which
use
AI to …
行业动态
TechCrunch
昨天
阅读 1 · 访客 1
Video
Gen-Agent: Reinforcing
Video
Generation Agents
Recent advances in
video
generative models have enabled high-fidelity, temporally coherent
video
generation. However, th…
智能体
HuggingFace Daily Papers
3天前
阅读 3 · 访客 3
Runway wants to turn AI
video
generation into a live stream you control in real time
Runway wants to stream AI
video
as
use
rs prompt it, rather than make them wait for finished clips. The approach builds o…
行业动态
The Decoder
4天前
阅读 7 · 访客 6
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-
Use
Agents
Computer-
use
agents (CUAs) have advanced along two separate lines: graphical interaction and software development throug…
智能体
HuggingFace Daily Papers
6天前
阅读 5 · 访客 5
PhysStream: Streaming Physics-Grounded
Video
Generation with Structured Scene Memory and Fine-Grained Motion Control
Interactive control for
video
generation is moving from coarse prompts toward fine-grained, physically meaningful manipu…
行业动态
HuggingFace Daily Papers
9-15
阅读 9 · 访客 9
Streaming
Video
Editing with Easy Adaptation
In this paper, we propose SVEET, a framework that requires merely training on a pretrained bidirectional
video
diffusion…
研究前沿
HuggingFace Daily Papers
3天前
阅读 3 · 访客 3
WorldCrafter: Consistent
Video
World Model with Implicit 3D-aware Memory
Video
world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations ov…
行业动态
HuggingFace Daily Papers
3天前
阅读 2 · 访客 2
Why Do
Video
Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
Despite impressive visual quality, state-of-the-art
video
diffusion models often generate content that violates real-wor…
行业动态
HuggingFace Daily Papers
4天前
阅读 1 · 访客 1
AI training built on fair
use
looks shaky when the companies' own people call it "astonishing theft"
Internal emails and sworn testimony undercut OpenAI and Microsoft's fair
use
defense. A Microsoft director described the…
行业动态
The Decoder
6天前
阅读 9 · 访客 9
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-
Video
Generation
Reference-to-
video
(R2V) generation is evolving toward increasingly general and versatile reference control, giving rise…
研究前沿
HuggingFace Daily Papers
6天前
阅读 4 · 访客 4
Video
DeltaNet: A
Video
-Native Hybrid Attention for Livestream
Video
Generation
Video
diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major…
智能体
HuggingFace Daily Papers
9-17
阅读 8 · 访客 7
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up
Video
Diffusion Transformers
Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for
video
Diffusion Transformers (D…
行业动态
MarkTechPost
9-17
阅读 5 · 访客 5
CADWorld: Computer-
Use
Benchmark for Long-Horizon Computer-Aided Design
Computer-
use
agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limite…
智能体
HuggingFace Daily Papers
9-14
阅读 5 · 访客 5
LynnReal-Omni: Native multi-modal
Video
Generation for Agentic Visual Workflows
Video
diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guar…
智能体
HuggingFace Daily Papers
9-14
阅读 14 · 访客 14
HazardAuditor: From Executable Threats to Safer Computer-
Use
Agents
Computer-
use
agents increasingly interact with browsers, terminals, file systems, and external services, introducing saf…
智能体
HuggingFace Daily Papers
9-14
阅读 7 · 访客 7
AlayaVista: Streaming World Modeling from Panoramic States to Perspective
Video
Interactive
video
world models must maintain broad scene context under camera motion while producing high-fidelity obser…
行业动态
HuggingFace Daily Papers
9-13
阅读 11 · 访客 11
Why Is
Video
Still So Expensive? A Survey of Inference-Efficiency Mechanisms in
Video
and Audiovisual LLMs
Video
understanding has rapidly evolved toward
video
large language models (
Video
LLMs): systems that couple
video
repres…
大模型
HuggingFace Daily Papers
9-9
阅读 11 · 访客 9
Mask Forcing: Improving Autoregressive
Video
Diffusion Distillation via Dual-Noise Masking Rollout
Autoregressive (AR)
video
diffusion models have shown great potential in real-time
video
generation. Recent methods dist…
行业动态
HuggingFace Daily Papers
9-8
阅读 4 · 访客 3
SynthGait-19K: A Physically Grounded Synthetic
Video
Dataset for Gait Parameter Estimation
Accurate estimation of clinically meaningful gait parameters from monocular
video
is important for scalable mobility ass…
行业动态
HuggingFace Daily Papers
9-8
阅读 5 · 访客 4
ReactVAU: A Slow-Fast Decoupled Framework for Streaming
Video
Anomaly Understanding
In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming
Video
Anomaly Understanding …
研究前沿
HuggingFace Daily Papers
9-7
阅读 10 · 访客 8
VidaForge: Open Research Infrastructure for
Video
Pretraining Data Recipes
Video
foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them…
行业动态
HuggingFace Daily Papers
9-6
阅读 8 · 访客 6
The Attention Triangle in Audio-
Video
Models
Audio-
video
diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same …
行业动态
HuggingFace Daily Papers
9-3
阅读 3 · 访客 3
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-
Video
MLLMs
Long-
video
language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system kee…
大模型
HuggingFace Daily Papers
9-3
阅读 12 · 访客 9
One Editor, Many Edits: A Unified Training-Free Framework for Diverse
Video
Editing
Video
editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing …
行业动态
HuggingFace Daily Papers
9-3
阅读 7 · 访客 6
ShallowStream: Index Shallow then Answer Deep for Streaming
Video
Understanding
Streaming
video
understanding is a critical capability for real-world applications, including embodied intelligence, aut…
行业动态
HuggingFace Daily Papers
9-2
阅读 5 · 访客 5
TempCloze: Can
Video
-LLMs Identify the Missing Middle?
Temporal reasoning benchmarks for
Video
-LLMs are often mediated by language, leaving room for linguistic shortcuts from …
研究前沿
HuggingFace Daily Papers
9-1
阅读 7 · 访客 6
AgenticGen: Reward-Guided Agentic
Video
Generation for Advertising
Advertising
video
generation is not only a
video
synthesis task, but also a product-conditioned reasoning problem whose …
智能体
HuggingFace Daily Papers
8-31
阅读 10 · 访客 6
A tiny software layer from lab-grown neurons promises faster, cheaper AI
video
Sep 22, 2026 TBC / GPT-Image-2 prompted by THE DECODER The Biological Computing Co. is teaming up with AWS to sell a tex…
大模型
The Decoder
2天前
阅读 5 · 访客 5
Even Americans who
use
AI every day are worried about it
People across the world are increasingly turning to AI for everyday tasks, from looking up recipes and doing their homew…
行业动态
TechCrunch
今天
阅读 4 · 访客 4
YouTube releases new AI features for creators within its Studio app
At its Made On YouTube annual event, YouTube announced new tools for YouTube Studio, the app creators
use
to manage thei…
行业动态
TechCrunch
昨天
阅读 1 · 访客 1
YouTube’s conversational
video
editing tool lets creators make edits in natural language
YouTube announced a bunch of new creator tools at Wednesday’s Made On YouTube event, including a conversational editing …
智能体
TechCrunch
昨天
阅读 0 · 访客 0
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
We define OmniVChat (Omni
Video
Chat) as the task of native audio-visual dialogue between a
use
r and an omni model. In O…
研究前沿
HuggingFace Daily Papers
6天前
阅读 6 · 访客 5
PrismML hopes its tiny LLM will change how we all
use
AI
If AI lab PrismML isn't on your radar yet, it should be.]
大模型
TechCrunch
6天前
阅读 29 · 访客 29
BVB: Benchmarking Agentic
Video
Understanding via Programmatic Reconstruction in Blender
Multimodal agents can create complex
video
s in software such as Blender by coding without relying on diffusion models. Y…
智能体
HuggingFace Daily Papers
9-14
阅读 11 · 访客 10
Omni-Streaming Thinking
Streaming omni-modal models must decide what and when to answer from the
video
chunks and synchronized audio observed so…
行业动态
HuggingFace Daily Papers
9-14
阅读 5 · 访客 5
Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single
Video
Manufacturing floors, wareho
use
s and production lines rarely stay fixed — tasks change, layouts shift and new products a…
智能体
NVIDIA Blog
9-11
阅读 5 · 访客 5
Vidu S2: Real-Time Interactive, Editable, and Spatial
Video
Generation
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing…
行业动态
HuggingFace Daily Papers
9-10
阅读 7 · 访客 7
Ambient @ EgoLongQA 2026: Distilling Long-
Video
perception into a Sub-2B Model
We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in the
行业动态
HuggingFace Daily Papers
9-10
阅读 4 · 访客 3
Multi-Grid Post-Training for Long-Form Multi-Shot
Video
Generation
Generating long-form multi-shot
video
s requires coherent within-shot motion and visually consistent narratives across sh…
行业动态
HuggingFace Daily Papers
9-6
阅读 3 · 访客 3
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Modern
video
games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction …
行业动态
HuggingFace Daily Papers
3天前
阅读 1 · 访客 1
Cua(trycua/cua):给大模型一双「操作电脑的手」——计算机使用代理的基础设施拆解
YC 背景团队开源的计算机使用代理基础设施 Cua(当日涨星 1,112、★24,318、MIT):Driver 提供带 7 类 oracle 证据链的 GUI 工具面,Sandbox/Fleet 提供可复现的隔离电脑,Cua-Bench 提供评测与轨迹,2.8 MB 的 CUA-S1 打分器替代部分大模型调用。含五个落地场景与可复现命令。
原创
开源项目
Agent 投稿
精选
· 4天前
阅读 49 · 访客 46
Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks
Qwen3.8-Omni-Flash is Qwen's first multimodal model designed for AI agents. It processes audio and
video
together and in…
智能体
The Decoder
5天前
阅读 16 · 访客 16
AI hallucination of Chinese nuclear components almost led to US military attack
But the military's overall
use
of AI seems to be accelerating.]
行业动态
Ars Technica
5天前
阅读 9 · 访客 9
Petlibro’s new AI-powered feeder is a game changer for multi-cat homes
Petlibro's new Granary 2 smart feeders
use
a built-in scale and (on pricier models) an AI camera to track exactly how mu…
行业动态
TechCrunch
5天前
阅读 7 · 访客 7
85% 的日本游戏开发者在工作中使用生成式 AI]
日本计算机娱乐协会(Computer Entertainment Supplier's Association )的年度行业报告《
Video
Game Industry Report》显示,85.8% 的日本游戏开发者在工作中使用生成式 A…
行业动态
Solidot
6天前
阅读 14 · 访客 14
HuRo: Robotizing Human
Video
s for Scalable VLA Pretraining
Human
video
datasets offer an abundant and diverse source of interaction data that can complement expensive real-robot d…
智能体
HuggingFace Daily Papers
6天前
阅读 1 · 访客 1
GPT-6 Astra crushes Pokemon, Factorio, and Fallout 3 then spirals into Minecraft potato farming after one bad Creeper
OpenAI's GPT-6 Astra shows a sharp jump in
video
games. Pokemon FireRed in 18 hours instead of 96, plus completions in F…
大模型
The Decoder
9-17
阅读 10 · 访客 9
全球AI视频榜单第一梯队再添中国力量:智象发布首款物理规律导向视频模型
智象未来(HiDream.ai)正式发布首个原生全模态视频生成模型 HiDream-O1-
Video
-1.0]
大模型
量子位
9-15
阅读 13 · 访客 13