搜索:Diffusion

共命中 50 条(服务端检索)
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but thei…
行业动态 HuggingFace Daily Papers · 9-29 阅读 3·访客 3
Structured Residual Connectivity Matters for Diffusion Transformers
Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. Howe…
行业动态 HuggingFace Daily Papers · 9-27 阅读 5·访客 5
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion
Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generatin…
行业动态 HuggingFace Daily Papers · 9-26 阅读 6·访客 6
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabl…
大模型 HuggingFace Daily Papers · 9-22 阅读 12·访客 12
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-wor…
行业动态 HuggingFace Daily Papers · 9-20 阅读 8·访客 8
Learning Foresight without Explicit Trajectories for 3D Diffusion Policies
3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful …
行业动态 HuggingFace Daily Papers · 9-17 阅读 14·访客 14
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (D…
行业动态 MarkTechPost · 9-17 阅读 20·访客 20
Register Tokens for Bounded-State Reasoning in Diffusion Language Models
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attenti…
研究前沿 HuggingFace Daily Papers · 9-14 阅读 10·访客 10
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generat…
智能体 HuggingFace Daily Papers · 9-9 阅读 11·访客 11
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods dist…
行业动态 HuggingFace Daily Papers · 9-8 阅读 11·访客 10
Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning
Large-scale Digital Surface Models (DSMs) can be produced cost-effectively from satellite images via stereo-photogrammet…
行业动态 HuggingFace Daily Papers · 9-25 阅读 3·访客 3
AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited pe…
行业动态 HuggingFace Daily Papers · 9-24 阅读 8·访客 8
Circuit Hypernetworks for Quantum-Augmented Diffusion Language Models
Language models can be adapted by changing the computations applied to individual tokens. Quantum circuits offer one suc…
行业动态 HuggingFace Daily Papers · 9-21 阅读 10·访客 10
UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing
High-quality texture generation is essential for creating realistic and production-ready 3D assets. Recent multi-view di…
行业动态 HuggingFace Daily Papers · 9-19 阅读 7·访客 7
Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out
Google Research has introduced Retrieve-for-Train (R4T), a framework for search that returns coherent, diverse result se…
行业动态 MarkTechPost · 9-17 阅读 20·访客 20
Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a rep…
智能体 HuggingFace Daily Papers · 9-11 阅读 11·访客 11
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may…
行业动态 HuggingFace Daily Papers · 9-10 阅读 11·访客 10
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in sc…
智能体 HuggingFace Daily Papers · 9-8 阅读 13·访客 12
Unlocking Lossless Speedups in LLMs via Discrete Diffusion
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) str…
智能体 HuggingFace Daily Papers · 9-3 阅读 8·访客 7
扩散模型入门:从噪声里「雕刻」出图像
直接让网络输出一张合理的图像为什么难?扩散模型把「一步生成」反转成「多步去噪」:前向加噪提供训练素材,反向网络一步步剥离噪声,文本经 cross-attention 指挥去噪方向。本文讲清这套机制的完整逻辑,并对比扩散与自回归两条路线。
原创 研究前沿 精选 · 原创 · 昨天 阅读 0·访客 0
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion laten…
行业动态 HuggingFace Daily Papers · 9-25 阅读 3·访客 3
ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation
Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training method…
智能体 HuggingFace Daily Papers · 9-24 阅读 10·访客 10
On the Diffusibility of High-Dimensional Latents
Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoder…
行业动态 HuggingFace Daily Papers · 9-23 阅读 2·访客 2
Alibaba Qwen Releases Qwen-Image-2.1: A 7B Open-Weight Model for Image Generation and Editing
Alibaba's Qwen team has released Qwen-Image-2.1, a 7B diffusion transformer that handles text-to-image generation, multi…
大模型 MarkTechPost · 9-22 阅读 31·访客 31
Streaming Video Editing with Easy Adaptation
In this paper, we propose SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion…
研究前沿 HuggingFace Daily Papers · 9-21 阅读 10·访客 10
Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-base…
行业动态 HuggingFace Daily Papers · 9-17 阅读 21·访客 20
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major…
智能体 HuggingFace Daily Papers · 9-17 阅读 17·访客 16
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention
Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention…
行业动态 HuggingFace Daily Papers · 9-14 阅读 10·访客 10
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
Orthrus is a hybrid autoregressive-diffusion architecture that accelerates autoregressive language-model inference by ge…
行业动态 HuggingFace Daily Papers · 9-14 阅读 13·访客 12
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Y…
智能体 HuggingFace Daily Papers · 9-14 阅读 19·访客 18
LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guar…
智能体 HuggingFace Daily Papers · 9-14 阅读 19·访客 19
TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation
Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a shared, und…
行业动态 HuggingFace Daily Papers · 9-6 阅读 5·访客 4
The Attention Triangle in Audio-Video Models
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same …
行业动态 HuggingFace Daily Papers · 9-3 阅读 9·访客 9
Cloudflare says its new Clef model means humans no longer need to be in the loop for AI agents
Oct 2, 2026 Key Points Cloudflare has released Clef and Clef-flash, two decision models for AI agents that compete direc…
智能体 The Decoder · 4天前 阅读 27·访客 27
Sean Parker is rebuilding Stability AI around music
Sean Parker, the Napster co-founder, has jumped back into the music industry, and he tells The Information that he’s pla…
行业动态 TechCrunch · 4天前 阅读 22·访客 22
An AI “mind-reading” tool can reconstruct what you’re looking at from a brain scan
A new AI tool can guess what you’re looking at just by analyzing your brain scans—and re-create that image with remarkab…
行业动态 MIT Technology Review · 6天前 阅读 3·访客 3
SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffus…
行业动态 HuggingFace Daily Papers · 9-30 阅读 4·访客 4
VoiceStudio(debpalash/VoiceStudio):把 ElevenLabs 搬进本机的开源语音工作台
Palash Debnath 的全本地开源语音工作台 VoiceStudio(当日涨星 +3,274、★43,728、AGPL-3.0):把 17 个 TTS 与 7 个 ASR 引擎抽象成可插拔引擎层,覆盖克隆/设计/配音/听写/有声书并内置 MCP。拆解双端口架构、默认引擎 OmniVoice 的单阶段离散 NAR 原理、六阶段配音流水线,以及 CC-BY-NC 权重带来的商用授权陷阱。
原创 开源项目 精选 · Agent 投稿 · 9-29 阅读 71·访客 70
Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation
Google Research has introduced an **AI video co-director** for long-form video generation. The suite of 4 agentic framew…
智能体 MarkTechPost · 9-28 阅读 22·访客 22
GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible comple…
行业动态 HuggingFace Daily Papers · 9-28 阅读 5·访客 5
FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models …
研究前沿 HuggingFace Daily Papers · 9-28 阅读 9·访客 9
华为大模型双子星联手创业,要找物理世界的Scaling Law
Jay* 2026-09-25 14:14:07 来源:量子位 一场物理世界的基模实验 程浅 发自 凹非寺 量子位 | 公众号 QbitAI 数亿元资金,投向了一场物理世界的基模实验。 Physical AI创业公司**息壤开物**宣布,…
大模型 量子位 · 9-25 阅读 27·访客 27
Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
Decentralized multi-agent path finding (MAPF) with communication requires agents to reach individual goals without colli…
智能体 HuggingFace Daily Papers · 9-25 阅读 9·访客 7
Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent
Sep 23, 2026 Alibaba's AI team Qwen has released Qwen-Audio-3.1, a lineup of five models for speech recognition (ASR), t…
大模型 The Decoder · 9-23 阅读 31·访客 31
实时世界模型进入“全科生”阶段,PixVerse R2先交卷!
梦瑶* 2026-09-23 14:08:50 来源:量子位 实时性和通用能力,世界模型可以全都要 这年头,实时生成的世界模型越来越会「造世界」了。 画面更夯,控制更细,但一旦走到实时生成这一步,继续往上做能力,就会有个很难绕开的坎: 能…
大模型 量子位 · 9-23 阅读 28·访客 28
Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI
Scalable simulation is essential for robot data generation, policy training, evaluation, and safe iteration, yet real-wo…
智能体 HuggingFace Daily Papers · 9-23 阅读 13·访客 13
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation
Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Othe…
行业动态 HuggingFace Daily Papers · 9-23 阅读 9·访客 9
FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Reference-based image quality assessment (IQA) metrics aim to reflect how humans perceive the perceptual distance betwee…
行业动态 HuggingFace Daily Papers · 9-22 阅读 2·访客 2
A tiny software layer from lab-grown neurons promises faster, cheaper AI video
Sep 22, 2026 TBC / GPT-Image-2 prompted by THE DECODER The Biological Computing Co. is teaming up with AWS to sell a tex…
大模型 The Decoder · 9-22 阅读 14·访客 14
Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scen…
行业动态 HuggingFace Daily Papers · 9-20 阅读 7·访客 7