BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Weixi, Liu, Chao, Liu, Sifei, Wang, William Yang, Vahdat, Arash, Nie, Weili |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compositional Text-to-Image Generation with Dense Blob Representations
by: Nie, Weili, et al.
Published: (2024)
by: Nie, Weili, et al.
Published: (2024)
Transition Matching Distillation for Fast Video Generation
by: Nie, Weili, et al.
Published: (2026)
by: Nie, Weili, et al.
Published: (2026)
On Equivariance and Fast Sampling in Video Diffusion Models Trained with Warped Noise
by: Liu, Chao, et al.
Published: (2025)
by: Liu, Chao, et al.
Published: (2025)
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
by: Li, Yaowei, et al.
Published: (2025)
by: Li, Yaowei, et al.
Published: (2025)
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
by: Xu, Dejia, et al.
Published: (2024)
by: Xu, Dejia, et al.
Published: (2024)
One-step Diffusion Models with $f$-Divergence Distribution Matching
by: Xu, Yilun, et al.
Published: (2025)
by: Xu, Yilun, et al.
Published: (2025)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
by: Feng, Weixi, et al.
Published: (2024)
by: Feng, Weixi, et al.
Published: (2024)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024)
by: Tang, Yolo Y., et al.
Published: (2024)
Fast Training of Diffusion Models with Masked Transformers
by: Zheng, Hongkai, et al.
Published: (2023)
by: Zheng, Hongkai, et al.
Published: (2023)
Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models
by: Daras, Giannis, et al.
Published: (2024)
by: Daras, Giannis, et al.
Published: (2024)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026)
by: Tzachor, Issar, et al.
Published: (2026)
Object-Shot Enhanced Grounding Network for Egocentric Video
by: Feng, Yisen, et al.
Published: (2025)
by: Feng, Yisen, et al.
Published: (2025)
Do Segmentation Models Understand Vascular Structure? A Blob-Based XAI Framework
by: Garret, Guillaume, et al.
Published: (2025)
by: Garret, Guillaume, et al.
Published: (2025)
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
by: Zhang, Haoyu, et al.
Published: (2023)
by: Zhang, Haoyu, et al.
Published: (2023)
Not-So-Optimal Transport Flows for 3D Point Cloud Generation
by: Hui, Ka-Hei, et al.
Published: (2025)
by: Hui, Ka-Hei, et al.
Published: (2025)
AdaVid: Adaptive Video-Language Pretraining
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
by: Xiao, Xinyu, et al.
Published: (2026)
by: Xiao, Xinyu, et al.
Published: (2026)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
by: Liang, Baoyu, et al.
Published: (2025)
by: Liang, Baoyu, et al.
Published: (2025)
HARIVO: Harnessing Text-to-Image Models for Video Generation
by: Kwon, Mingi, et al.
Published: (2024)
by: Kwon, Mingi, et al.
Published: (2024)
Asynchronous Blob Tracker for Event Cameras
by: Wang, Ziwei, et al.
Published: (2023)
by: Wang, Ziwei, et al.
Published: (2023)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
Truncated Consistency Models
by: Lee, Sangyun, et al.
Published: (2024)
by: Lee, Sangyun, et al.
Published: (2024)
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
by: Yang, Yiming, et al.
Published: (2025)
by: Yang, Yiming, et al.
Published: (2025)
VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control
by: Jiang, Lifan, et al.
Published: (2025)
by: Jiang, Lifan, et al.
Published: (2025)
SafeVid: Toward Safety Aligned Video Large Multimodal Models
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
Video Text Preservation with Synthetic Text-Rich Videos
by: Liu, Ziyang, et al.
Published: (2025)
by: Liu, Ziyang, et al.
Published: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
by: Fang, Ye, et al.
Published: (2025)
by: Fang, Ye, et al.
Published: (2025)
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
by: Lin, Rui, et al.
Published: (2026)
by: Lin, Rui, et al.
Published: (2026)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing
by: Chowdhury, Rohit, et al.
Published: (2025)
by: Chowdhury, Rohit, et al.
Published: (2025)
EdgeVidSum: Real-Time Personalized Video Summarization at the Edge
by: Mujtaba, Ghulam, et al.
Published: (2025)
by: Mujtaba, Ghulam, et al.
Published: (2025)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
by: Poppi, Tobia, et al.
Published: (2026)
by: Poppi, Tobia, et al.
Published: (2026)
VidDoS: Universal Denial-of-Service Attack on Video-based Large Language Models
by: Tang, Duoxun, et al.
Published: (2026)
by: Tang, Duoxun, et al.
Published: (2026)
Elucidating Optimal Reward-Diversity Tradeoffs in Text-to-Image Diffusion Models
by: Jena, Rohit, et al.
Published: (2024)
by: Jena, Rohit, et al.
Published: (2024)
Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
by: Chen, Haijier, et al.
Published: (2025)
by: Chen, Haijier, et al.
Published: (2025)
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
by: He, Zhihao, et al.
Published: (2026)
by: He, Zhihao, et al.
Published: (2026)
DiffiT: Diffusion Vision Transformers for Image Generation
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
by: Cheng, Sijie, et al.
Published: (2024)
by: Cheng, Sijie, et al.
Published: (2024)
Reproducibility Study of "ITI-GEN: Inclusive Text-to-Image Generation"
by: Fernández, Daniel Gallo, et al.
Published: (2024)
by: Fernández, Daniel Gallo, et al.
Published: (2024)
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
by: Vandersanden, Jente, et al.
Published: (2026)
by: Vandersanden, Jente, et al.
Published: (2026)
Similar Items
-
Compositional Text-to-Image Generation with Dense Blob Representations
by: Nie, Weili, et al.
Published: (2024) -
Transition Matching Distillation for Fast Video Generation
by: Nie, Weili, et al.
Published: (2026) -
On Equivariance and Fast Sampling in Video Diffusion Models Trained with Warped Noise
by: Liu, Chao, et al.
Published: (2025) -
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
by: Li, Yaowei, et al.
Published: (2025) -
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
by: Xu, Dejia, et al.
Published: (2024)