Training-Free Efficient Video Generation via Dynamic Token Carving
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yuechen, Xing, Jinbo, Xia, Bin, Liu, Shaoteng, Peng, Bohao, Tao, Xin, Wan, Pengfei, Lo, Eric, Jia, Jiaya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers
by: Zhang, Yuechen, et al.
Published: (2025)
by: Zhang, Yuechen, et al.
Published: (2025)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
by: Huang, Jiehui, et al.
Published: (2025)
by: Huang, Jiehui, et al.
Published: (2025)
Prompt Highlighter: Interactive Control for Multi-Modal LLMs
by: Zhang, Yuechen, et al.
Published: (2023)
by: Zhang, Yuechen, et al.
Published: (2023)
DreamVE: Unified Instruction-based Image and Video Editing
by: Xia, Bin, et al.
Published: (2025)
by: Xia, Bin, et al.
Published: (2025)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
by: Li, Yanwei, et al.
Published: (2024)
by: Li, Yanwei, et al.
Published: (2024)
GRADEO: Towards Human-Like Evaluation for Text-to-Video Generation via Multi-Step Reasoning
by: Mou, Zhun, et al.
Published: (2025)
by: Mou, Zhun, et al.
Published: (2025)
Scalable Language Model with Generalized Continual Learning
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
by: Qu, Tianyuan, et al.
Published: (2025)
by: Qu, Tianyuan, et al.
Published: (2025)
DreamOmni: Unified Image Generation and Editing
by: Xia, Bin, et al.
Published: (2024)
by: Xia, Bin, et al.
Published: (2024)
TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
by: Tan, Xudong, et al.
Published: (2025)
by: Tan, Xudong, et al.
Published: (2025)
DreamOmni2: Multimodal Instruction-based Editing and Generation
by: Xia, Bin, et al.
Published: (2025)
by: Xia, Bin, et al.
Published: (2025)
LLMGA: Multimodal Large Language Model based Generation Assistant
by: Xia, Bin, et al.
Published: (2023)
by: Xia, Bin, et al.
Published: (2023)
Generative Video Propagation
by: Liu, Shaoteng, et al.
Published: (2024)
by: Liu, Shaoteng, et al.
Published: (2024)
LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
by: Yang, Senqiao, et al.
Published: (2023)
by: Yang, Senqiao, et al.
Published: (2023)
VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
by: Liu, Yuqi, et al.
Published: (2025)
by: Liu, Yuqi, et al.
Published: (2025)
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
by: Song, Baiyang, et al.
Published: (2026)
by: Song, Baiyang, et al.
Published: (2026)
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
DreamOmni3: Scribble-based Editing and Generation
by: Xia, Bin, et al.
Published: (2025)
by: Xia, Bin, et al.
Published: (2025)
ConceptCarve: Dynamic Realization of Evidence
by: Caplan, Eylon, et al.
Published: (2025)
by: Caplan, Eylon, et al.
Published: (2025)
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
by: Zhong, Tianxiong, et al.
Published: (2025)
by: Zhong, Tianxiong, et al.
Published: (2025)
A General and Efficient Training for Transformer via Token Expansion
by: Huang, Wenxuan, et al.
Published: (2024)
by: Huang, Wenxuan, et al.
Published: (2024)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
by: Li, Haodong, et al.
Published: (2026)
by: Li, Haodong, et al.
Published: (2026)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
by: Wang, Chengyao, et al.
Published: (2025)
by: Wang, Chengyao, et al.
Published: (2025)
OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding
by: Wang, Chengyao, et al.
Published: (2024)
by: Wang, Chengyao, et al.
Published: (2024)
MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens
by: Kim, Youngrae, et al.
Published: (2026)
by: Kim, Youngrae, et al.
Published: (2026)
Enhancing Micro Gesture Recognition for Emotion Understanding via Context-aware Visual-Text Contrastive Learning
by: Li, Deng, et al.
Published: (2024)
by: Li, Deng, et al.
Published: (2024)
RL-GPT: Integrating Reinforcement Learning and Code-as-policy
by: Liu, Shaoteng, et al.
Published: (2024)
by: Liu, Shaoteng, et al.
Published: (2024)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
by: Yang, Senqiao, et al.
Published: (2025)
by: Yang, Senqiao, et al.
Published: (2025)
MoTCoder: Elevating Large Language Models with Modular of Thought for Challenging Programming Tasks
by: Li, Jingyao, et al.
Published: (2023)
by: Li, Jingyao, et al.
Published: (2023)
TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Acceleration
by: Zhu, Haowei, et al.
Published: (2026)
by: Zhu, Haowei, et al.
Published: (2026)
TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
by: Hu, Junshan, et al.
Published: (2025)
by: Hu, Junshan, et al.
Published: (2025)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
by: Zhong, Zhisheng, et al.
Published: (2024)
by: Zhong, Zhisheng, et al.
Published: (2024)
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
by: Wang, Ruotong, et al.
Published: (2025)
by: Wang, Ruotong, et al.
Published: (2025)
StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation
by: Jiao, Guanlong, et al.
Published: (2026)
by: Jiao, Guanlong, et al.
Published: (2026)
SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation
by: Liu, YaoYang, et al.
Published: (2026)
by: Liu, YaoYang, et al.
Published: (2026)
Training-free and Adaptive Sparse Attention for Efficient Long Video Generation
by: Xia, Yifei, et al.
Published: (2025)
by: Xia, Yifei, et al.
Published: (2025)
Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing
by: Goel, Raghavv, et al.
Published: (2026)
by: Goel, Raghavv, et al.
Published: (2026)
Similar Items
-
MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers
by: Zhang, Yuechen, et al.
Published: (2025) -
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
by: Peng, Bohao, et al.
Published: (2024) -
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
by: Huang, Jiehui, et al.
Published: (2025) -
Prompt Highlighter: Interactive Control for Multi-Modal LLMs
by: Zhang, Yuechen, et al.
Published: (2023) -
DreamVE: Unified Instruction-based Image and Video Editing
by: Xia, Bin, et al.
Published: (2025)