The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Aoxiong, Shen, Kai, Leng, Yichong, Tan, Xu, Zhou, Xinyu, Li, Juncheng, Tang, Siliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
von: Yin, Aoxiong, et al.
Veröffentlicht: (2024)
von: Yin, Aoxiong, et al.
Veröffentlicht: (2024)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
von: Shen, Kai, et al.
Veröffentlicht: (2024)
von: Shen, Kai, et al.
Veröffentlicht: (2024)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
von: Wang, Chenglin, et al.
Veröffentlicht: (2026)
von: Wang, Chenglin, et al.
Veröffentlicht: (2026)
Deciphering Oracle Bone Language with Diffusion Models
von: Guan, Haisu, et al.
Veröffentlicht: (2024)
von: Guan, Haisu, et al.
Veröffentlicht: (2024)
Pandora: Towards General World Model with Natural Language Actions and Video States
von: Xiang, Jiannan, et al.
Veröffentlicht: (2024)
von: Xiang, Jiannan, et al.
Veröffentlicht: (2024)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
von: Yang, Mingyue, et al.
Veröffentlicht: (2025)
von: Yang, Mingyue, et al.
Veröffentlicht: (2025)
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2023)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2023)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
Video Understanding with Large Language Models: A Survey
von: Tang, Yolo Y., et al.
Veröffentlicht: (2023)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2023)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
von: Qian, Long, et al.
Veröffentlicht: (2024)
von: Qian, Long, et al.
Veröffentlicht: (2024)
Open World Scene Graph Generation using Vision Language Models
von: Dutta, Amartya, et al.
Veröffentlicht: (2025)
von: Dutta, Amartya, et al.
Veröffentlicht: (2025)
Task-Aware Resolution Optimization for Visual Large Language Models
von: Luo, Weiqing, et al.
Veröffentlicht: (2025)
von: Luo, Weiqing, et al.
Veröffentlicht: (2025)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
Enhancing Large Vision Language Models with Self-Training on Image Comprehension
von: Deng, Yihe, et al.
Veröffentlicht: (2024)
von: Deng, Yihe, et al.
Veröffentlicht: (2024)
Scaling Concept With Text-Guided Diffusion Models
von: Huang, Chao, et al.
Veröffentlicht: (2024)
von: Huang, Chao, et al.
Veröffentlicht: (2024)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
Survey of Video Diffusion Models: Foundations, Implementations, and Applications
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
Fast Thinking for Large Language Models
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
von: Liu, Xinxin, et al.
Veröffentlicht: (2026)
von: Liu, Xinxin, et al.
Veröffentlicht: (2026)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
von: Wang, Peng, et al.
Veröffentlicht: (2024)
von: Wang, Peng, et al.
Veröffentlicht: (2024)
Visual In-Context Learning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
von: Gao, Qiyue, et al.
Veröffentlicht: (2025)
von: Gao, Qiyue, et al.
Veröffentlicht: (2025)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
Diffusion-RPO: Aligning Diffusion Models through Relative Preference Optimization
von: Gu, Yi, et al.
Veröffentlicht: (2024)
von: Gu, Yi, et al.
Veröffentlicht: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
VideoNorms: Benchmarking Cultural Awareness of Video Language Models
von: Varimalla, Nikhil Reddy, et al.
Veröffentlicht: (2025)
von: Varimalla, Nikhil Reddy, et al.
Veröffentlicht: (2025)
Efficient Tuning and Inference for Large Language Models on Textual Graphs
von: Zhu, Yun, et al.
Veröffentlicht: (2024)
von: Zhu, Yun, et al.
Veröffentlicht: (2024)
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
von: You, Zebin, et al.
Veröffentlicht: (2025)
von: You, Zebin, et al.
Veröffentlicht: (2025)
Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
von: Huang, Haoyang, et al.
Veröffentlicht: (2025)
von: Huang, Haoyang, et al.
Veröffentlicht: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
Sentence-Level or Token-Level? A Comprehensive Study on Knowledge Distillation
von: Wei, Jingxuan, et al.
Veröffentlicht: (2024)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2024)
MoonCast: High-Quality Zero-Shot Podcast Generation
von: Ju, Zeqian, et al.
Veröffentlicht: (2025)
von: Ju, Zeqian, et al.
Veröffentlicht: (2025)
The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
von: Yin, Aoxiong, et al.
Veröffentlicht: (2024) -
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
von: Ye, Zhen, et al.
Veröffentlicht: (2026) -
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
von: Shen, Kai, et al.
Veröffentlicht: (2024) -
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024) -
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
von: Wang, Chenglin, et al.
Veröffentlicht: (2026)