Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Jialu, Yu, Shoubin, Lin, Han, Cho, Jaemin, Yoon, Jaehong, Bansal, Mohit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
di: Yu, Shoubin, et al.
Pubblicazione: (2024)
di: Yu, Shoubin, et al.
Pubblicazione: (2024)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
di: Lee, Daeun, et al.
Pubblicazione: (2024)
di: Lee, Daeun, et al.
Pubblicazione: (2024)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
di: Li, Jialu, et al.
Pubblicazione: (2024)
di: Li, Jialu, et al.
Pubblicazione: (2024)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
di: Zala, Abhay, et al.
Pubblicazione: (2024)
di: Zala, Abhay, et al.
Pubblicazione: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
di: Lee, Daeun, et al.
Pubblicazione: (2025)
di: Lee, Daeun, et al.
Pubblicazione: (2025)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
di: Lin, Han, et al.
Pubblicazione: (2023)
di: Lin, Han, et al.
Pubblicazione: (2023)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
di: Wang, Zun, et al.
Pubblicazione: (2024)
di: Wang, Zun, et al.
Pubblicazione: (2024)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
di: Wang, Zun, et al.
Pubblicazione: (2025)
di: Wang, Zun, et al.
Pubblicazione: (2025)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
di: Zala, Abhay, et al.
Pubblicazione: (2023)
di: Zala, Abhay, et al.
Pubblicazione: (2023)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
di: Wang, Ziyang, et al.
Pubblicazione: (2024)
di: Wang, Ziyang, et al.
Pubblicazione: (2024)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
di: Wan, David, et al.
Pubblicazione: (2024)
di: Wan, David, et al.
Pubblicazione: (2024)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
di: Wang, Zun, et al.
Pubblicazione: (2026)
di: Wang, Zun, et al.
Pubblicazione: (2026)
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
di: Huang, Yidong, et al.
Pubblicazione: (2025)
di: Huang, Yidong, et al.
Pubblicazione: (2025)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
di: Lin, Han, et al.
Pubblicazione: (2025)
di: Lin, Han, et al.
Pubblicazione: (2025)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
di: Sung, Yi-Lin, et al.
Pubblicazione: (2023)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2023)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2026)
di: Wang, Ziyang, et al.
Pubblicazione: (2026)
Hierarchy-Aware Multimodal Unlearning for Medical AI
di: Wu, Fengli, et al.
Pubblicazione: (2025)
di: Wu, Fengli, et al.
Pubblicazione: (2025)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
di: Wan, David, et al.
Pubblicazione: (2025)
di: Wan, David, et al.
Pubblicazione: (2025)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
di: Niu, Tianyi, et al.
Pubblicazione: (2025)
di: Niu, Tianyi, et al.
Pubblicazione: (2025)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
di: Pothiraj, Atin, et al.
Pubblicazione: (2025)
di: Pothiraj, Atin, et al.
Pubblicazione: (2025)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
di: Huang, Yidong, et al.
Pubblicazione: (2026)
di: Huang, Yidong, et al.
Pubblicazione: (2026)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
di: Li, Jialu, et al.
Pubblicazione: (2024)
di: Li, Jialu, et al.
Pubblicazione: (2024)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
di: Cho, Jaemin, et al.
Pubblicazione: (2024)
di: Cho, Jaemin, et al.
Pubblicazione: (2024)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
di: Khan, Zaid, et al.
Pubblicazione: (2025)
di: Khan, Zaid, et al.
Pubblicazione: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
di: Sivakumaran, Nithin, et al.
Pubblicazione: (2025)
di: Sivakumaran, Nithin, et al.
Pubblicazione: (2025)
Multimodal Representation Learning by Alternating Unimodal Adaptation
di: Zhang, Xiaohui, et al.
Pubblicazione: (2023)
di: Zhang, Xiaohui, et al.
Pubblicazione: (2023)
Glider: Global and Local Instruction-Driven Expert Router
di: Li, Pingzhi, et al.
Pubblicazione: (2024)
di: Li, Pingzhi, et al.
Pubblicazione: (2024)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
di: Lee, Daeun, et al.
Pubblicazione: (2026)
di: Lee, Daeun, et al.
Pubblicazione: (2026)
Balancing Faithfulness and Performance in Reasoning via Multi-Listener Soft Execution
di: Sivakumaran, Nithin, et al.
Pubblicazione: (2026)
di: Sivakumaran, Nithin, et al.
Pubblicazione: (2026)
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
di: Cho, Jaemin, et al.
Pubblicazione: (2023)
One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
di: Khan, Zaid, et al.
Pubblicazione: (2025)
di: Khan, Zaid, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
di: Yu, Shoubin, et al.
Pubblicazione: (2024) -
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
di: Yoon, Jaehong, et al.
Pubblicazione: (2024) -
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
di: Lee, Daeun, et al.
Pubblicazione: (2024) -
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
di: Li, Jialu, et al.
Pubblicazione: (2024) -
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
di: Zala, Abhay, et al.
Pubblicazione: (2024)