Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Daeun, Yoon, Jaehong, Cho, Jaemin, Bansal, Mohit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
por: Lee, Daeun, et al.
Publicado: (2025)
por: Lee, Daeun, et al.
Publicado: (2025)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
por: Li, Jialu, et al.
Publicado: (2025)
por: Li, Jialu, et al.
Publicado: (2025)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
por: Yoon, Jaehong, et al.
Publicado: (2024)
por: Yoon, Jaehong, et al.
Publicado: (2024)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
por: Wang, Zun, et al.
Publicado: (2026)
por: Wang, Zun, et al.
Publicado: (2026)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
por: Yu, Shoubin, et al.
Publicado: (2024)
por: Yu, Shoubin, et al.
Publicado: (2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
por: Wang, Zun, et al.
Publicado: (2024)
por: Wang, Zun, et al.
Publicado: (2024)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
por: Lin, Han, et al.
Publicado: (2023)
por: Lin, Han, et al.
Publicado: (2023)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
por: Wang, Zun, et al.
Publicado: (2025)
por: Wang, Zun, et al.
Publicado: (2025)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
por: Yu, Shoubin, et al.
Publicado: (2025)
por: Yu, Shoubin, et al.
Publicado: (2025)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
por: Wang, Ziyang, et al.
Publicado: (2025)
por: Wang, Ziyang, et al.
Publicado: (2025)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
por: Sung, Yi-Lin, et al.
Publicado: (2023)
por: Sung, Yi-Lin, et al.
Publicado: (2023)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
por: Wang, Ziyang, et al.
Publicado: (2024)
por: Wang, Ziyang, et al.
Publicado: (2024)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
por: Zala, Abhay, et al.
Publicado: (2023)
por: Zala, Abhay, et al.
Publicado: (2023)
Hierarchy-Aware Multimodal Unlearning for Medical AI
por: Wu, Fengli, et al.
Publicado: (2025)
por: Wu, Fengli, et al.
Publicado: (2025)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
por: Pothiraj, Atin, et al.
Publicado: (2025)
por: Pothiraj, Atin, et al.
Publicado: (2025)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
por: Niu, Tianyi, et al.
Publicado: (2025)
por: Niu, Tianyi, et al.
Publicado: (2025)
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
por: Huang, Yidong, et al.
Publicado: (2025)
por: Huang, Yidong, et al.
Publicado: (2025)
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
por: Cho, Jaemin, et al.
Publicado: (2024)
por: Cho, Jaemin, et al.
Publicado: (2024)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
por: Lin, Han, et al.
Publicado: (2025)
por: Lin, Han, et al.
Publicado: (2025)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
por: Yoon, Jaehong, et al.
Publicado: (2024)
por: Yoon, Jaehong, et al.
Publicado: (2024)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
por: Huang, Yidong, et al.
Publicado: (2026)
por: Huang, Yidong, et al.
Publicado: (2026)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
por: Wang, Ziyang, et al.
Publicado: (2026)
por: Wang, Ziyang, et al.
Publicado: (2026)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
por: Lee, Daeun, et al.
Publicado: (2025)
por: Lee, Daeun, et al.
Publicado: (2025)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
por: Wan, David, et al.
Publicado: (2024)
por: Wan, David, et al.
Publicado: (2024)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
por: Lee, Daeun, et al.
Publicado: (2026)
por: Lee, Daeun, et al.
Publicado: (2026)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
por: Yu, Shoubin, et al.
Publicado: (2026)
por: Yu, Shoubin, et al.
Publicado: (2026)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
por: Cho, Jaemin, et al.
Publicado: (2023)
por: Cho, Jaemin, et al.
Publicado: (2023)
TimeRefine: Temporal Grounding with Time Refining Video LLM
por: Wang, Xizi, et al.
Publicado: (2024)
por: Wang, Xizi, et al.
Publicado: (2024)
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
por: Cho, Jaemin, et al.
Publicado: (2023)
por: Cho, Jaemin, et al.
Publicado: (2023)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
por: Sivakumaran, Nithin, et al.
Publicado: (2025)
por: Sivakumaran, Nithin, et al.
Publicado: (2025)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
por: Wang, Zun, et al.
Publicado: (2024)
por: Wang, Zun, et al.
Publicado: (2024)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
por: Zala, Abhay, et al.
Publicado: (2024)
por: Zala, Abhay, et al.
Publicado: (2024)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
por: Lin, Han, et al.
Publicado: (2024)
por: Lin, Han, et al.
Publicado: (2024)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
por: Yu, Shoubin, et al.
Publicado: (2025)
por: Yu, Shoubin, et al.
Publicado: (2025)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
por: Han, Donghoon, et al.
Publicado: (2024)
por: Han, Donghoon, et al.
Publicado: (2024)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
por: Wang, Ziyang, et al.
Publicado: (2025)
por: Wang, Ziyang, et al.
Publicado: (2025)
Are Video Reasoning Models Ready to Go Outside?
por: He, Yangfan, et al.
Publicado: (2026)
por: He, Yangfan, et al.
Publicado: (2026)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
Ejemplares similares
-
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
por: Lee, Daeun, et al.
Publicado: (2025) -
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
por: Li, Jialu, et al.
Publicado: (2025) -
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
por: Li, Jialu, et al.
Publicado: (2024) -
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
por: Yoon, Jaehong, et al.
Publicado: (2024) -
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
por: Wang, Zun, et al.
Publicado: (2026)