RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
Fuente:
arXiv
Salvato in:
| Autori principali: | Yoon, Jaehong, Yu, Shoubin, Bansal, Mohit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
di: Yu, Shoubin, et al.
Pubblicazione: (2024)
di: Yu, Shoubin, et al.
Pubblicazione: (2024)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
di: Li, Jialu, et al.
Pubblicazione: (2025)
di: Li, Jialu, et al.
Pubblicazione: (2025)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
di: Yu, Shoubin, et al.
Pubblicazione: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
di: Wang, Ziyang, et al.
Pubblicazione: (2024)
di: Wang, Ziyang, et al.
Pubblicazione: (2024)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
di: Lee, Daeun, et al.
Pubblicazione: (2024)
di: Lee, Daeun, et al.
Pubblicazione: (2024)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
di: Lee, Daeun, et al.
Pubblicazione: (2025)
di: Lee, Daeun, et al.
Pubblicazione: (2025)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
di: Yoon, Jaehong, et al.
Pubblicazione: (2024)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2026)
di: Wang, Ziyang, et al.
Pubblicazione: (2026)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
di: Wang, Zun, et al.
Pubblicazione: (2024)
di: Wang, Zun, et al.
Pubblicazione: (2024)
SELMA: Learning and Merging Skill-Specific Text-to-Image Experts with Auto-Generated Data
di: Li, Jialu, et al.
Pubblicazione: (2024)
di: Li, Jialu, et al.
Pubblicazione: (2024)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
di: Sung, Yi-Lin, et al.
Pubblicazione: (2023)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2023)
Hierarchy-Aware Multimodal Unlearning for Medical AI
di: Wu, Fengli, et al.
Pubblicazione: (2025)
di: Wu, Fengli, et al.
Pubblicazione: (2025)
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
di: Huang, Yidong, et al.
Pubblicazione: (2025)
di: Huang, Yidong, et al.
Pubblicazione: (2025)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
di: Lee, Daeun, et al.
Pubblicazione: (2026)
di: Lee, Daeun, et al.
Pubblicazione: (2026)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
di: Yu, Shoubin, et al.
Pubblicazione: (2026)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
di: Wang, Zun, et al.
Pubblicazione: (2026)
di: Wang, Zun, et al.
Pubblicazione: (2026)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
di: Sivakumaran, Nithin, et al.
Pubblicazione: (2025)
di: Sivakumaran, Nithin, et al.
Pubblicazione: (2025)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
di: Li, Jialu, et al.
Pubblicazione: (2024)
di: Li, Jialu, et al.
Pubblicazione: (2024)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
di: Wang, Zun, et al.
Pubblicazione: (2025)
di: Wang, Zun, et al.
Pubblicazione: (2025)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
di: Lin, Han, et al.
Pubblicazione: (2023)
di: Lin, Han, et al.
Pubblicazione: (2023)
STAR: A Benchmark for Situated Reasoning in Real-World Videos
di: Wu, Bo, et al.
Pubblicazione: (2024)
di: Wu, Bo, et al.
Pubblicazione: (2024)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
di: Wang, Zun, et al.
Pubblicazione: (2024)
di: Wang, Zun, et al.
Pubblicazione: (2024)
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
di: Huang, Yidong, et al.
Pubblicazione: (2026)
di: Huang, Yidong, et al.
Pubblicazione: (2026)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
di: Deng, Andong, et al.
Pubblicazione: (2025)
di: Deng, Andong, et al.
Pubblicazione: (2025)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
di: Lee, Daeun, et al.
Pubblicazione: (2025)
di: Lee, Daeun, et al.
Pubblicazione: (2025)
Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
di: Lin, Yiqi, et al.
Pubblicazione: (2026)
di: Lin, Yiqi, et al.
Pubblicazione: (2026)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
di: Gao, Xiangbo, et al.
Pubblicazione: (2026)
di: Gao, Xiangbo, et al.
Pubblicazione: (2026)
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
di: Zhang, Kai, et al.
Pubblicazione: (2023)
di: Zhang, Kai, et al.
Pubblicazione: (2023)
Are Video Reasoning Models Ready to Go Outside?
di: He, Yangfan, et al.
Pubblicazione: (2026)
di: He, Yangfan, et al.
Pubblicazione: (2026)
EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation
di: Wu, Xinyi, et al.
Pubblicazione: (2026)
di: Wu, Xinyi, et al.
Pubblicazione: (2026)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
di: Deng, Andong, et al.
Pubblicazione: (2024)
di: Deng, Andong, et al.
Pubblicazione: (2024)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
di: Pothiraj, Atin, et al.
Pubblicazione: (2025)
di: Pothiraj, Atin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
di: Yu, Shoubin, et al.
Pubblicazione: (2024) -
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
di: Yu, Shoubin, et al.
Pubblicazione: (2025) -
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
di: Li, Jialu, et al.
Pubblicazione: (2025) -
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
di: Wang, Ziyang, et al.
Pubblicazione: (2025) -
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
di: Yu, Shoubin, et al.
Pubblicazione: (2025)