VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Shoubin, Liu, Difan, Ma, Ziqiao, Hong, Yicong, Zhou, Yang, Tan, Hao, Chai, Joyce, Bansal, Mohit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
by: Yoon, Jaehong, et al.
Published: (2024)
by: Yoon, Jaehong, et al.
Published: (2024)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
by: Yu, Shoubin, et al.
Published: (2024)
by: Yu, Shoubin, et al.
Published: (2024)
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models
by: Ma, Ziqiao, et al.
Published: (2023)
by: Ma, Ziqiao, et al.
Published: (2023)
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
by: Yu, Shoubin, et al.
Published: (2025)
by: Yu, Shoubin, et al.
Published: (2025)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
by: Wang, Ziyang, et al.
Published: (2024)
by: Wang, Ziyang, et al.
Published: (2024)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
by: Lee, Daeun, et al.
Published: (2026)
by: Lee, Daeun, et al.
Published: (2026)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
by: Li, Jialu, et al.
Published: (2025)
by: Li, Jialu, et al.
Published: (2025)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
Babysit A Language Model From Scratch: Interactive Language Learning by Trials and Demonstrations
by: Ma, Ziqiao, et al.
Published: (2024)
by: Ma, Ziqiao, et al.
Published: (2024)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
by: Wang, Ziyang, et al.
Published: (2026)
by: Wang, Ziyang, et al.
Published: (2026)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
by: Liang, Yiming, et al.
Published: (2026)
by: Liang, Yiming, et al.
Published: (2026)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
The Mechanistic Emergence of Symbol Grounding in Language Models
by: Wu, Shuyu, et al.
Published: (2025)
by: Wu, Shuyu, et al.
Published: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models
by: Ma, Ziqiao, et al.
Published: (2023)
by: Ma, Ziqiao, et al.
Published: (2023)
Balancing Faithfulness and Performance in Reasoning via Multi-Listener Soft Execution
by: Sivakumaran, Nithin, et al.
Published: (2026)
by: Sivakumaran, Nithin, et al.
Published: (2026)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
by: Zhou, Gengze, et al.
Published: (2024)
by: Zhou, Gengze, et al.
Published: (2024)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
by: Wang, Zun, et al.
Published: (2024)
by: Wang, Zun, et al.
Published: (2024)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
by: Prasad, Archiki, et al.
Published: (2023)
by: Prasad, Archiki, et al.
Published: (2023)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
DriVLMe: Enhancing LLM-based Autonomous Driving Agents with Embodied and Social Experiences
by: Huang, Yidong, et al.
Published: (2024)
by: Huang, Yidong, et al.
Published: (2024)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
by: Deng, Andong, et al.
Published: (2024)
by: Deng, Andong, et al.
Published: (2024)
STAR: A Benchmark for Situated Reasoning in Real-World Videos
by: Wu, Bo, et al.
Published: (2024)
by: Wu, Bo, et al.
Published: (2024)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
by: Zhou, Gengze, et al.
Published: (2025)
by: Zhou, Gengze, et al.
Published: (2025)
4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
by: Ma, Ziqiao, et al.
Published: (2025)
by: Ma, Ziqiao, et al.
Published: (2025)
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
by: Ma, Ziqiao, et al.
Published: (2025)
by: Ma, Ziqiao, et al.
Published: (2025)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
by: Lee, Daeun, et al.
Published: (2024)
by: Lee, Daeun, et al.
Published: (2024)
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
by: Yoon, Jaehong, et al.
Published: (2024)
by: Yoon, Jaehong, et al.
Published: (2024)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
by: Lin, Han, et al.
Published: (2023)
by: Lin, Han, et al.
Published: (2023)
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?
by: Hase, Peter, et al.
Published: (2024)
by: Hase, Peter, et al.
Published: (2024)
Conflict-Resolving and Sharpness-Aware Minimization for Generalized Knowledge Editing with Multiple Updates
by: Nguyen, Duy, et al.
Published: (2026)
by: Nguyen, Duy, et al.
Published: (2026)
VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
by: Vaishnav, Mohit, et al.
Published: (2026)
by: Vaishnav, Mohit, et al.
Published: (2026)
TimeRefine: Temporal Grounding with Time Refining Video LLM
by: Wang, Xizi, et al.
Published: (2024)
by: Wang, Xizi, et al.
Published: (2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
by: Wang, Zun, et al.
Published: (2024)
by: Wang, Zun, et al.
Published: (2024)
Similar Items
-
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
by: Yoon, Jaehong, et al.
Published: (2024) -
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
by: Yu, Shoubin, et al.
Published: (2024) -
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models
by: Ma, Ziqiao, et al.
Published: (2023) -
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation
by: Yu, Shoubin, et al.
Published: (2025) -
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
by: Wang, Ziyang, et al.
Published: (2025)