LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Bryan, Li, Yuliang, Lv, Zhaoyang, Xia, Haijun, Xu, Yan, Sodhi, Raj |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
by: Yeh, Catherine, et al.
Published: (2026)
by: Yeh, Catherine, et al.
Published: (2026)
DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
by: Tian, Ye, et al.
Published: (2025)
by: Tian, Ye, et al.
Published: (2025)
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
by: Kyaw, Alexander Htet, et al.
Published: (2025)
by: Kyaw, Alexander Htet, et al.
Published: (2025)
Laugh at Your Own Pace: Basic Performance Evaluation of Language Learning Assistance by Adjustment of Video Playback Speeds Based on Laughter Detection
by: Nishida, Naoto, et al.
Published: (2025)
by: Nishida, Naoto, et al.
Published: (2025)
Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justice
by: Wasi, Azmine Toushik, et al.
Published: (2025)
by: Wasi, Azmine Toushik, et al.
Published: (2025)
Livia: An Emotion-Aware AR Companion Powered by Modular AI Agents and Progressive Memory Compression
by: Xi, Rui, et al.
Published: (2025)
by: Xi, Rui, et al.
Published: (2025)
Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model
by: Chen, Xiaolin, et al.
Published: (2022)
by: Chen, Xiaolin, et al.
Published: (2022)
Counterfactual Reasoning Using Predicted Latent Personality Dimensions for Optimizing Persuasion Outcome
by: Zeng, Donghuo, et al.
Published: (2024)
by: Zeng, Donghuo, et al.
Published: (2024)
ANVIL: Analogies and Videos for Lecturers
by: Noviello, Yuri, et al.
Published: (2026)
by: Noviello, Yuri, et al.
Published: (2026)
MapStory: Prototyping Editable Map Animations with LLM Agents
by: Gunturu, Aditya, et al.
Published: (2025)
by: Gunturu, Aditya, et al.
Published: (2025)
IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models
by: Rahimi, Hamed, et al.
Published: (2026)
by: Rahimi, Hamed, et al.
Published: (2026)
Modular Conversational Agents for Surveys and Interviews
by: Yu, Jiangbo, et al.
Published: (2024)
by: Yu, Jiangbo, et al.
Published: (2024)
Photoshop Batch Rendering Using Actions for Stylistic Video Editing
by: De La Fuente, Tessa
Published: (2025)
by: De La Fuente, Tessa
Published: (2025)
DreamLLM-3D: Affective Dream Reliving using Large Language Model and 3D Generative AI
by: Liu, Pinyao, et al.
Published: (2025)
by: Liu, Pinyao, et al.
Published: (2025)
Code2Video: A Code-centric Paradigm for Educational Video Generation
by: Chen, Yanzhe, et al.
Published: (2025)
by: Chen, Yanzhe, et al.
Published: (2025)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025)
by: Zhao, Baoquan, et al.
Published: (2025)
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
Focus360: Guiding User Attention in Immersive Videos for VR
by: Silva, Paulo Vitor S., et al.
Published: (2026)
by: Silva, Paulo Vitor S., et al.
Published: (2026)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
by: Wan, Ninghao, et al.
Published: (2026)
by: Wan, Ninghao, et al.
Published: (2026)
Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports
by: Du, Changde, et al.
Published: (2025)
by: Du, Changde, et al.
Published: (2025)
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
by: Chai, Yuxiang, et al.
Published: (2024)
by: Chai, Yuxiang, et al.
Published: (2024)
An Empirical Evaluation of AI-Powered Non-Player Characters' Perceived Realism and Performance in Virtual Reality Environments
by: Korkiakoski, Mikko, et al.
Published: (2025)
by: Korkiakoski, Mikko, et al.
Published: (2025)
Evaluating the Usability of Microgestures for Text Editing Tasks in Virtual Reality
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Through the Looking-Glass: AI-Mediated Video Communication Reduces Interpersonal Trust and Confidence in Judgments
by: Fernández, Nelson Navajas, et al.
Published: (2026)
by: Fernández, Nelson Navajas, et al.
Published: (2026)
ICE: Interactive 3D Game Character Editing via Dialogue
by: Wu, Haoqian, et al.
Published: (2024)
by: Wu, Haoqian, et al.
Published: (2024)
Simulacra Naturae: Generative Ecosystem driven by Agent-Based Simulations and Brain Organoid Collective Intelligence
by: Manoudaki, Nefeli, et al.
Published: (2025)
by: Manoudaki, Nefeli, et al.
Published: (2025)
Applying LLM-Powered Virtual Humans to Child Interviews in Child-Centered Design
by: Li, Linshi, et al.
Published: (2025)
by: Li, Linshi, et al.
Published: (2025)
Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization
by: Wen, Zhibin, et al.
Published: (2024)
by: Wen, Zhibin, et al.
Published: (2024)
Save It for the "Hot" Day: An LLM-Empowered Visual Analytics System for Heat Risk Management
by: Li, Haobo, et al.
Published: (2024)
by: Li, Haobo, et al.
Published: (2024)
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
by: Lin, Ronghao, et al.
Published: (2025)
by: Lin, Ronghao, et al.
Published: (2025)
VideoMap: Supporting Video Editing Exploration, Brainstorming, and Prototyping in the Latent Space
by: Lin, David Chuan-En, et al.
Published: (2022)
by: Lin, David Chuan-En, et al.
Published: (2022)
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
by: Tsangko, Iosif, et al.
Published: (2026)
by: Tsangko, Iosif, et al.
Published: (2026)
Hue4U: Real-Time Personalized Color Correction in Augmented Reality
by: Qin, Jingwen, et al.
Published: (2025)
by: Qin, Jingwen, et al.
Published: (2025)
Editing Physiological Signals in Videos Using Latent Representations
by: Zhou, Tianwen, et al.
Published: (2025)
by: Zhou, Tianwen, et al.
Published: (2025)
FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Contexts
by: Kawamura, Kazuki, et al.
Published: (2024)
by: Kawamura, Kazuki, et al.
Published: (2024)
Development of Immersive Virtual and Augmented Reality-Based Joint Attention Training Platform for Children with Autism
by: Samantaray, Ashirbad, et al.
Published: (2025)
by: Samantaray, Ashirbad, et al.
Published: (2025)
Video-Mediated Emotion Disclosure: Expressions of Fear, Sadness, and Joy by People with Schizophrenia on YouTube
by: Liu, Jiaying Lizzy, et al.
Published: (2025)
by: Liu, Jiaying Lizzy, et al.
Published: (2025)
MV-Crafter: An Intelligent System for Music-guided Video Generation
by: Chen, Chuer, et al.
Published: (2025)
by: Chen, Chuer, et al.
Published: (2025)
Memento: Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding Navigation
by: Ghosh, Indrajeet, et al.
Published: (2025)
by: Ghosh, Indrajeet, et al.
Published: (2025)
Similar Items
-
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
by: Yeh, Catherine, et al.
Published: (2026) -
DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
by: Tian, Ye, et al.
Published: (2025) -
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
by: Kyaw, Alexander Htet, et al.
Published: (2025) -
Laugh at Your Own Pace: Basic Performance Evaluation of Language Learning Assistance by Adjustment of Video Playback Speeds Based on Laughter Detection
by: Nishida, Naoto, et al.
Published: (2025) -
Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justice
by: Wasi, Azmine Toushik, et al.
Published: (2025)