Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baid, Ami, Xue, Zihui, Grauman, Kristen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Personal Visual Context Learning in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
von: Adebi, Daniel, et al.
Veröffentlicht: (2025)
von: Adebi, Daniel, et al.
Veröffentlicht: (2025)
Learning Object State Changes in Videos: An Open-World Perspective
von: Xue, Zihui, et al.
Veröffentlicht: (2023)
von: Xue, Zihui, et al.
Veröffentlicht: (2023)
Progress-Aware Video Frame Captioning
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
Seeing the Arrow of Time in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
von: Chen, Changan, et al.
Veröffentlicht: (2024)
von: Chen, Changan, et al.
Veröffentlicht: (2024)
Detours for Navigating Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
von: Xue, Zihui, et al.
Veröffentlicht: (2024)
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos
von: Luo, Mi, et al.
Veröffentlicht: (2024)
von: Luo, Mi, et al.
Veröffentlicht: (2024)
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
von: Luo, Mi, et al.
Veröffentlicht: (2025)
von: Luo, Mi, et al.
Veröffentlicht: (2025)
SPOC: Spatially-Progressing Object State Change Segmentation in Video
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
von: Mandikal, Priyanka, et al.
Veröffentlicht: (2025)
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2024)
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2024)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
von: An, Joungbin, et al.
Veröffentlicht: (2025)
von: An, Joungbin, et al.
Veröffentlicht: (2025)
Seeing without Pixels: Perception from Camera Trajectories
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
Learning Skill-Attributes for Transferable Assessment in Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2025)
Vid2Coach: Transforming How-To Videos into Task Assistants
von: Huh, Mina, et al.
Veröffentlicht: (2025)
von: Huh, Mina, et al.
Veröffentlicht: (2025)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
von: An, Joungbin, et al.
Veröffentlicht: (2026)
von: An, Joungbin, et al.
Veröffentlicht: (2026)
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2026)
von: Somayazulu, Arjun, et al.
Veröffentlicht: (2026)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
EgoExo-WM: Unlocking Exo Video for Ego World Models
von: Tran, Danny, et al.
Veröffentlicht: (2026)
von: Tran, Danny, et al.
Veröffentlicht: (2026)
FIction: 4D Future Interaction Prediction from Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
Complementary and Contrastive Learning for Audio-Visual Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
von: Kong, Zhe, et al.
Veröffentlicht: (2025)
von: Kong, Zhe, et al.
Veröffentlicht: (2025)
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
von: Wu, Chi Hsuan, et al.
Veröffentlicht: (2025)
AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2025)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2025)
SportSkills: Physical Skill Learning from Sports Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
von: Su, Zibo, et al.
Veröffentlicht: (2025)
von: Su, Zibo, et al.
Veröffentlicht: (2025)
Efficient Audio-Visual Fusion for Video Classification
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
von: Wang, Kai, et al.
Veröffentlicht: (2024)
von: Wang, Kai, et al.
Veröffentlicht: (2024)
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
von: Ishikawa, Yuchi, et al.
Veröffentlicht: (2025)
von: Ishikawa, Yuchi, et al.
Veröffentlicht: (2025)
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy
von: Ni, Minheng, et al.
Veröffentlicht: (2024)
von: Ni, Minheng, et al.
Veröffentlicht: (2024)
MistExit: Learning to Exit for Early Mistake Detection in Procedural Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2026)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2026)
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2024)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
von: Yu, Li, et al.
Veröffentlicht: (2025)
von: Yu, Li, et al.
Veröffentlicht: (2025)
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis
von: Wu, Xuecheng, et al.
Veröffentlicht: (2025)
von: Wu, Xuecheng, et al.
Veröffentlicht: (2025)
ExpertAF: Expert Actionable Feedback from Video
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2024)
AudioScenic: Audio-Driven Video Scene Editing
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Personal Visual Context Learning in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2026) -
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
von: Adebi, Daniel, et al.
Veröffentlicht: (2025) -
Learning Object State Changes in Videos: An Open-World Perspective
von: Xue, Zihui, et al.
Veröffentlicht: (2023) -
Progress-Aware Video Frame Captioning
von: Xue, Zihui, et al.
Veröffentlicht: (2024) -
Seeing the Arrow of Time in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2025)