Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zinuo, Zhang, Xian, Guo, Yongxin, Bennamoun, Mohammed, Boussaid, Farid, Dwivedi, Girish, Gong, Luqi, Ke, Qiuhong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning
by: Li, Zinuo, et al.
Published: (2026)
by: Li, Zinuo, et al.
Published: (2026)
3D Brain and Heart Volume Generative Models: A Survey
by: Liu, Yanbin, et al.
Published: (2022)
by: Liu, Yanbin, et al.
Published: (2022)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
by: Zhang, Xian, et al.
Published: (2025)
by: Zhang, Xian, et al.
Published: (2025)
Hybrid Transformer-Mamba Architecture for Weakly Supervised Volumetric Medical Segmentation
by: Lyu, Yiheng, et al.
Published: (2025)
by: Lyu, Yiheng, et al.
Published: (2025)
LatentMove: Towards Complex Human Movement Video Generation
by: Taghipour, Ashkan, et al.
Published: (2025)
by: Taghipour, Ashkan, et al.
Published: (2025)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
by: Taghipour, Ashkan, et al.
Published: (2026)
by: Taghipour, Ashkan, et al.
Published: (2026)
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
by: Taghipour, Ashkan, et al.
Published: (2024)
by: Taghipour, Ashkan, et al.
Published: (2024)
DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition
by: Zhu, Jingmin, et al.
Published: (2025)
by: Zhu, Jingmin, et al.
Published: (2025)
Generalized Closed-form Formulae for Feature-based Subpixel Alignment in Patch-based Matching
by: Jospin, Laurent Valentin, et al.
Published: (2021)
by: Jospin, Laurent Valentin, et al.
Published: (2021)
Admitting Ignorance Helps the Video Question Answering Models to Answer
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting
by: Taghipour, Ashkan, et al.
Published: (2025)
by: Taghipour, Ashkan, et al.
Published: (2025)
Dynamic Neural Surfaces for Elastic 4D Shape Representation and Analysis
by: Nizamani, Awais, et al.
Published: (2025)
by: Nizamani, Awais, et al.
Published: (2025)
Auxiliary Tasks Enhanced Dual-affinity Learning for Weakly Supervised Semantic Segmentation
by: Xu, Lian, et al.
Published: (2024)
by: Xu, Lian, et al.
Published: (2024)
Fact or Fake? Assessing the Role of Deepfake Detectors in Multimodal Misinformation Detection
by: Sagar, A S M Sharifuzzaman, et al.
Published: (2026)
by: Sagar, A S M Sharifuzzaman, et al.
Published: (2026)
Box It to Bind It: Unified Layout Control and Attribute Binding in T2I Diffusion Models
by: Taghipour, Ashkan, et al.
Published: (2024)
by: Taghipour, Ashkan, et al.
Published: (2024)
WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
by: Zhang, Binbin, et al.
Published: (2025)
by: Zhang, Binbin, et al.
Published: (2025)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
by: Wu, Wenxuan, et al.
Published: (2025)
by: Wu, Wenxuan, et al.
Published: (2025)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Cocktail-Party Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
by: Yan, Canxiang, et al.
Published: (2025)
by: Yan, Canxiang, et al.
Published: (2025)
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
by: Gong, Kaixiong, et al.
Published: (2024)
by: Gong, Kaixiong, et al.
Published: (2024)
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
by: Romero-Díaz, Jacobo, et al.
Published: (2025)
by: Romero-Díaz, Jacobo, et al.
Published: (2025)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
by: Lin, Junyan, et al.
Published: (2026)
by: Lin, Junyan, et al.
Published: (2026)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
by: Zhang, William, et al.
Published: (2024)
by: Zhang, William, et al.
Published: (2024)
A Riemannian Approach for Spatiotemporal Analysis and Generation of 4D Tree-shaped Structures
by: Khanam, Tahmina, et al.
Published: (2024)
by: Khanam, Tahmina, et al.
Published: (2024)
A Riemannian Framework for the Elastic Analysis of the Spatiotemporal Variability in the Shape and Structure of Tree-like 4D Objects
by: Khanam, Tahmina, et al.
Published: (2025)
by: Khanam, Tahmina, et al.
Published: (2025)
SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
by: Wang, Ning, et al.
Published: (2026)
by: Wang, Ning, et al.
Published: (2026)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
by: Manakul, Potsawee, et al.
Published: (2025)
by: Manakul, Potsawee, et al.
Published: (2025)
ViSpeR: Multilingual Audio-Visual Speech Recognition
by: Narayan, Sanath, et al.
Published: (2024)
by: Narayan, Sanath, et al.
Published: (2024)
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
by: Khorrami, Khazar, et al.
Published: (2021)
by: Khorrami, Khazar, et al.
Published: (2021)
UIFormer: A Unified Transformer-based Framework for Incremental Few-Shot Object Detection and Instance Segmentation
by: Zhang, Chengyuan, et al.
Published: (2024)
by: Zhang, Chengyuan, et al.
Published: (2024)
Listen, Attend, Understand: a Regularization Technique for Stable E2E Speech Translation Training on High Variance labels
by: Diarra, Yacouba, et al.
Published: (2026)
by: Diarra, Yacouba, et al.
Published: (2026)
URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
by: Shi, Yongxin, et al.
Published: (2025)
by: Shi, Yongxin, et al.
Published: (2025)
Can Speech LLMs Think while Listening?
by: Shih, Yi-Jen, et al.
Published: (2025)
by: Shih, Yi-Jen, et al.
Published: (2025)
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
by: Huang, Ailin, et al.
Published: (2025)
by: Huang, Ailin, et al.
Published: (2025)
Language-based Audio Moment Retrieval
by: Munakata, Hokuto, et al.
Published: (2024)
by: Munakata, Hokuto, et al.
Published: (2024)
Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation
by: Zhu, Jingmin, et al.
Published: (2025)
by: Zhu, Jingmin, et al.
Published: (2025)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
by: Ginjala, Srishti, et al.
Published: (2026)
by: Ginjala, Srishti, et al.
Published: (2026)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
by: Goncalves, Lucas, et al.
Published: (2024)
by: Goncalves, Lucas, et al.
Published: (2024)
Similar Items
-
STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning
by: Li, Zinuo, et al.
Published: (2026) -
3D Brain and Heart Volume Generative Models: A Survey
by: Liu, Yanbin, et al.
Published: (2022) -
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
by: Zhang, Xian, et al.
Published: (2025) -
Hybrid Transformer-Mamba Architecture for Weakly Supervised Volumetric Medical Segmentation
by: Lyu, Yiheng, et al.
Published: (2025) -
LatentMove: Towards Complex Human Movement Video Generation
by: Taghipour, Ashkan, et al.
Published: (2025)