FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhi-Song, Courant, Robin, Kalogeiton, Vicky |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MCDubber: Multimodal Context-Aware Expressive Video Dubbing
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
Video-Guided Foley Sound Generation with Multimodal Controls
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
On the Audio Hallucinations in Large Audio-Video Language Models
von: Nishimura, Taichi, et al.
Veröffentlicht: (2024)
von: Nishimura, Taichi, et al.
Veröffentlicht: (2024)
Towards Accurate Lip-to-Speech Synthesis in-the-Wild
von: Hegde, Sindhu, et al.
Veröffentlicht: (2024)
von: Hegde, Sindhu, et al.
Veröffentlicht: (2024)
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
von: Jiang, Xilin, et al.
Veröffentlicht: (2026)
von: Jiang, Xilin, et al.
Veröffentlicht: (2026)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
Interpretable Convolutional SyncNet
von: Park, Sungjoon, et al.
Veröffentlicht: (2024)
von: Park, Sungjoon, et al.
Veröffentlicht: (2024)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
von: Zhao, Fuzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Fuzheng, et al.
Veröffentlicht: (2024)
Audio-Assisted Face Video Restoration with Temporal and Identity Complementary Learning
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
von: Chen, Changan, et al.
Veröffentlicht: (2024)
von: Chen, Changan, et al.
Veröffentlicht: (2024)
VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2024)
Video-to-Audio Generation with Hidden Alignment
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
Temporally Aligned Audio for Video with Autoregression
von: Viertola, Ilpo, et al.
Veröffentlicht: (2024)
von: Viertola, Ilpo, et al.
Veröffentlicht: (2024)
Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation
von: Wang, Baisen, et al.
Veröffentlicht: (2024)
von: Wang, Baisen, et al.
Veröffentlicht: (2024)
Digit Recognition using Multimodal Spiking Neural Networks
von: Bjorndahl, William, et al.
Veröffentlicht: (2024)
von: Bjorndahl, William, et al.
Veröffentlicht: (2024)
Diffusion Models as Masked Audio-Video Learners
von: Nunez, Elvis, et al.
Veröffentlicht: (2023)
von: Nunez, Elvis, et al.
Veröffentlicht: (2023)
Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
von: Yu, Fei, et al.
Veröffentlicht: (2024)
von: Yu, Fei, et al.
Veröffentlicht: (2024)
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
von: You, Wenhao, et al.
Veröffentlicht: (2025)
von: You, Wenhao, et al.
Veröffentlicht: (2025)
Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
von: Yang, Qi, et al.
Veröffentlicht: (2024)
von: Yang, Qi, et al.
Veröffentlicht: (2024)
Read, Watch and Scream! Sound Generation from Text and Video
von: Jeong, Yujin, et al.
Veröffentlicht: (2024)
von: Jeong, Yujin, et al.
Veröffentlicht: (2024)
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
von: Rai, Aashish, et al.
Veröffentlicht: (2024)
von: Rai, Aashish, et al.
Veröffentlicht: (2024)
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
von: Oorloff, Trevine, et al.
Veröffentlicht: (2024)
von: Oorloff, Trevine, et al.
Veröffentlicht: (2024)
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
von: Hong, Joanna, et al.
Veröffentlicht: (2025)
von: Hong, Joanna, et al.
Veröffentlicht: (2025)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
Taming Data and Transformers for Audio Generation
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2024)
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2024)
Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
von: Jiang, Xilin, et al.
Veröffentlicht: (2025)
von: Jiang, Xilin, et al.
Veröffentlicht: (2025)
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
von: Min, Anna, et al.
Veröffentlicht: (2025)
von: Min, Anna, et al.
Veröffentlicht: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
von: Guan, Yiwen, et al.
Veröffentlicht: (2024)
von: Guan, Yiwen, et al.
Veröffentlicht: (2024)
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
von: Low, Chetwin, et al.
Veröffentlicht: (2025)
von: Low, Chetwin, et al.
Veröffentlicht: (2025)
Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper
von: Chen, Gehui, et al.
Veröffentlicht: (2025)
von: Chen, Gehui, et al.
Veröffentlicht: (2025)
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
AutoMV: An Automatic Multi-Agent System for Music Video Generation
von: Tang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaoxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MCDubber: Multimodal Context-Aware Expressive Video Dubbing
von: Zhao, Yuan, et al.
Veröffentlicht: (2024) -
Video-Guided Foley Sound Generation with Multimodal Controls
von: Chen, Ziyang, et al.
Veröffentlicht: (2024) -
On the Audio Hallucinations in Large Audio-Video Language Models
von: Nishimura, Taichi, et al.
Veröffentlicht: (2024) -
Towards Accurate Lip-to-Speech Synthesis in-the-Wild
von: Hegde, Sindhu, et al.
Veröffentlicht: (2024) -
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)