TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Yubeen, Lee, Sangeun, Park, Chaewon, Cha, Junyeop, Park, Eunil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stage-Adaptive Reliability Modeling for Continuous Valence-Arousal Estimation
von: Lee, Yubeen, et al.
Veröffentlicht: (2026)
von: Lee, Yubeen, et al.
Veröffentlicht: (2026)
EmoVLM-KD: Fusing Distilled Expertise with Vision-Language Models for Visual Emotion Analysis
von: Lee, SangEun, et al.
Veröffentlicht: (2025)
von: Lee, SangEun, et al.
Veröffentlicht: (2025)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
von: Li, Zongyi, et al.
Veröffentlicht: (2025)
von: Li, Zongyi, et al.
Veröffentlicht: (2025)
Audio-Visual Separation with Hierarchical Fusion and Representation Alignment
von: Hu, Han, et al.
Veröffentlicht: (2025)
von: Hu, Han, et al.
Veröffentlicht: (2025)
EMO100DB: An Open Dataset of Improvised Songs with Emotion Data
von: Hwang, Daeun, et al.
Veröffentlicht: (2025)
von: Hwang, Daeun, et al.
Veröffentlicht: (2025)
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
von: Yang, Meng, et al.
Veröffentlicht: (2026)
von: Yang, Meng, et al.
Veröffentlicht: (2026)
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
von: Xu, Xinmeng, et al.
Veröffentlicht: (2026)
von: Xu, Xinmeng, et al.
Veröffentlicht: (2026)
MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
von: Shao, Yongqi, et al.
Veröffentlicht: (2025)
von: Shao, Yongqi, et al.
Veröffentlicht: (2025)
Interpretable Convolutional SyncNet
von: Park, Sungjoon, et al.
Veröffentlicht: (2024)
von: Park, Sungjoon, et al.
Veröffentlicht: (2024)
Unified Microphone Conversion: Many-to-Many Device Mapping via Feature-wise Linear Modulation
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2024)
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2024)
Microphone Conversion: Mitigating Device Variability in Sound Event Classification
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2024)
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2024)
Towards Practical Real-Time Low-Latency Music Source Separation
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
von: Wu, Junyu, et al.
Veröffentlicht: (2025)
Timed text extraction from Taiwanese Kua-á-hì TV series
von: Huang, Tzu-Hung, et al.
Veröffentlicht: (2026)
von: Huang, Tzu-Hung, et al.
Veröffentlicht: (2026)
Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
von: Feng, Zhou, et al.
Veröffentlicht: (2025)
von: Feng, Zhou, et al.
Veröffentlicht: (2025)
Gesture2Music: A Low-Latency Real-Time Framework for Continuous Gesture-Driven Music Generation
von: Jeyaraj, Rathinaraja, et al.
Veröffentlicht: (2025)
von: Jeyaraj, Rathinaraja, et al.
Veröffentlicht: (2025)
Sema: Semantic Transport for Real-Time Multimodal Agents
von: Meng, Jiaying, et al.
Veröffentlicht: (2026)
von: Meng, Jiaying, et al.
Veröffentlicht: (2026)
Solving Copyright Infringement on Short Video Platforms: Novel Datasets and an Audio Restoration Deep Learning Pipeline
von: Oh, Minwoo, et al.
Veröffentlicht: (2025)
von: Oh, Minwoo, et al.
Veröffentlicht: (2025)
A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text
von: Bang, Hayeon, et al.
Veröffentlicht: (2024)
von: Bang, Hayeon, et al.
Veröffentlicht: (2024)
Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering
von: Koh, Junyoung, et al.
Veröffentlicht: (2026)
von: Koh, Junyoung, et al.
Veröffentlicht: (2026)
Multimodal Fusion Method with Spatiotemporal Sequences and Relationship Learning for Valence-Arousal Estimation
von: Yu, Jun, et al.
Veröffentlicht: (2024)
von: Yu, Jun, et al.
Veröffentlicht: (2024)
PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
von: Bang, Hayeon, et al.
Veröffentlicht: (2025)
von: Bang, Hayeon, et al.
Veröffentlicht: (2025)
SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
von: Desai, Shail, et al.
Veröffentlicht: (2025)
von: Desai, Shail, et al.
Veröffentlicht: (2025)
Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition
von: Koo, Inyong, et al.
Veröffentlicht: (2026)
von: Koo, Inyong, et al.
Veröffentlicht: (2026)
Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations
von: Fu, Jingwen, et al.
Veröffentlicht: (2025)
von: Fu, Jingwen, et al.
Veröffentlicht: (2025)
HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations
von: Wang, Chunshi, et al.
Veröffentlicht: (2025)
von: Wang, Chunshi, et al.
Veröffentlicht: (2025)
Efficient Transformer-Based Piano Transcription With Sparse Attention Mechanisms
von: Wei, Weixing, et al.
Veröffentlicht: (2025)
von: Wei, Weixing, et al.
Veröffentlicht: (2025)
MusicWeaver: Composer-Style Structural Editing and Minute-Scale Coherent Music Generation
von: Wang, Xuanchen, et al.
Veröffentlicht: (2025)
von: Wang, Xuanchen, et al.
Veröffentlicht: (2025)
MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Stage-Adaptive Reliability Modeling for Continuous Valence-Arousal Estimation
von: Lee, Yubeen, et al.
Veröffentlicht: (2026) -
EmoVLM-KD: Fusing Distilled Expertise with Vision-Language Models for Visual Emotion Analysis
von: Lee, SangEun, et al.
Veröffentlicht: (2025) -
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
von: Choi, Suhwan, et al.
Veröffentlicht: (2025) -
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
von: Wang, Junbo, et al.
Veröffentlicht: (2025) -
Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)