MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
Fuente:
arXiv
Saved in:
| Main Authors: | Ahire, Vrushank, Shah, Kunal, Khan, Mudasir Nazir, Pakhale, Nikhil, Sookha, Lownish Rai, Ganaie, M. A., Dhall, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative Techniques
by: Sookha, Lownish Rai, et al.
Published: (2025)
by: Sookha, Lownish Rai, et al.
Published: (2025)
MIP-GAF: A MLLM-annotated Benchmark for Most Important Person Localization and Group Context Understanding
by: Madan, Surbhi, et al.
Published: (2024)
by: Madan, Surbhi, et al.
Published: (2024)
SFANet: Spatial-Frequency Attention Network for Deepfake Detection
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
Deep Learning-Based Tracking and Lineage Reconstruction of Ligament Breakup
by: Ahire, Vrushank, et al.
Published: (2026)
by: Ahire, Vrushank, et al.
Published: (2026)
LayLens: Improving Deepfake Understanding through Simplified Explanations
by: Narang, Abhijeet, et al.
Published: (2025)
by: Narang, Abhijeet, et al.
Published: (2025)
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
by: Gupta, Parul, et al.
Published: (2025)
by: Gupta, Parul, et al.
Published: (2025)
TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation
by: Lee, Yubeen, et al.
Published: (2025)
by: Lee, Yubeen, et al.
Published: (2025)
Multi-modal Speech Emotion Recognition via Feature Distribution Adaptation Network
by: Li, Shaokai, et al.
Published: (2024)
by: Li, Shaokai, et al.
Published: (2024)
Stage-Adaptive Reliability Modeling for Continuous Valence-Arousal Estimation
by: Lee, Yubeen, et al.
Published: (2026)
by: Lee, Yubeen, et al.
Published: (2026)
Granular Ball Twin Support Vector Machine with Universum Data
by: Ganaie, M. A., et al.
Published: (2024)
by: Ganaie, M. A., et al.
Published: (2024)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
by: Zhang, Haoyang, et al.
Published: (2025)
by: Zhang, Haoyang, et al.
Published: (2025)
MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening
by: Ahire, Vrushank, et al.
Published: (2026)
by: Ahire, Vrushank, et al.
Published: (2026)
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
by: Li, Deng, et al.
Published: (2024)
by: Li, Deng, et al.
Published: (2024)
Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
by: Hu, Xiaowan, et al.
Published: (2024)
by: Hu, Xiaowan, et al.
Published: (2024)
Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
by: Shi, Haoxiang, et al.
Published: (2024)
by: Shi, Haoxiang, et al.
Published: (2024)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
by: Choi, Suhwan, et al.
Published: (2025)
by: Choi, Suhwan, et al.
Published: (2025)
Visual Grounding with Multi-modal Conditional Adaptation
by: Yao, Ruilin, et al.
Published: (2024)
by: Yao, Ruilin, et al.
Published: (2024)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
by: Yang, Zequn, et al.
Published: (2024)
by: Yang, Zequn, et al.
Published: (2024)
Granular Ball K-Class Twin Support Vector Classifier
by: Ganaie, M. A., et al.
Published: (2024)
by: Ganaie, M. A., et al.
Published: (2024)
A Unified Framework for EEG Seizure Detection Using Universum-Integrated Generalized Eigenvalues Proximal Support Vector Machine
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
by: Chow, Wei, et al.
Published: (2024)
by: Chow, Wei, et al.
Published: (2024)
Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
by: Zhang, Zhihao, et al.
Published: (2023)
by: Zhang, Zhihao, et al.
Published: (2023)
Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
by: Zhou, Zhiyuan, et al.
Published: (2026)
by: Zhou, Zhiyuan, et al.
Published: (2026)
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
by: Li, Zeju, et al.
Published: (2024)
by: Li, Zeju, et al.
Published: (2024)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
by: Dai, Guangyu, et al.
Published: (2025)
by: Dai, Guangyu, et al.
Published: (2025)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
by: Tang, Yolo Yunlong, et al.
Published: (2022)
by: Tang, Yolo Yunlong, et al.
Published: (2022)
IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
by: Yuan, Bo, et al.
Published: (2024)
by: Yuan, Bo, et al.
Published: (2024)
Multi-scale Attention Guided Pose Transfer
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
by: Ding, Peng, et al.
Published: (2024)
by: Ding, Peng, et al.
Published: (2024)
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
by: Gao, Lancheng, et al.
Published: (2025)
by: Gao, Lancheng, et al.
Published: (2025)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
by: Huang, Dawei, et al.
Published: (2025)
by: Huang, Dawei, et al.
Published: (2025)
ReFiNe: Recursive Field Networks for Cross-modal Multi-scene Representation
by: Zakharov, Sergey, et al.
Published: (2024)
by: Zakharov, Sergey, et al.
Published: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
FortisAVQA and MAVEN: a Benchmark Dataset and Debiasing Framework for Robust Multimodal Reasoning
by: Ma, Jie, et al.
Published: (2025)
by: Ma, Jie, et al.
Published: (2025)
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
MLANet: Multi-Level Attention Network with Sub-instruction for Continuous Vision-and-Language Navigation
by: He, Zongtao, et al.
Published: (2023)
by: He, Zongtao, et al.
Published: (2023)
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
Similar Items
-
A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative Techniques
by: Sookha, Lownish Rai, et al.
Published: (2025) -
MIP-GAF: A MLLM-annotated Benchmark for Most Important Person Localization and Group Context Understanding
by: Madan, Surbhi, et al.
Published: (2024) -
SFANet: Spatial-Frequency Attention Network for Deepfake Detection
by: Ahire, Vrushank, et al.
Published: (2025) -
Deep Learning-Based Tracking and Lineage Reconstruction of Ligament Breakup
by: Ahire, Vrushank, et al.
Published: (2026) -
LayLens: Improving Deepfake Understanding through Simplified Explanations
by: Narang, Abhijeet, et al.
Published: (2025)