AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Chaubey, Ashutosh, Pang, Jiacheng, Siniukov, Maksim, Soleymani, Mohammad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization
by: Chaubey, Ashutosh, et al.
Published: (2026)
by: Chaubey, Ashutosh, et al.
Published: (2026)
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
by: Chaubey, Ashutosh, et al.
Published: (2025)
by: Chaubey, Ashutosh, et al.
Published: (2025)
GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization
by: Jin, Zhangyu, et al.
Published: (2026)
by: Jin, Zhangyu, et al.
Published: (2026)
DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion
by: Siniukov, Maksim, et al.
Published: (2025)
by: Siniukov, Maksim, et al.
Published: (2025)
Revisiting Emotions Representation for Recognition in the Wild
by: Neto, Joao Baptista Cardia, et al.
Published: (2026)
by: Neto, Joao Baptista Cardia, et al.
Published: (2026)
Exploring Emotion Expression Recognition in Older Adults Interacting with a Virtual Coach
by: Palmero, Cristina, et al.
Published: (2023)
by: Palmero, Cristina, et al.
Published: (2023)
Improve accessibility for Low Vision and Blind people using Machine Learning and Computer Vision
by: Shukurov, Jasur
Published: (2024)
by: Shukurov, Jasur
Published: (2024)
Customizable Avatars with Dynamic Facial Action Coded Expressions (CADyFACE) for Improved User Engagement
by: Witherow, Megan A., et al.
Published: (2024)
by: Witherow, Megan A., et al.
Published: (2024)
An Efficient and Streaming Audio Visual Active Speaker Detection System
by: Kundu, Arnav, et al.
Published: (2024)
by: Kundu, Arnav, et al.
Published: (2024)
Learning Confident Classifiers in the Presence of Label Noise
by: Hashmi, Asma Ahmed, et al.
Published: (2023)
by: Hashmi, Asma Ahmed, et al.
Published: (2023)
How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception
by: Shahzad, Sahibzada Adil, et al.
Published: (2024)
by: Shahzad, Sahibzada Adil, et al.
Published: (2024)
ShelfHelp: Empowering Humans to Perform Vision-Independent Manipulation Tasks with a Socially Assistive Robotic Cane
by: Agrawal, Shivendra, et al.
Published: (2024)
by: Agrawal, Shivendra, et al.
Published: (2024)
EmoGene: Audio-Driven Emotional 3D Talking-Head Generation
by: Wang, Wenqing, et al.
Published: (2024)
by: Wang, Wenqing, et al.
Published: (2024)
Deep Learning Based Approach to Enhanced Recognition of Emotions and Behavioral Patterns of Autistic Children
by: R, Nelaka K. A., et al.
Published: (2025)
by: R, Nelaka K. A., et al.
Published: (2025)
FERGI: Automatic Scoring of User Preferences for Text-to-Image Generation from Spontaneous Facial Expression Reaction
by: Feng, Shuangquan, et al.
Published: (2023)
by: Feng, Shuangquan, et al.
Published: (2023)
Generalization of CNNs on Relational Reasoning with Bar Charts
by: Cui, Zhenxing, et al.
Published: (2025)
by: Cui, Zhenxing, et al.
Published: (2025)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
by: Xenos, Alexandros, et al.
Published: (2024)
by: Xenos, Alexandros, et al.
Published: (2024)
Voting-based Multimodal Automatic Deception Detection
by: Touma, Lana, et al.
Published: (2023)
by: Touma, Lana, et al.
Published: (2023)
Temporal Structure Matters for Efficient Test-Time Adaptation in Wearable Human Activity Recognition
by: Zhou, Zishu, et al.
Published: (2026)
by: Zhou, Zishu, et al.
Published: (2026)
From Model Uncertainty to Human Attention: Localization-Aware Visual Cues for Scalable Annotation Review
by: Sbeyti, Moussa Kassem, et al.
Published: (2026)
by: Sbeyti, Moussa Kassem, et al.
Published: (2026)
Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
by: González-González, Manuela, et al.
Published: (2026)
by: González-González, Manuela, et al.
Published: (2026)
DeltaDorsal: Enhancing Hand Pose Estimation with Dorsal Features in Egocentric Views
by: Huang, William, et al.
Published: (2026)
by: Huang, William, et al.
Published: (2026)
CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
by: Si, Yonghao, et al.
Published: (2026)
by: Si, Yonghao, et al.
Published: (2026)
DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification
by: Zimmermann, Robert, et al.
Published: (2026)
by: Zimmermann, Robert, et al.
Published: (2026)
Interpretable facial dynamics as behavioral and perceptual traces of deepfakes
by: Murphy, Timothy Joseph, et al.
Published: (2026)
by: Murphy, Timothy Joseph, et al.
Published: (2026)
SurfaceXR: Fusing Smartwatch IMUs and Egocentric Hand Pose for Seamless Surface Interactions
by: Xu, Vasco, et al.
Published: (2026)
by: Xu, Vasco, et al.
Published: (2026)
Resource-Efficient Gesture Recognition through Convexified Attention
by: Schwartz, Daniel, et al.
Published: (2026)
by: Schwartz, Daniel, et al.
Published: (2026)
Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches
by: Zhang, He, et al.
Published: (2025)
by: Zhang, He, et al.
Published: (2025)
Intelligent Power Grid Design Review via Active Perception-Enabled Multimodal Large Language Models
by: Tan, Taoliang, et al.
Published: (2025)
by: Tan, Taoliang, et al.
Published: (2025)
CHiQPM: Calibrated Hierarchical Interpretable Image Classification
by: Norrenbrock, Thomas, et al.
Published: (2025)
by: Norrenbrock, Thomas, et al.
Published: (2025)
No Need to Sacrifice Data Quality for Quantity: Crowd-Informed Machine Annotation for Cost-Effective Understanding of Visual Data
by: Klugmann, Christopher, et al.
Published: (2024)
by: Klugmann, Christopher, et al.
Published: (2024)
Helios: An extremely low power event-based gesture recognition for always-on smart eyewear
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
SLIMBRAIN: Augmented Reality Real-Time Acquisition and Processing System For Hyperspectral Classification Mapping with Depth Information for In-Vivo Surgical Procedures
by: Sancho, Jaime, et al.
Published: (2024)
by: Sancho, Jaime, et al.
Published: (2024)
An Evaluation of Hybrid Annotation Workflows on High-Ambiguity Spatiotemporal Video Footage
by: Gutiérrez, Juan, et al.
Published: (2025)
by: Gutiérrez, Juan, et al.
Published: (2025)
VocalEyes: Enhancing Environmental Perception for the Visually Impaired through Vision-Language Models and Distance-Aware Object Detection
by: Chavan, Kunal, et al.
Published: (2025)
by: Chavan, Kunal, et al.
Published: (2025)
ExeChecker: Where Did I Go Wrong?
by: Gu, Yiwen, et al.
Published: (2024)
by: Gu, Yiwen, et al.
Published: (2024)
Evaluating the Utility of Conformal Prediction Sets for AI-Advised Image Labeling
by: Zhang, Dongping, et al.
Published: (2024)
by: Zhang, Dongping, et al.
Published: (2024)
A General Model for Detecting Learner Engagement: Implementation and Evaluation
by: Malekshahi, Somayeh, et al.
Published: (2024)
by: Malekshahi, Somayeh, et al.
Published: (2024)
Using a CNN Model to Assess Paintings' Creativity
by: Zhang, Zhehan, et al.
Published: (2024)
by: Zhang, Zhehan, et al.
Published: (2024)
Methodology to Deploy CNN-Based Computer Vision Models on Immersive Wearable Devices
by: Malek, Kaveh, et al.
Published: (2024)
by: Malek, Kaveh, et al.
Published: (2024)
Similar Items
-
MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization
by: Chaubey, Ashutosh, et al.
Published: (2026) -
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
by: Chaubey, Ashutosh, et al.
Published: (2025) -
GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization
by: Jin, Zhangyu, et al.
Published: (2026) -
DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion
by: Siniukov, Maksim, et al.
Published: (2025) -
Revisiting Emotions Representation for Recognition in the Wild
by: Neto, Joao Baptista Cardia, et al.
Published: (2026)