UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zheng, Wang, Shuo, Wang, Junhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-user activity recognition using deep domain adaptation with temporal relation information
by: Ye, Xiaozhou, et al.
Published: (2024)
by: Ye, Xiaozhou, et al.
Published: (2024)
Accurate online action and gesture recognition system using detectors and Deep SPD Siamese Networks
by: Akremi, Mohamed Sanim, et al.
Published: (2025)
by: Akremi, Mohamed Sanim, et al.
Published: (2025)
Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
by: Kang, Wan Ju, et al.
Published: (2025)
by: Kang, Wan Ju, et al.
Published: (2025)
Cross-user activity recognition via temporal relation optimal transport
by: Ye, Xiaozhou, et al.
Published: (2024)
by: Ye, Xiaozhou, et al.
Published: (2024)
Multi-face emotion detection for effective Human-Robot Interaction
by: Yahyaoui, Mohamed Ala, et al.
Published: (2025)
by: Yahyaoui, Mohamed Ala, et al.
Published: (2025)
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
by: Zhao, Haoyu, et al.
Published: (2025)
by: Zhao, Haoyu, et al.
Published: (2025)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
Milmer: a Framework for Multiple Instance Learning based Multimodal Emotion Recognition
by: Wang, Zaitian, et al.
Published: (2025)
by: Wang, Zaitian, et al.
Published: (2025)
LCE: A Framework for Explainability of DNNs for Ultrasound Image Based on Concept Discovery
by: Kong, Weiji, et al.
Published: (2024)
by: Kong, Weiji, et al.
Published: (2024)
ScreenAgent: A Vision Language Model-driven Computer Control Agent
by: Niu, Runliang, et al.
Published: (2024)
by: Niu, Runliang, et al.
Published: (2024)
Do Vision Language Models Understand Human Engagement in Games?
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
by: Zheng, Kaizhi, et al.
Published: (2024)
by: Zheng, Kaizhi, et al.
Published: (2024)
HandS3C: 3D Hand Mesh Reconstruction with State Space Spatial Channel Attention from RGB images
by: Jiao, Zixun, et al.
Published: (2024)
by: Jiao, Zixun, et al.
Published: (2024)
OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
by: Luo, Cheng, et al.
Published: (2025)
by: Luo, Cheng, et al.
Published: (2025)
SASG-DA: Sparse-Aware Semantic-Guided Diffusion Augmentation For Myoelectric Gesture Recognition
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
Exploring Gaze Pattern Differences Between Autistic and Neurotypical Children: Clustering, Visualisation, and Prediction
by: Shi, Weiyan, et al.
Published: (2024)
by: Shi, Weiyan, et al.
Published: (2024)
Evaluating multimodal emotion recognition in proactive conversational agents: A user study
by: Dragut, Adnana, et al.
Published: (2026)
by: Dragut, Adnana, et al.
Published: (2026)
Code2World: A GUI World Model via Renderable Code Generation
by: Zheng, Yuhao, et al.
Published: (2026)
by: Zheng, Yuhao, et al.
Published: (2026)
Modeling Subjective Urban Perception with Human Gaze
by: Che, Lin, et al.
Published: (2026)
by: Che, Lin, et al.
Published: (2026)
Triple Spectral Fusion for Sensor-based Human Activity Recognition
by: Zhang, Ye, et al.
Published: (2026)
by: Zhang, Ye, et al.
Published: (2026)
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
by: Xie, Tianbao, et al.
Published: (2025)
by: Xie, Tianbao, et al.
Published: (2025)
DiffMesh: A Motion-aware Diffusion Framework for Human Mesh Recovery from Videos
by: Zheng, Ce, et al.
Published: (2023)
by: Zheng, Ce, et al.
Published: (2023)
Adaptive 3D UI Placement in Mixed Reality Using Deep Reinforcement Learning
by: Lu, Feiyu, et al.
Published: (2025)
by: Lu, Feiyu, et al.
Published: (2025)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
by: Hu, Siyuan, et al.
Published: (2025)
by: Hu, Siyuan, et al.
Published: (2025)
Generative Augmented Reality: Paradigms, Technologies, and Future Applications
by: Liang, Chen, et al.
Published: (2025)
by: Liang, Chen, et al.
Published: (2025)
Human-like object concept representations emerge naturally in multimodal large language models
by: Du, Changde, et al.
Published: (2024)
by: Du, Changde, et al.
Published: (2024)
How good are humans at detecting AI-generated images? Learnings from an experiment
by: Roca, Thomas, et al.
Published: (2025)
by: Roca, Thomas, et al.
Published: (2025)
WebAccessVL: Violation-Aware VLM for Web Accessibility
by: Zheng, Amber Yijia, et al.
Published: (2025)
by: Zheng, Amber Yijia, et al.
Published: (2025)
VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and Analysis
by: Yan, Xinyuan, et al.
Published: (2025)
by: Yan, Xinyuan, et al.
Published: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
ViEEG: Hierarchical Visual Neural Representation for EEG Brain Decoding
by: Liu, Minxu, et al.
Published: (2025)
by: Liu, Minxu, et al.
Published: (2025)
From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
by: Zhao, Yiming, et al.
Published: (2026)
by: Zhao, Yiming, et al.
Published: (2026)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
by: Ouyang, Mingyu, et al.
Published: (2026)
by: Ouyang, Mingyu, et al.
Published: (2026)
Achieving Effective Virtual Reality Interactions via Acoustic Gesture Recognition based on Large Language Models
by: Zhang, Xijie, et al.
Published: (2025)
by: Zhang, Xijie, et al.
Published: (2025)
K-Sort Arena: Efficient and Reliable Benchmarking for Generative Models via K-wise Human Preferences
by: Li, Zhikai, et al.
Published: (2024)
by: Li, Zhikai, et al.
Published: (2024)
Decomposing and Fusing Intra- and Inter-Sensor Spatio-Temporal Signal for Multi-Sensor Wearable Human Activity Recognition
by: Xie, Haoyu, et al.
Published: (2025)
by: Xie, Haoyu, et al.
Published: (2025)
Learning High-Quality Navigation and Zooming on Omnidirectional Images in Virtual Reality
by: Cao, Zidong, et al.
Published: (2024)
by: Cao, Zidong, et al.
Published: (2024)
From Image Generation to Infrastructure Design: a Multi-agent Pipeline for Street Design Generation
by: Wang, Chenguang, et al.
Published: (2025)
by: Wang, Chenguang, et al.
Published: (2025)
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
by: Wang, Kohou, et al.
Published: (2025)
by: Wang, Kohou, et al.
Published: (2025)
Similar Items
-
Cross-user activity recognition using deep domain adaptation with temporal relation information
by: Ye, Xiaozhou, et al.
Published: (2024) -
Accurate online action and gesture recognition system using detectors and Deep SPD Siamese Networks
by: Akremi, Mohamed Sanim, et al.
Published: (2025) -
Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
by: Kang, Wan Ju, et al.
Published: (2025) -
Cross-user activity recognition via temporal relation optimal transport
by: Ye, Xiaozhou, et al.
Published: (2024) -
Multi-face emotion detection for effective Human-Robot Interaction
by: Yahyaoui, Mohamed Ala, et al.
Published: (2025)