Multimodal Representation Learning Techniques for Comprehensive Facial State Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Kaiwen, Ge, Xuri, Fu, Junchen, Peng, Jun, Jose, Joemon M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
Focal-RegionFace: Generating Fine-Grained Multi-attribute Descriptions for Arbitrarily Selected Face Focal Regions
by: Zheng, Kaiwen, et al.
Published: (2026)
by: Zheng, Kaiwen, et al.
Published: (2026)
Efficient and Effective Adaptation of Multimodal Foundation Models in Sequential Recommendation
by: Fu, Junchen, et al.
Published: (2024)
by: Fu, Junchen, et al.
Published: (2024)
IISAN: Efficiently Adapting Multimodal Representation for Sequential Recommendation with Decoupled PEFT
by: Fu, Junchen, et al.
Published: (2024)
by: Fu, Junchen, et al.
Published: (2024)
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
by: Fu, Junchen, et al.
Published: (2026)
by: Fu, Junchen, et al.
Published: (2026)
LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation
by: Fu, Junchen, et al.
Published: (2025)
by: Fu, Junchen, et al.
Published: (2025)
MGRR-Net: Multi-level Graph Relational Reasoning Network for Facial Action Units Detection
by: Ge, Xuri, et al.
Published: (2022)
by: Ge, Xuri, et al.
Published: (2022)
Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
by: Shi, Tong, et al.
Published: (2024)
by: Shi, Tong, et al.
Published: (2024)
CFIR: Fast and Effective Long-Text To Image Retrieval for Large Corpora
by: Long, Zijun, et al.
Published: (2024)
by: Long, Zijun, et al.
Published: (2024)
3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
Self-Supervised Facial Representation Learning with Facial Region Awareness
by: Gao, Zheng, et al.
Published: (2024)
by: Gao, Zheng, et al.
Published: (2024)
Face-GPS: A Comprehensive Technique for Quantifying Facial Muscle Dynamics in Videos
by: Kim, Juni, et al.
Published: (2024)
by: Kim, Juni, et al.
Published: (2024)
A Generalist FaceX via Learning Unified Facial Representation
by: Han, Yue, et al.
Published: (2023)
by: Han, Yue, et al.
Published: (2023)
Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions
by: Sun, Licai, et al.
Published: (2025)
by: Sun, Licai, et al.
Published: (2025)
HpEIS: Learning Hand Pose Embeddings for Multimedia Interactive Systems
by: Xu, Songpei, et al.
Published: (2024)
by: Xu, Songpei, et al.
Published: (2024)
A Survey on Deep Learning for Polyp Segmentation: Techniques, Challenges and Future Trends
by: Mei, Jiaxin, et al.
Published: (2023)
by: Mei, Jiaxin, et al.
Published: (2023)
EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction
by: Ge, Chengjie, et al.
Published: (2025)
by: Ge, Chengjie, et al.
Published: (2025)
Optical Flow Techniques for Facial Expression Analysis -- a Practical Evaluation Study
by: Allaert, Benjamin, et al.
Published: (2019)
by: Allaert, Benjamin, et al.
Published: (2019)
Improved Techniques for GAN based Facial Inpainting
by: Lahiri, Avisek, et al.
Published: (2018)
by: Lahiri, Avisek, et al.
Published: (2018)
Learning Spatially Decoupled Color Representations for Facial Image Colorization
by: Zhu, Hangyan, et al.
Published: (2024)
by: Zhu, Hangyan, et al.
Published: (2024)
Representation Learning and Identity Adversarial Training for Facial Behavior Understanding
by: Ning, Mang, et al.
Published: (2024)
by: Ning, Mang, et al.
Published: (2024)
A Generative Framework for Self-Supervised Facial Representation Learning
by: He, Ruian, et al.
Published: (2023)
by: He, Ruian, et al.
Published: (2023)
Multimodal Engagement Analysis from Facial Videos in the Classroom
by: Sümer, Ömer, et al.
Published: (2021)
by: Sümer, Ömer, et al.
Published: (2021)
Revisiting Face Forgery Detection: From Facial Representation to Forgery Detection
by: Guo, Zonghui, et al.
Published: (2024)
by: Guo, Zonghui, et al.
Published: (2024)
CREM: Compression-Driven Representation Enhancement for Multimodal Retrieval and Comprehension
by: Liu, Lihao, et al.
Published: (2026)
by: Liu, Lihao, et al.
Published: (2026)
Learning Semantic Facial Descriptors for Accurate Face Animation
by: Zhu, Lei, et al.
Published: (2025)
by: Zhu, Lei, et al.
Published: (2025)
Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities
by: Bao, Jindi, et al.
Published: (2026)
by: Bao, Jindi, et al.
Published: (2026)
A survey on Graph Deep Representation Learning for Facial Expression Recognition
by: Gueuret, Théo, et al.
Published: (2024)
by: Gueuret, Théo, et al.
Published: (2024)
Contrastive Learning of Person-independent Representations for Facial Action Unit Detection
by: Li, Yong, et al.
Published: (2024)
by: Li, Yong, et al.
Published: (2024)
MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image Retrieval
by: Ge, Xuri, et al.
Published: (2026)
by: Ge, Xuri, et al.
Published: (2026)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
by: Ge, Yuying, et al.
Published: (2024)
by: Ge, Yuying, et al.
Published: (2024)
15M Multimodal Facial Image-Text Dataset
by: Dai, Dawei, et al.
Published: (2024)
by: Dai, Dawei, et al.
Published: (2024)
Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
by: Hu, Jiewen, et al.
Published: (2025)
by: Hu, Jiewen, et al.
Published: (2025)
M-ErasureBench: A Comprehensive Multimodal Evaluation Benchmark for Concept Erasure in Diffusion Models
by: Weng, Ju-Hsuan, et al.
Published: (2025)
by: Weng, Ju-Hsuan, et al.
Published: (2025)
DrFER: Learning Disentangled Representations for 3D Facial Expression Recognition
by: Li, Hebeizi, et al.
Published: (2024)
by: Li, Hebeizi, et al.
Published: (2024)
Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition
by: Liang, Zeyu, et al.
Published: (2025)
by: Liang, Zeyu, et al.
Published: (2025)
Promptable Representation Distribution Learning and Data Augmentation for Gigapixel Histopathology WSI Analysis
by: Tang, Kunming, et al.
Published: (2024)
by: Tang, Kunming, et al.
Published: (2024)
Analysis of Bias in Deep Learning Facial Beauty Regressors
by: Hamel, Chandon, et al.
Published: (2025)
by: Hamel, Chandon, et al.
Published: (2025)
Similar Items
-
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
by: Ge, Xuri, et al.
Published: (2024) -
Focal-RegionFace: Generating Fine-Grained Multi-attribute Descriptions for Arbitrarily Selected Face Focal Regions
by: Zheng, Kaiwen, et al.
Published: (2026) -
Efficient and Effective Adaptation of Multimodal Foundation Models in Sequential Recommendation
by: Fu, Junchen, et al.
Published: (2024) -
IISAN: Efficiently Adapting Multimodal Representation for Sequential Recommendation with Decoupled PEFT
by: Fu, Junchen, et al.
Published: (2024) -
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
by: Fu, Junchen, et al.
Published: (2026)