Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ge, Xuri, Fu, Junchen, Chen, Fuhai, An, Shan, Sebe, Nicu, Jose, Joemon M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multimodal Representation Learning Techniques for Comprehensive Facial State Analysis
por: Zheng, Kaiwen, et al.
Publicado: (2025)
por: Zheng, Kaiwen, et al.
Publicado: (2025)
MGRR-Net: Multi-level Graph Relational Reasoning Network for Facial Action Units Detection
por: Ge, Xuri, et al.
Publicado: (2022)
por: Ge, Xuri, et al.
Publicado: (2022)
3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting
por: Ge, Xuri, et al.
Publicado: (2024)
por: Ge, Xuri, et al.
Publicado: (2024)
Rethinking the Learning Paradigm for Facial Expression Recognition
por: Wang, Weijie, et al.
Publicado: (2022)
por: Wang, Weijie, et al.
Publicado: (2022)
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching
por: Ge, Xuri, et al.
Publicado: (2024)
por: Ge, Xuri, et al.
Publicado: (2024)
Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion
por: Zhang, Jiaru, et al.
Publicado: (2026)
por: Zhang, Jiaru, et al.
Publicado: (2026)
Focal-RegionFace: Generating Fine-Grained Multi-attribute Descriptions for Arbitrarily Selected Face Focal Regions
por: Zheng, Kaiwen, et al.
Publicado: (2026)
por: Zheng, Kaiwen, et al.
Publicado: (2026)
Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
por: Shi, Tong, et al.
Publicado: (2024)
por: Shi, Tong, et al.
Publicado: (2024)
An Effective End-to-End Solution for Multimodal Action Recognition
por: Wang, Songping, et al.
Publicado: (2025)
por: Wang, Songping, et al.
Publicado: (2025)
IISAN: Efficiently Adapting Multimodal Representation for Sequential Recommendation with Decoupled PEFT
por: Fu, Junchen, et al.
Publicado: (2024)
por: Fu, Junchen, et al.
Publicado: (2024)
Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues
por: Fu, Junchen, et al.
Publicado: (2026)
por: Fu, Junchen, et al.
Publicado: (2026)
Action Unit Enhance Dynamic Facial Expression Recognition
por: Liu, Feng, et al.
Publicado: (2025)
por: Liu, Feng, et al.
Publicado: (2025)
Towards Unified Facial Action Unit Recognition Framework by Large Language Models
por: Hu, Guohong, et al.
Publicado: (2024)
por: Hu, Guohong, et al.
Publicado: (2024)
CFIR: Fast and Effective Long-Text To Image Retrieval for Large Corpora
por: Long, Zijun, et al.
Publicado: (2024)
por: Long, Zijun, et al.
Publicado: (2024)
Efficient and Effective Adaptation of Multimodal Foundation Models in Sequential Recommendation
por: Fu, Junchen, et al.
Publicado: (2024)
por: Fu, Junchen, et al.
Publicado: (2024)
LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation
por: Fu, Junchen, et al.
Publicado: (2025)
por: Fu, Junchen, et al.
Publicado: (2025)
LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving
por: Song, Nan, et al.
Publicado: (2025)
por: Song, Nan, et al.
Publicado: (2025)
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
por: Yu, Haiyang, et al.
Publicado: (2025)
por: Yu, Haiyang, et al.
Publicado: (2025)
Hierarchical Vision-Language Interaction for Facial Action Unit Detection
por: Li, Yong, et al.
Publicado: (2026)
por: Li, Yong, et al.
Publicado: (2026)
Contrastive Learning of Person-independent Representations for Facial Action Unit Detection
por: Li, Yong, et al.
Publicado: (2024)
por: Li, Yong, et al.
Publicado: (2024)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
por: Fu, Haoyu, et al.
Publicado: (2025)
por: Fu, Haoyu, et al.
Publicado: (2025)
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
por: Zhou, Jiaming, et al.
Publicado: (2023)
por: Zhou, Jiaming, et al.
Publicado: (2023)
Vision+X: A Survey on Multimodal Learning in the Light of Data
por: Zhu, Ye, et al.
Publicado: (2022)
por: Zhu, Ye, et al.
Publicado: (2022)
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
por: Zhong, Yufeng, et al.
Publicado: (2026)
por: Zhong, Yufeng, et al.
Publicado: (2026)
Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving
por: Sun, Jiangxin, et al.
Publicado: (2026)
por: Sun, Jiangxin, et al.
Publicado: (2026)
Towards Localized Fine-Grained Control for Facial Expression Generation
por: Varanka, Tuomas, et al.
Publicado: (2024)
por: Varanka, Tuomas, et al.
Publicado: (2024)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
por: Cheng, Hao, et al.
Publicado: (2024)
por: Cheng, Hao, et al.
Publicado: (2024)
Guided Interpretable Facial Expression Recognition via Spatial Action Unit Cues
por: Belharbi, Soufiane, et al.
Publicado: (2024)
por: Belharbi, Soufiane, et al.
Publicado: (2024)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
por: Li, Jinlong, et al.
Publicado: (2024)
por: Li, Jinlong, et al.
Publicado: (2024)
High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained Expressions
por: He, Yikang, et al.
Publicado: (2026)
por: He, Yikang, et al.
Publicado: (2026)
Causal Intervention for Subject-Deconfounded Facial Action Unit Recognition
por: Chen, Yingjie, et al.
Publicado: (2022)
por: Chen, Yingjie, et al.
Publicado: (2022)
End-to-End Facial Expression Detection in Long Videos
por: Fang, Yini, et al.
Publicado: (2025)
por: Fang, Yini, et al.
Publicado: (2025)
VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
por: Seong, Hyunki, et al.
Publicado: (2025)
por: Seong, Hyunki, et al.
Publicado: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
por: Li, Jinlong, et al.
Publicado: (2026)
por: Li, Jinlong, et al.
Publicado: (2026)
End-to-End Chess Recognition
por: Masouris, Athanasios, et al.
Publicado: (2023)
por: Masouris, Athanasios, et al.
Publicado: (2023)
An End-to-End Two-Stream Network Based on RGB Flow and Representation Flow for Human Action Recognition
por: Lai, Song-Jiang, et al.
Publicado: (2024)
por: Lai, Song-Jiang, et al.
Publicado: (2024)
Democratizing Fine-grained Visual Recognition with Large Language Models
por: Liu, Mingxuan, et al.
Publicado: (2024)
por: Liu, Mingxuan, et al.
Publicado: (2024)
GeoConv: Geodesic Guided Convolution for Facial Action Unit Recognition
por: Chen, Yuedong, et al.
Publicado: (2020)
por: Chen, Yuedong, et al.
Publicado: (2020)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
por: Zhou, Xingcheng, et al.
Publicado: (2025)
por: Zhou, Xingcheng, et al.
Publicado: (2025)
Action Images: End-to-End Policy Learning via Multiview Video Generation
por: Zhen, Haoyu, et al.
Publicado: (2026)
por: Zhen, Haoyu, et al.
Publicado: (2026)
Ejemplares similares
-
Multimodal Representation Learning Techniques for Comprehensive Facial State Analysis
por: Zheng, Kaiwen, et al.
Publicado: (2025) -
MGRR-Net: Multi-level Graph Relational Reasoning Network for Facial Action Units Detection
por: Ge, Xuri, et al.
Publicado: (2022) -
3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting
por: Ge, Xuri, et al.
Publicado: (2024) -
Rethinking the Learning Paradigm for Facial Expression Recognition
por: Wang, Weijie, et al.
Publicado: (2022) -
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching
por: Ge, Xuri, et al.
Publicado: (2024)