Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mobbs, Rebecca, Makris, Dimitrios, Argyriou, Vasileios |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluation of Environmental Conditions on Object Detection using Oriented Bounding Boxes for AR Applications
von: Li, Vladislav, et al.
Veröffentlicht: (2023)
von: Li, Vladislav, et al.
Veröffentlicht: (2023)
Next-Generation License Plate Detection and Recognition System using YOLOv8
von: Amin, Arslan, et al.
Veröffentlicht: (2025)
von: Amin, Arslan, et al.
Veröffentlicht: (2025)
Reasoning is a Modality
von: Liu, Zhiguang, et al.
Veröffentlicht: (2026)
von: Liu, Zhiguang, et al.
Veröffentlicht: (2026)
Sora as a World Model? A Complete Survey on Text-to-Video Generation
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024)
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024)
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
von: Li, Tao, et al.
Veröffentlicht: (2024)
von: Li, Tao, et al.
Veröffentlicht: (2024)
Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
von: Agarwal, Rachit, et al.
Veröffentlicht: (2026)
von: Agarwal, Rachit, et al.
Veröffentlicht: (2026)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
A Multi-Modal Deep Learning Based Approach for House Price Prediction
von: Hasan, Md Hasebul, et al.
Veröffentlicht: (2024)
von: Hasan, Md Hasebul, et al.
Veröffentlicht: (2024)
Siamese Networks for Cat Re-Identification: Exploring Neural Models for Cat Instance Recognition
von: Trein, Tobias, et al.
Veröffentlicht: (2025)
von: Trein, Tobias, et al.
Veröffentlicht: (2025)
TexTailor: Customized Text-aligned Texturing via Effective Resampling
von: Lee, Suin, et al.
Veröffentlicht: (2025)
von: Lee, Suin, et al.
Veröffentlicht: (2025)
From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models
von: Atighehchian, Parmida, et al.
Veröffentlicht: (2026)
von: Atighehchian, Parmida, et al.
Veröffentlicht: (2026)
Rotation-Adaptive Point Cloud Domain Generalization via Intricate Orientation Learning
von: Liu, Bangzhen, et al.
Veröffentlicht: (2025)
von: Liu, Bangzhen, et al.
Veröffentlicht: (2025)
HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition
von: Chen, Honghui, et al.
Veröffentlicht: (2024)
von: Chen, Honghui, et al.
Veröffentlicht: (2024)
Instruction-based Image Editing with Planning, Reasoning, and Generation
von: Ji, Liya, et al.
Veröffentlicht: (2026)
von: Ji, Liya, et al.
Veröffentlicht: (2026)
CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier
von: Ou, Ziyang
Veröffentlicht: (2025)
von: Ou, Ziyang
Veröffentlicht: (2025)
Deep Learning for automated multi-scale functional field boundaries extraction using multi-date Sentinel-2 and PlanetScope imagery: Case Study of Netherlands and Pakistan
von: Zahid, Saba, et al.
Veröffentlicht: (2024)
von: Zahid, Saba, et al.
Veröffentlicht: (2024)
Supervised Contrastive Learning for Few-Shot AI-Generated Image Detection and Attribution
von: Urueña, Jaime Álvarez, et al.
Veröffentlicht: (2025)
von: Urueña, Jaime Álvarez, et al.
Veröffentlicht: (2025)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
von: Masrourisaadat, Nila, et al.
Veröffentlicht: (2024)
von: Masrourisaadat, Nila, et al.
Veröffentlicht: (2024)
Scalable Face Security Vision Foundation Model for Deepfake, Diffusion, and Spoofing Detection
von: Wang, Gaojian, et al.
Veröffentlicht: (2025)
von: Wang, Gaojian, et al.
Veröffentlicht: (2025)
Region Mixup
von: Saha, Saptarshi, et al.
Veröffentlicht: (2024)
von: Saha, Saptarshi, et al.
Veröffentlicht: (2024)
Neuron-based explanations of neural networks sacrifice completeness and interpretability
von: Dey, Nolan, et al.
Veröffentlicht: (2020)
von: Dey, Nolan, et al.
Veröffentlicht: (2020)
3D Adaptive Structural Convolution Network for Domain-Invariant Point Cloud Recognition
von: Kim, Younggun, et al.
Veröffentlicht: (2024)
von: Kim, Younggun, et al.
Veröffentlicht: (2024)
Enhancing Long-Term Re-Identification Robustness Using Synthetic Data: A Comparative Analysis
von: Pionzewski, Christian, et al.
Veröffentlicht: (2025)
von: Pionzewski, Christian, et al.
Veröffentlicht: (2025)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2024)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2024)
Data Augmentation for Image Classification using Generative AI
von: Rahat, Fazle, et al.
Veröffentlicht: (2024)
von: Rahat, Fazle, et al.
Veröffentlicht: (2024)
SITUATE -- Synthetic Object Counting Dataset for VLM training
von: Peinl, René, et al.
Veröffentlicht: (2026)
von: Peinl, René, et al.
Veröffentlicht: (2026)
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
Image Segmentation and Classification of E-waste for Training Robots for Waste Segregation
von: Tripathi, Prakriti
Veröffentlicht: (2025)
von: Tripathi, Prakriti
Veröffentlicht: (2025)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
von: Holm, Felix, et al.
Veröffentlicht: (2025)
von: Holm, Felix, et al.
Veröffentlicht: (2025)
Appearance-based gaze estimation enhanced with synthetic images using deep neural networks
von: Herashchenko, Dmytro, et al.
Veröffentlicht: (2023)
von: Herashchenko, Dmytro, et al.
Veröffentlicht: (2023)
MFTF: Mask-free Training-free Object Level Layout Control Diffusion Model
von: Yang, Shan
Veröffentlicht: (2024)
von: Yang, Shan
Veröffentlicht: (2024)
Attentive VQ-VAE
von: Hoyos, Angello, et al.
Veröffentlicht: (2023)
von: Hoyos, Angello, et al.
Veröffentlicht: (2023)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective
von: Luo, Wang, et al.
Veröffentlicht: (2025)
von: Luo, Wang, et al.
Veröffentlicht: (2025)
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
von: Safdar, Aon, et al.
Veröffentlicht: (2025)
von: Safdar, Aon, et al.
Veröffentlicht: (2025)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
von: Feng, Yigui, et al.
Veröffentlicht: (2026)
von: Feng, Yigui, et al.
Veröffentlicht: (2026)
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
von: Ma, Jie, et al.
Veröffentlicht: (2023)
von: Ma, Jie, et al.
Veröffentlicht: (2023)
Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based Customization
von: Liu, Yisu, et al.
Veröffentlicht: (2024)
von: Liu, Yisu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Evaluation of Environmental Conditions on Object Detection using Oriented Bounding Boxes for AR Applications
von: Li, Vladislav, et al.
Veröffentlicht: (2023) -
Next-Generation License Plate Detection and Recognition System using YOLOv8
von: Amin, Arslan, et al.
Veröffentlicht: (2025) -
Reasoning is a Modality
von: Liu, Zhiguang, et al.
Veröffentlicht: (2026) -
Sora as a World Model? A Complete Survey on Text-to-Video Generation
von: Puspitasari, Fachrina Dewi, et al.
Veröffentlicht: (2024) -
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
von: Li, Tao, et al.
Veröffentlicht: (2024)