Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Zhou, Yang, Yueyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Graph Embedding with Mel-spectrograms for Underwater Acoustic Target Recognition
von: Feng, Sheng, et al.
Veröffentlicht: (2025)
von: Feng, Sheng, et al.
Veröffentlicht: (2025)
Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues
von: Ba, Zhongjie, et al.
Veröffentlicht: (2026)
von: Ba, Zhongjie, et al.
Veröffentlicht: (2026)
Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
von: Shen, Nanhan, et al.
Veröffentlicht: (2026)
von: Shen, Nanhan, et al.
Veröffentlicht: (2026)
DOA Estimation with Lightweight Network on LLM-Aided Simulated Acoustic Scenes
von: Li, Haowen, et al.
Veröffentlicht: (2025)
von: Li, Haowen, et al.
Veröffentlicht: (2025)
DDSC: Dynamic Dual-Signal Curriculum for Data-Efficient Acoustic Scene Classification under Domain Shift
von: Zhang, Peihong, et al.
Veröffentlicht: (2025)
von: Zhang, Peihong, et al.
Veröffentlicht: (2025)
An Entropy-Guided Curriculum Learning Strategy for Data-Efficient Acoustic Scene Classification under Domain Shift
von: Zhang, Peihong, et al.
Veröffentlicht: (2025)
von: Zhang, Peihong, et al.
Veröffentlicht: (2025)
Towards Explicit Acoustic Evidence Perception in Audio LLMs for Speech Deepfake Detection
von: Guo, Xiaoxuan, et al.
Veröffentlicht: (2026)
von: Guo, Xiaoxuan, et al.
Veröffentlicht: (2026)
AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks
von: Liang, Yun, et al.
Veröffentlicht: (2024)
von: Liang, Yun, et al.
Veröffentlicht: (2024)
AUDRON: A Deep Learning Framework with Fused Acoustic Signatures for Drone Type Recognition
von: Chatterjee, Rajdeep, et al.
Veröffentlicht: (2025)
von: Chatterjee, Rajdeep, et al.
Veröffentlicht: (2025)
The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning
von: Hummel, Hilde I., et al.
Veröffentlicht: (2025)
von: Hummel, Hilde I., et al.
Veröffentlicht: (2025)
TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation
von: Yang, Xiaoda, et al.
Veröffentlicht: (2026)
von: Yang, Xiaoda, et al.
Veröffentlicht: (2026)
BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing
von: Hammami, Hamze, et al.
Veröffentlicht: (2026)
von: Hammami, Hamze, et al.
Veröffentlicht: (2026)
Diffusion-based Surrogate Model for Time-varying Underwater Acoustic Channels
von: Li, Kexin, et al.
Veröffentlicht: (2025)
von: Li, Kexin, et al.
Veröffentlicht: (2025)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems
von: Raymondaud, Quentin, et al.
Veröffentlicht: (2024)
von: Raymondaud, Quentin, et al.
Veröffentlicht: (2024)
DialogGraph-LLM: Graph-Informed LLMs for End-to-End Audio Dialogue Intent Recognition
von: Liu, HongYu, et al.
Veröffentlicht: (2025)
von: Liu, HongYu, et al.
Veröffentlicht: (2025)
Smart Passive Acoustic Monitoring: Embedding a Classifier on AudioMoth Microcontroller
von: Lerbourg, Louis, et al.
Veröffentlicht: (2026)
von: Lerbourg, Louis, et al.
Veröffentlicht: (2026)
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
von: Gao, Zhuoyue, et al.
Veröffentlicht: (2026)
von: Gao, Zhuoyue, et al.
Veröffentlicht: (2026)
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
Khala: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation
von: Liu, Jiafeng, et al.
Veröffentlicht: (2026)
von: Liu, Jiafeng, et al.
Veröffentlicht: (2026)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
von: Dietrich, Juergen
Veröffentlicht: (2026)
von: Dietrich, Juergen
Veröffentlicht: (2026)
Chord Recognition with Deep Learning
von: Mackenzie, Pierre
Veröffentlicht: (2025)
von: Mackenzie, Pierre
Veröffentlicht: (2025)
Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
ChladniSonify: A Visual-Acoustic Mapping Method for Chladni Patterns in New Media Art Creation
von: Liu, Yakun, et al.
Veröffentlicht: (2026)
von: Liu, Yakun, et al.
Veröffentlicht: (2026)
Adaptive Knowledge Distillation using a Device-Aware Teacher for Low-Complexity Acoustic Scene Classification
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2025)
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2025)
Improving Audio Event Recognition with Consistency Regularization
von: Sadhu, Shanmuka, et al.
Veröffentlicht: (2025)
von: Sadhu, Shanmuka, et al.
Veröffentlicht: (2025)
SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation
von: Wang, Hongrui, et al.
Veröffentlicht: (2026)
von: Wang, Hongrui, et al.
Veröffentlicht: (2026)
Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
von: Park, Hansol, et al.
Veröffentlicht: (2025)
von: Park, Hansol, et al.
Veröffentlicht: (2025)
MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
von: Sailor, Hardik B., et al.
Veröffentlicht: (2025)
von: Sailor, Hardik B., et al.
Veröffentlicht: (2025)
AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis
von: Bhosale, Swapnil, et al.
Veröffentlicht: (2024)
von: Bhosale, Swapnil, et al.
Veröffentlicht: (2024)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
RAS: a Reliability Oriented Metric for Automatic Speech Recognition
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
Speech Emotion Recognition via Entropy-Aware Score Selection
von: Chua, ChenYi, et al.
Veröffentlicht: (2025)
von: Chua, ChenYi, et al.
Veröffentlicht: (2025)
A Multi-task Learning Balanced Attention Convolutional Neural Network Model for Few-shot Underwater Acoustic Target Recognition
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language
von: Wiafe, Isaac, et al.
Veröffentlicht: (2026)
von: Wiafe, Isaac, et al.
Veröffentlicht: (2026)
Deep Learning for Speech Emotion Recognition: A CNN Approach Utilizing Mel Spectrograms
von: Penumajji, Niketa
Veröffentlicht: (2025)
von: Penumajji, Niketa
Veröffentlicht: (2025)
Multi-Accent Mandarin Dry-Vocal Singing Dataset: Benchmark for Singing Accent Recognition
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database
von: Xiao, Qing, et al.
Veröffentlicht: (2025)
von: Xiao, Qing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Graph Embedding with Mel-spectrograms for Underwater Acoustic Target Recognition
von: Feng, Sheng, et al.
Veröffentlicht: (2025) -
Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues
von: Ba, Zhongjie, et al.
Veröffentlicht: (2026) -
Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
von: Shen, Nanhan, et al.
Veröffentlicht: (2026) -
DOA Estimation with Lightweight Network on LLM-Aided Simulated Acoustic Scenes
von: Li, Haowen, et al.
Veröffentlicht: (2025) -
DDSC: Dynamic Dual-Signal Curriculum for Data-Efficient Acoustic Scene Classification under Domain Shift
von: Zhang, Peihong, et al.
Veröffentlicht: (2025)