Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Salehi, Pegah, Sheshkal, Sajad Amouei, Thambawita, Vajira, Gautam, Sushant, Sabet, Saeed S., Johansen, Dag, Riegler, Michael A., Halvorsen, Pål |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
by: Salehi, Pegah, et al.
Published: (2025)
by: Salehi, Pegah, et al.
Published: (2025)
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
by: Gautam, Sushant, et al.
Published: (2026)
by: Gautam, Sushant, et al.
Published: (2026)
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
by: Niizumi, Daisuke, et al.
Published: (2024)
by: Niizumi, Daisuke, et al.
Published: (2024)
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection
by: Niizumi, Daisuke, et al.
Published: (2024)
by: Niizumi, Daisuke, et al.
Published: (2024)
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
by: Niizumi, Daisuke, et al.
Published: (2024)
by: Niizumi, Daisuke, et al.
Published: (2024)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
by: Zheng, Zihao, et al.
Published: (2025)
by: Zheng, Zihao, et al.
Published: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
FakeSound: Deepfake General Audio Detection
by: Xie, Zeyu, et al.
Published: (2024)
by: Xie, Zeyu, et al.
Published: (2024)
STAR: Speech-to-Audio Generation via Representation Learning
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
Towards Pre-training an Effective Respiratory Audio Foundation Model
by: Niizumi, Daisuke, et al.
Published: (2025)
by: Niizumi, Daisuke, et al.
Published: (2025)
Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis
by: Niizumi, Daisuke, et al.
Published: (2025)
by: Niizumi, Daisuke, et al.
Published: (2025)
M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP
by: Niizumi, Daisuke, et al.
Published: (2025)
by: Niizumi, Daisuke, et al.
Published: (2025)
A Speech Enhancement Method Using Fast Fourier Transform and Convolutional Autoencoder
by: Kow, Pu-Yun, et al.
Published: (2025)
by: Kow, Pu-Yun, et al.
Published: (2025)
ML-ASPA: A Contemplation of Machine Learning-based Acoustic Signal Processing Analysis for Sounds, & Strains Emerging Technology
by: Ali, Ratul, et al.
Published: (2023)
by: Ali, Ratul, et al.
Published: (2023)
Can Sound Replace Vision in LLaVA With Token Substitution?
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
BowelRCNN: Region-based Convolutional Neural Network System for Bowel Sound Auscultation
by: Matynia, Igor, et al.
Published: (2025)
by: Matynia, Igor, et al.
Published: (2025)
Spectral oversubtraction? An approach for speech enhancement after robot ego speech filtering in semi-real-time
by: Li, Yue, et al.
Published: (2024)
by: Li, Yue, et al.
Published: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
by: Bovbjerg, Holger Severin, et al.
Published: (2025)
by: Bovbjerg, Holger Severin, et al.
Published: (2025)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
SARA: Stress Test Reasoning in Audio Deepfake Detection
by: Nguyen, Binh, et al.
Published: (2026)
by: Nguyen, Binh, et al.
Published: (2026)
Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and Localization
by: Wu, Junyan, et al.
Published: (2024)
by: Wu, Junyan, et al.
Published: (2024)
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
by: Niizumi, Daisuke, et al.
Published: (2026)
by: Niizumi, Daisuke, et al.
Published: (2026)
MATER: Multi-level Acoustic and Textual Emotion Representation for Interpretable Speech Emotion Recognition
by: Jon, Hyo Jin, et al.
Published: (2025)
by: Jon, Hyo Jin, et al.
Published: (2025)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
by: Xie, Zeyu, et al.
Published: (2025)
by: Xie, Zeyu, et al.
Published: (2025)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
by: Guichoux, Téo, et al.
Published: (2025)
by: Guichoux, Téo, et al.
Published: (2025)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
Sound Safeguarding for Acoustic Measurement Using Any Sounds: Tools and Applications
by: Kawahara, Hideki, et al.
Published: (2025)
by: Kawahara, Hideki, et al.
Published: (2025)
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Non-Verbal Vocalisations and their Challenges: Emotion, Privacy, Sparseness, and Real Life
by: Batliner, Anton, et al.
Published: (2025)
by: Batliner, Anton, et al.
Published: (2025)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
by: Cheripally, Sowmya
Published: (2024)
by: Cheripally, Sowmya
Published: (2024)
FastAST: Accelerating Audio Spectrogram Transformer via Token Merging and Cross-Model Knowledge Distillation
by: Behera, Swarup Ranjan, et al.
Published: (2024)
by: Behera, Swarup Ranjan, et al.
Published: (2024)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
Similar Items
-
Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
by: Salehi, Pegah, et al.
Published: (2025) -
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
by: Gautam, Sushant, et al.
Published: (2026) -
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
by: Gautam, Sushant, et al.
Published: (2025) -
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
by: Gautam, Sushant, et al.
Published: (2025) -
Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model
by: Vosoughi, Ali, et al.
Published: (2025)