Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Phukan, Orchid Chetia, Girish, Akhtar, Mohd Mujtaba, Behera, Swarup Ranjan, Mallick, Priyabrata, Reddy, Pailla Balakrishna, Buduru, Arun Balaji, Sharma, Rajesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Towards Multilingual Audio-Visual Question Answering
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Rethinking Cross-Corpus Speech Emotion Recognition Benchmarking: Are Paralinguistic Pre-Trained Representations Sufficient?
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations
von: Girish, et al.
Veröffentlicht: (2025)
von: Girish, et al.
Veröffentlicht: (2025)
Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages
von: Girish, et al.
Veröffentlicht: (2026)
von: Girish, et al.
Veröffentlicht: (2026)
Are Multimodal Foundation Models All That Is Needed for Emofake Detection?
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2025)
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2025)
Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Towards Fusion of Neural Audio Codec-based Representations with Spectral for Heart Murmur Classification via Bandit-based Cross-Attention Mechanism
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Towards Neural Audio Codec Source Parsing
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2025)
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2025)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Towards Machine Unlearning for Paralinguistic Speech Processing
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System
von: Choudhury, Nitin, et al.
Veröffentlicht: (2026)
von: Choudhury, Nitin, et al.
Veröffentlicht: (2026)
SONIC: Synergizing VisiON Foundation Models for Stress RecogNItion from ECG signals
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Modality-Order Matters! A Novel Hierarchical Feature Fusion Method for CoSAm: A Code-Switched Autism Corpus
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2024)
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2024)
NeuRO: An Application for Code-Switched Autism Detection in Children
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2024)
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2024)
The Reasonable Effectiveness of Speaker Embeddings for Violence Detection
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
CoLLAB: A Collaborative Approach for Multilingual Abuse Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
AVR: Synergizing Foundation Models for Audio-Visual Humor Detection
von: Sharma, Sarthak, et al.
Veröffentlicht: (2024)
von: Sharma, Sarthak, et al.
Veröffentlicht: (2024)
AQUALLM: Audio Question Answering Data Generation Using Large Language Models
von: Behera, Swarup Ranjan, et al.
Veröffentlicht: (2023)
von: Behera, Swarup Ranjan, et al.
Veröffentlicht: (2023)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
von: Koshal, Devyani, et al.
Veröffentlicht: (2024)
von: Koshal, Devyani, et al.
Veröffentlicht: (2024)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
A Lightweight Feature Fusion Architecture For Resource-Constrained Crowd Counting
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
von: Zhang, You, et al.
Veröffentlicht: (2024)
von: Zhang, You, et al.
Veröffentlicht: (2024)
ComFeAT: Combination of Neural and Spectral Features for Improved Depression Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
VoxMed: One-Step Respiratory Disease Classifier using Digital Stethoscope Sounds
von: Mundra, Paridhi, et al.
Veröffentlicht: (2024)
von: Mundra, Paridhi, et al.
Veröffentlicht: (2024)
FOCA: Multimodal Malware Classification via Hyperbolic Cross-Attention
von: Choudhury, Nitin, et al.
Veröffentlicht: (2026)
von: Choudhury, Nitin, et al.
Veröffentlicht: (2026)
ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025) -
SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025) -
Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025) -
HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025) -
Towards Multilingual Audio-Visual Question Answering
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)