SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Pandey, Prabhat, Swaminathan, Rupak Vignesh, Girish, K V Vijay, Sen, Arunasish, Xie, Jian, Strimel, Grant P., Schwarz, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VCNAC: A Variable-Channel Neural Audio Codec for Mono, Stereo, and Surround Sound
by: Grötschla, Florian, et al.
Published: (2026)
by: Grötschla, Florian, et al.
Published: (2026)
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
by: Girish, et al.
Published: (2026)
by: Girish, et al.
Published: (2026)
Streaming Speech-to-Confusion Network Speech Recognition
by: Filimonov, Denis, et al.
Published: (2023)
by: Filimonov, Denis, et al.
Published: (2023)
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
by: He, Haorui, et al.
Published: (2024)
by: He, Haorui, et al.
Published: (2024)
Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models
by: Phukan, Orchid Chetia, et al.
Published: (2025)
by: Phukan, Orchid Chetia, et al.
Published: (2025)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
by: Pandey, Rahul, et al.
Published: (2023)
by: Pandey, Rahul, et al.
Published: (2023)
Promptformer: Prompted Conformer Transducer for ASR
by: Duarte-Torres, Sergio, et al.
Published: (2024)
by: Duarte-Torres, Sergio, et al.
Published: (2024)
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
by: Huang, Wen, et al.
Published: (2025)
by: Huang, Wen, et al.
Published: (2025)
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
by: Guo, Xin, et al.
Published: (2026)
by: Guo, Xin, et al.
Published: (2026)
ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark
by: Chen, Yun, et al.
Published: (2025)
by: Chen, Yun, et al.
Published: (2025)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
by: Chen, Huakang, et al.
Published: (2026)
by: Chen, Huakang, et al.
Published: (2026)
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
by: He, Haorui, et al.
Published: (2025)
by: He, Haorui, et al.
Published: (2025)
DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation
by: Male, Prabash Reddy, et al.
Published: (2025)
by: Male, Prabash Reddy, et al.
Published: (2025)
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
by: Lee, Beomseok, et al.
Published: (2024)
by: Lee, Beomseok, et al.
Published: (2024)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
by: Zhou, Yixuan, et al.
Published: (2024)
by: Zhou, Yixuan, et al.
Published: (2024)
Multilingual Speech Recognition Using Discrete Tokens with a Two-step Training Strategy
by: Li, Zehan, et al.
Published: (2025)
by: Li, Zehan, et al.
Published: (2025)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
by: Mujtaba, Dena, et al.
Published: (2025)
by: Mujtaba, Dena, et al.
Published: (2025)
Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
by: Mu, Bingshen, et al.
Published: (2025)
by: Mu, Bingshen, et al.
Published: (2025)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
by: Wagner, Dominik, et al.
Published: (2025)
by: Wagner, Dominik, et al.
Published: (2025)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
by: Zheng, Xiuwen, et al.
Published: (2024)
by: Zheng, Xiuwen, et al.
Published: (2024)
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
by: Zhang, Leying, et al.
Published: (2024)
by: Zhang, Leying, et al.
Published: (2024)
Learning Fine-Grained Controllability on Speech Generation via Efficient Fine-Tuning
by: Chien, Chung-Ming, et al.
Published: (2024)
by: Chien, Chung-Ming, et al.
Published: (2024)
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
by: Garg, Ashi, et al.
Published: (2025)
by: Garg, Ashi, et al.
Published: (2025)
Multi-Iteration Multi-Stage Fine-Tuning of Transformers for Sound Event Detection with Heterogeneous Datasets
by: Schmid, Florian, et al.
Published: (2024)
by: Schmid, Florian, et al.
Published: (2024)
Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
by: Nakata, Wataru, et al.
Published: (2025)
by: Nakata, Wataru, et al.
Published: (2025)
Towards Attribution of Generators and Emotional Manipulation in Cross-Lingual Synthetic Speech using Geometric Learning
by: Girish, et al.
Published: (2025)
by: Girish, et al.
Published: (2025)
Bridging Attribution and Open-Set Detection using Graph-Augmented Instance Learning in Synthetic Speech
by: Akhtar, Mohd Mujtaba, et al.
Published: (2026)
by: Akhtar, Mohd Mujtaba, et al.
Published: (2026)
EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
by: Park, Joonyong, et al.
Published: (2025)
by: Park, Joonyong, et al.
Published: (2025)
Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
by: Vecino, Biel Tura, et al.
Published: (2025)
by: Vecino, Biel Tura, et al.
Published: (2025)
Curved Worlds, Clear Boundaries: Generalizing Speech Deepfake Detection using Hyperbolic and Spherical Geometry Spaces
by: Sheth, Farhan, et al.
Published: (2025)
by: Sheth, Farhan, et al.
Published: (2025)
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
by: Langman, Ryan, et al.
Published: (2025)
by: Langman, Ryan, et al.
Published: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
by: Qi, Tianhua, et al.
Published: (2026)
by: Qi, Tianhua, et al.
Published: (2026)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
by: Denisov, Pavel, et al.
Published: (2024)
by: Denisov, Pavel, et al.
Published: (2024)
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
Benchmarking Humans and Machines on Complex Multilingual Speech Understanding Tasks
by: Kankanala, Sai Samrat, et al.
Published: (2025)
by: Kankanala, Sai Samrat, et al.
Published: (2025)
Open-Source System for Multilingual Translation and Cloned Speech Synthesis
by: Cámara, Mateo, et al.
Published: (2025)
by: Cámara, Mateo, et al.
Published: (2025)
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
by: Lu, Ke-Han, et al.
Published: (2024)
by: Lu, Ke-Han, et al.
Published: (2024)
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
by: de Groot, Dimme, et al.
Published: (2026)
by: de Groot, Dimme, et al.
Published: (2026)
Parameter-Efficient Fine-Tuning of Foundation Models for CLP Speech Classification
by: Bhattacharjee, Susmita, et al.
Published: (2025)
by: Bhattacharjee, Susmita, et al.
Published: (2025)
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
by: Jing, Ruihao, et al.
Published: (2025)
by: Jing, Ruihao, et al.
Published: (2025)
Similar Items
-
VCNAC: A Variable-Channel Neural Audio Codec for Mono, Stereo, and Surround Sound
by: Grötschla, Florian, et al.
Published: (2026) -
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
by: Girish, et al.
Published: (2026) -
Streaming Speech-to-Confusion Network Speech Recognition
by: Filimonov, Denis, et al.
Published: (2023) -
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
by: He, Haorui, et al.
Published: (2024) -
Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models
by: Phukan, Orchid Chetia, et al.
Published: (2025)