LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Meghanani, Amit, Hain, Thomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2026)
di: Meghanani, Amit, et al.
Pubblicazione: (2026)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
di: Park, Chanho, et al.
Pubblicazione: (2023)
di: Park, Chanho, et al.
Pubblicazione: (2023)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
di: Park, Chanho, et al.
Pubblicazione: (2024)
di: Park, Chanho, et al.
Pubblicazione: (2024)
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
di: Liu, Alexander H., et al.
Pubblicazione: (2024)
di: Liu, Alexander H., et al.
Pubblicazione: (2024)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
di: Chang, Kalvin, et al.
Pubblicazione: (2024)
di: Chang, Kalvin, et al.
Pubblicazione: (2024)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
SpeechAlign: Aligning Speech Generation to Human Preferences
di: Zhang, Dong, et al.
Pubblicazione: (2024)
di: Zhang, Dong, et al.
Pubblicazione: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
Implicit Self-supervised Language Representation for Spoken Language Diarization
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
Task-Agnostic Structured Pruning of Speech Representation Models
di: Wang, Haoyu, et al.
Pubblicazione: (2023)
di: Wang, Haoyu, et al.
Pubblicazione: (2023)
Simultaneous Speech-to-Speech Translation Without Aligned Data
di: Labiausse, Tom, et al.
Pubblicazione: (2026)
di: Labiausse, Tom, et al.
Pubblicazione: (2026)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
di: Farhadipour, Aref, et al.
Pubblicazione: (2023)
di: Farhadipour, Aref, et al.
Pubblicazione: (2023)
Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models
di: Gok, Alican, et al.
Pubblicazione: (2025)
di: Gok, Alican, et al.
Pubblicazione: (2025)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
di: Osakuade, Opeyemi, et al.
Pubblicazione: (2024)
di: Osakuade, Opeyemi, et al.
Pubblicazione: (2024)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
Is Self-Supervised Learning Enough to Fill in the Gap? A Study on Speech Inpainting
di: Asaad, Ihab, et al.
Pubblicazione: (2024)
di: Asaad, Ihab, et al.
Pubblicazione: (2024)
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
di: Du, Yichao, et al.
Pubblicazione: (2024)
di: Du, Yichao, et al.
Pubblicazione: (2024)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
di: Subramanian, Aswin Shanmugam, et al.
Pubblicazione: (2025)
di: Subramanian, Aswin Shanmugam, et al.
Pubblicazione: (2025)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2023)
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2023)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
di: Venkateswaran, Nitin, et al.
Pubblicazione: (2025)
di: Venkateswaran, Nitin, et al.
Pubblicazione: (2025)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
di: Hernandez, Abner, et al.
Pubblicazione: (2026)
di: Hernandez, Abner, et al.
Pubblicazione: (2026)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
di: Han, HyoJung, et al.
Pubblicazione: (2024)
di: Han, HyoJung, et al.
Pubblicazione: (2024)
Self-supervised Speech Models for Word-Level Stuttered Speech Detection
di: Shih, Yi-Jen, et al.
Pubblicazione: (2024)
di: Shih, Yi-Jen, et al.
Pubblicazione: (2024)
FASA: a Flexible and Automatic Speech Aligner for Extracting High-quality Aligned Children Speech Data
di: Liu, Dancheng, et al.
Pubblicazione: (2024)
di: Liu, Dancheng, et al.
Pubblicazione: (2024)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2025)
di: Peng, Junyi, et al.
Pubblicazione: (2025)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
di: Jang, Kangwook, et al.
Pubblicazione: (2023)
di: Jang, Kangwook, et al.
Pubblicazione: (2023)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
di: Close, George, et al.
Pubblicazione: (2024)
di: Close, George, et al.
Pubblicazione: (2024)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
di: Wei, Kun, et al.
Pubblicazione: (2023)
di: Wei, Kun, et al.
Pubblicazione: (2023)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
di: Liu, Alexander H., et al.
Pubblicazione: (2025)
di: Liu, Alexander H., et al.
Pubblicazione: (2025)
LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
di: Parcollet, Titouan, et al.
Pubblicazione: (2023)
di: Parcollet, Titouan, et al.
Pubblicazione: (2023)
Scaling Analysis of Interleaved Speech-Text Language Models
di: Maimon, Gallil, et al.
Pubblicazione: (2025)
di: Maimon, Gallil, et al.
Pubblicazione: (2025)
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
Convexity-based Pruning of Speech Representation Models
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
Configurable Multilingual ASR with Speech Summary Representations
di: Zhu, Harrison, et al.
Pubblicazione: (2024)
di: Zhu, Harrison, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024) -
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024) -
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2026) -
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
di: Park, Chanho, et al.
Pubblicazione: (2023) -
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)