RankUp: Boosting Semi-Supervised Regression with an Auxiliary Ranking Classifier
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Pin-Yen, Fu, Szu-Wei, Tsao, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
di: Fernandez-Lopez, Adriana, et al.
Pubblicazione: (2024)
di: Fernandez-Lopez, Adriana, et al.
Pubblicazione: (2024)
Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025)
ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
di: Gupta, Pranav, et al.
Pubblicazione: (2024)
di: Gupta, Pranav, et al.
Pubblicazione: (2024)
Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization
di: Guo, Yuxin, et al.
Pubblicazione: (2024)
di: Guo, Yuxin, et al.
Pubblicazione: (2024)
Universal Speech Enhancement with Regression and Generative Mamba
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
ESARM: 3D Emotional Speech-to-Animation via Reward Model from Automatically-Ranked Demonstrations
di: Zhang, Xulong, et al.
Pubblicazione: (2024)
di: Zhang, Xulong, et al.
Pubblicazione: (2024)
How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
di: Lu, Ke-Han, et al.
Pubblicazione: (2026)
di: Lu, Ke-Han, et al.
Pubblicazione: (2026)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
di: Yen, Hao, et al.
Pubblicazione: (2024)
di: Yen, Hao, et al.
Pubblicazione: (2024)
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild Data
di: Ding, Wen, et al.
Pubblicazione: (2025)
di: Ding, Wen, et al.
Pubblicazione: (2025)
Measuring Sound Symbolism in Audio-visual Models
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2024)
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2024)
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
di: Wei, Sheng-Lun, et al.
Pubblicazione: (2026)
di: Wei, Sheng-Lun, et al.
Pubblicazione: (2026)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
di: Khan, Muhammad Salman, et al.
Pubblicazione: (2024)
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
di: Lu, Ke-Han, et al.
Pubblicazione: (2025)
di: Lu, Ke-Han, et al.
Pubblicazione: (2025)
Boosting Large Language Model for Speech Synthesis: An Empirical Study
di: Hao, Hongkun, et al.
Pubblicazione: (2023)
di: Hao, Hongkun, et al.
Pubblicazione: (2023)
Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis
di: Ortega-Beltrán, Elena, et al.
Pubblicazione: (2024)
di: Ortega-Beltrán, Elena, et al.
Pubblicazione: (2024)
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
di: Wu, Yihan, et al.
Pubblicazione: (2024)
di: Wu, Yihan, et al.
Pubblicazione: (2024)
ANIM-400K: A Large-Scale Dataset for Automated End-To-End Dubbing of Video
di: Cai, Kevin, et al.
Pubblicazione: (2024)
di: Cai, Kevin, et al.
Pubblicazione: (2024)
Qwen2.5-Omni Technical Report
di: Xu, Jin, et al.
Pubblicazione: (2025)
di: Xu, Jin, et al.
Pubblicazione: (2025)
Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder
di: Li, Zhengyang, et al.
Pubblicazione: (2026)
di: Li, Zhengyang, et al.
Pubblicazione: (2026)
Unsupervised Out-of-Distribution Dialect Detection with Mahalanobis Distance
di: Das, Sourya Dipta, et al.
Pubblicazione: (2023)
di: Das, Sourya Dipta, et al.
Pubblicazione: (2023)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
di: Yusuyin, Saierdaer, et al.
Pubblicazione: (2024)
di: Yusuyin, Saierdaer, et al.
Pubblicazione: (2024)
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
di: Bafna, Niyati, et al.
Pubblicazione: (2025)
di: Bafna, Niyati, et al.
Pubblicazione: (2025)
Semi-Autoregressive Streaming ASR With Label Context
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
di: Wang, Yujin, et al.
Pubblicazione: (2022)
di: Wang, Yujin, et al.
Pubblicazione: (2022)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
di: Wang, Haoyu, et al.
Pubblicazione: (2022)
di: Wang, Haoyu, et al.
Pubblicazione: (2022)
Bayesian Low-Rank Factorization for Robust Model Adaptation
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
di: Rajasekhar, Gnana Praveen, et al.
Pubblicazione: (2025)
di: Rajasekhar, Gnana Praveen, et al.
Pubblicazione: (2025)
Effective Context in Neural Speech Models
di: Meng, Yen, et al.
Pubblicazione: (2025)
di: Meng, Yen, et al.
Pubblicazione: (2025)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
di: Ye, Zhen, et al.
Pubblicazione: (2026)
di: Ye, Zhen, et al.
Pubblicazione: (2026)
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
di: Park, Sooyoung, et al.
Pubblicazione: (2025)
di: Park, Sooyoung, et al.
Pubblicazione: (2025)
SemiPL: A Semi-supervised Method for Event Sound Source Localization
di: Li, Yue, et al.
Pubblicazione: (2024)
di: Li, Yue, et al.
Pubblicazione: (2024)
Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression
di: Wu, Wen, et al.
Pubblicazione: (2023)
di: Wu, Wen, et al.
Pubblicazione: (2023)
Interface Design for Self-Supervised Speech Models
di: Shih, Yi-Jen, et al.
Pubblicazione: (2024)
di: Shih, Yi-Jen, et al.
Pubblicazione: (2024)
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
di: Rehman, Yasar Abbas Ur, et al.
Pubblicazione: (2025)
di: Rehman, Yasar Abbas Ur, et al.
Pubblicazione: (2025)
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
di: Poli, Maxime, et al.
Pubblicazione: (2025)
di: Poli, Maxime, et al.
Pubblicazione: (2025)
Self-Supervised Learning for Multi-Channel Neural Transducer
di: Kojima, Atsushi
Pubblicazione: (2024)
di: Kojima, Atsushi
Pubblicazione: (2024)
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
di: Goel, Arnav, et al.
Pubblicazione: (2024)
di: Goel, Arnav, et al.
Pubblicazione: (2024)
Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
di: Altinok, Duygu
Pubblicazione: (2025)
di: Altinok, Duygu
Pubblicazione: (2025)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
di: Fernandez-Lopez, Adriana, et al.
Pubblicazione: (2024) -
Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
di: Huang, Sung-Feng, et al.
Pubblicazione: (2025) -
ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
di: Gupta, Pranav, et al.
Pubblicazione: (2024) -
Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization
di: Guo, Yuxin, et al.
Pubblicazione: (2024) -
Universal Speech Enhancement with Regression and Generative Mamba
di: Chao, Rong, et al.
Pubblicazione: (2025)