Logit Distillation on Manifolds: Mapping by Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Yiru, Wang, Junling, Singh, Nishant Kumar, Wu, Luohong, Yan, Haoran |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models
por: Sienkiewicz, Bruno, et al.
Publicado: (2026)
por: Sienkiewicz, Bruno, et al.
Publicado: (2026)
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
por: Novack, Zachary, et al.
Publicado: (2024)
por: Novack, Zachary, et al.
Publicado: (2024)
AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation
por: Lu, Yuxin, et al.
Publicado: (2026)
por: Lu, Yuxin, et al.
Publicado: (2026)
Multi-Task Learning for Lung sound & Lung disease classification
por: K V, Suma, et al.
Publicado: (2024)
por: K V, Suma, et al.
Publicado: (2024)
Real-Time Voicemail Detection in Telephony Audio Using Temporal Speech Activity Features
por: Saurav, Kumar
Publicado: (2026)
por: Saurav, Kumar
Publicado: (2026)
Myna: Masking-Based Contrastive Learning of Musical Representations
por: Yonay, Ori, et al.
Publicado: (2025)
por: Yonay, Ori, et al.
Publicado: (2025)
Dual Knowledge Distillation for Efficient Sound Event Detection
por: Xiao, Yang, et al.
Publicado: (2024)
por: Xiao, Yang, et al.
Publicado: (2024)
DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching
por: Chen, Wei, et al.
Publicado: (2025)
por: Chen, Wei, et al.
Publicado: (2025)
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
por: Wang, Lu, et al.
Publicado: (2025)
por: Wang, Lu, et al.
Publicado: (2025)
AudioMosaic: Contrastive Masked Audio Representation Learning
por: Huang, Hanxun, et al.
Publicado: (2026)
por: Huang, Hanxun, et al.
Publicado: (2026)
Linguistic and Audio Embedding-Based Machine Learning for Alzheimer's Dementia and Mild Cognitive Impairment Detection: Insights from the PROCESS Challenge
por: Devahi, Adharsha Sam Edwin Sam, et al.
Publicado: (2025)
por: Devahi, Adharsha Sam Edwin Sam, et al.
Publicado: (2025)
SAND Challenge: Four Approaches for Dysartria Severity Classification
por: Deshpande, Gauri, et al.
Publicado: (2025)
por: Deshpande, Gauri, et al.
Publicado: (2025)
QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems
por: Wang, Chien-Chun, et al.
Publicado: (2025)
por: Wang, Chien-Chun, et al.
Publicado: (2025)
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
por: Agrawal, Saurabh, et al.
Publicado: (2025)
por: Agrawal, Saurabh, et al.
Publicado: (2025)
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
por: Liao, Huan, et al.
Publicado: (2025)
por: Liao, Huan, et al.
Publicado: (2025)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
por: Chi, Hyung Gun, et al.
Publicado: (2025)
por: Chi, Hyung Gun, et al.
Publicado: (2025)
Preference-Based Learning in Audio Applications: A Systematic Analysis
por: Broukhim, Aaron, et al.
Publicado: (2025)
por: Broukhim, Aaron, et al.
Publicado: (2025)
A Human-Inspired Decoupled Architecture for Efficient Audio Representation Learning
por: Kawano, Harunori, et al.
Publicado: (2026)
por: Kawano, Harunori, et al.
Publicado: (2026)
Recognizing Ornaments in Vocal Indian Art Music with Active Annotation
por: Kumar, Sumit, et al.
Publicado: (2025)
por: Kumar, Sumit, et al.
Publicado: (2025)
Tri-MTL: A Triple Multitask Learning Approach for Respiratory Disease Diagnosis
por: Kim, June-Woo, et al.
Publicado: (2025)
por: Kim, June-Woo, et al.
Publicado: (2025)
Phase-Aware Deep Learning with Complex-Valued CNNs for Audio Signal Applications
por: Agrawal, Naman
Publicado: (2025)
por: Agrawal, Naman
Publicado: (2025)
AUDRON: A Deep Learning Framework with Fused Acoustic Signatures for Drone Type Recognition
por: Chatterjee, Rajdeep, et al.
Publicado: (2025)
por: Chatterjee, Rajdeep, et al.
Publicado: (2025)
A$^2$-LLM: An End-to-end Conversational Audio Avatar Large Language Model
por: Hu, Xiaolin, et al.
Publicado: (2026)
por: Hu, Xiaolin, et al.
Publicado: (2026)
Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards
por: Fang, Linghan, et al.
Publicado: (2026)
por: Fang, Linghan, et al.
Publicado: (2026)
Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation
por: Si, Chen, et al.
Publicado: (2025)
por: Si, Chen, et al.
Publicado: (2025)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
por: Della Libera, Luca, et al.
Publicado: (2026)
por: Della Libera, Luca, et al.
Publicado: (2026)
Explainable Multi-Modal Deep Learning for Automatic Detection of Lung Diseases from Respiratory Audio Signals
por: Saky, S M Asiful Islam, et al.
Publicado: (2025)
por: Saky, S M Asiful Islam, et al.
Publicado: (2025)
Sample-Efficient Diffusion for Text-To-Speech Synthesis
por: Lovelace, Justin, et al.
Publicado: (2024)
por: Lovelace, Justin, et al.
Publicado: (2024)
TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation
por: Mazumdar, Soumya, et al.
Publicado: (2026)
por: Mazumdar, Soumya, et al.
Publicado: (2026)
Hookpad Aria: A Copilot for Songwriters
por: Donahue, Chris, et al.
Publicado: (2025)
por: Donahue, Chris, et al.
Publicado: (2025)
Of All StrIPEs: Investigating Structure-informed Positional Encoding for Efficient Music Generation
por: Agarwal, Manvi, et al.
Publicado: (2025)
por: Agarwal, Manvi, et al.
Publicado: (2025)
Presto! Distilling Steps and Layers for Accelerating Music Generation
por: Novack, Zachary, et al.
Publicado: (2024)
por: Novack, Zachary, et al.
Publicado: (2024)
AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
por: Wang, Qizhou, et al.
Publicado: (2025)
por: Wang, Qizhou, et al.
Publicado: (2025)
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
por: Tripathi, Suraj, et al.
Publicado: (2019)
por: Tripathi, Suraj, et al.
Publicado: (2019)
DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
por: Chen, Yang, et al.
Publicado: (2024)
por: Chen, Yang, et al.
Publicado: (2024)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
por: Della Libera, Luca, et al.
Publicado: (2025)
por: Della Libera, Luca, et al.
Publicado: (2025)
Peak-Controlled Logits Poisoning Attack in Federated Distillation
por: Tang, Yuhan, et al.
Publicado: (2024)
por: Tang, Yuhan, et al.
Publicado: (2024)
Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
por: Singh, Karamvir
Publicado: (2025)
por: Singh, Karamvir
Publicado: (2025)
Learning When to Think While Listening in Large Audio-Language Models
por: Song, Zhiyuan, et al.
Publicado: (2026)
por: Song, Zhiyuan, et al.
Publicado: (2026)
Beyond Fixed Frames: Dynamic Character-Aligned Speech Tokenization
por: Della Libera, Luca, et al.
Publicado: (2026)
por: Della Libera, Luca, et al.
Publicado: (2026)
Ejemplares similares
-
ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models
por: Sienkiewicz, Bruno, et al.
Publicado: (2026) -
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
por: Novack, Zachary, et al.
Publicado: (2024) -
AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation
por: Lu, Yuxin, et al.
Publicado: (2026) -
Multi-Task Learning for Lung sound & Lung disease classification
por: K V, Suma, et al.
Publicado: (2024) -
Real-Time Voicemail Detection in Telephony Audio Using Temporal Speech Activity Features
por: Saurav, Kumar
Publicado: (2026)