Gespeichert in:
| Hauptverfasser: | Onda, Kentaro, Futami, Hayato, Kashiwagi, Yosuke, Tsunoo, Emiru, Watanabe, Shinji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.19781 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Differentiable K-means for Fully-optimized Discrete Token-based ASR
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2025)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2025)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
von: Arora, Siddhant, et al.
Veröffentlicht: (2026)
von: Arora, Siddhant, et al.
Veröffentlicht: (2026)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
von: Futami, Hayato, et al.
Veröffentlicht: (2024)
von: Futami, Hayato, et al.
Veröffentlicht: (2024)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2025)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2025)
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Advanced Modeling of Interlanguage Speech Intelligibility Benefit with L1-L2 Multi-Task Learning Using Differentiable K-Means for Accent-Robust Discrete Token-Based ASR
von: Onda, Kentaro, et al.
Veröffentlicht: (2026)
von: Onda, Kentaro, et al.
Veröffentlicht: (2026)
Task Arithmetic for Language Expansion in Speech Translation
von: Cheng, Yao-Fei, et al.
Veröffentlicht: (2024)
von: Cheng, Yao-Fei, et al.
Veröffentlicht: (2024)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
von: Nakata, Wataru, et al.
Veröffentlicht: (2026)
von: Nakata, Wataru, et al.
Veröffentlicht: (2026)
LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
von: Ohmura, Junki, et al.
Veröffentlicht: (2025)
von: Ohmura, Junki, et al.
Veröffentlicht: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
von: Wang, Pu, et al.
Veröffentlicht: (2026)
von: Wang, Pu, et al.
Veröffentlicht: (2026)
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
von: Someki, Masao, et al.
Veröffentlicht: (2023)
von: Someki, Masao, et al.
Veröffentlicht: (2023)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2026)
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
von: Onda, Kentaro, et al.
Veröffentlicht: (2024)
von: Onda, Kentaro, et al.
Veröffentlicht: (2024)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
von: Ivry, Amir, et al.
Veröffentlicht: (2026)
von: Ivry, Amir, et al.
Veröffentlicht: (2026)
EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
von: Wang, Dingdong, et al.
Veröffentlicht: (2026)
von: Wang, Dingdong, et al.
Veröffentlicht: (2026)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Geolocation-Aware Robust Spoken Language Identification
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
Sign-to-Speech Prosody Transfer via Sign Reconstruction-based GAN
von: Manabe, Toranosuke, et al.
Veröffentlicht: (2026)
von: Manabe, Toranosuke, et al.
Veröffentlicht: (2026)
DuoTok: Source-Aware Dual-Track Tokenization for Multi-Track Music Language Modeling
von: Lin, Rui, et al.
Veröffentlicht: (2025)
von: Lin, Rui, et al.
Veröffentlicht: (2025)
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
von: Li, Jialu, et al.
Veröffentlicht: (2023)
von: Li, Jialu, et al.
Veröffentlicht: (2023)
CantoASR: Prosody-Aware ASR-LALM Collaboration for Low-Resource Cantonese
von: Chen, Dazhong, et al.
Veröffentlicht: (2025)
von: Chen, Dazhong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Differentiable K-means for Fully-optimized Discrete Token-based ASR
von: Onda, Kentaro, et al.
Veröffentlicht: (2025) -
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024) -
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023) -
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2025) -
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)