Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Someki, Masao, Eng, Nicholas, Higuchi, Yosuke, Watanabe, Shinji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On-device Streaming Discrete Speech Units
por: Choi, Kwanghee, et al.
Publicado: (2025)
por: Choi, Kwanghee, et al.
Publicado: (2025)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
por: Tsunoo, Emiru, et al.
Publicado: (2023)
por: Tsunoo, Emiru, et al.
Publicado: (2023)
Differentiable K-means for Fully-optimized Discrete Token-based ASR
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
por: Sudo, Yui, et al.
Publicado: (2025)
por: Sudo, Yui, et al.
Publicado: (2025)
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
por: Someki, Masao, et al.
Publicado: (2024)
por: Someki, Masao, et al.
Publicado: (2024)
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
por: Ivry, Amir, et al.
Publicado: (2026)
por: Ivry, Amir, et al.
Publicado: (2026)
Semi-Autoregressive Streaming ASR With Label Context
por: Arora, Siddhant, et al.
Publicado: (2023)
por: Arora, Siddhant, et al.
Publicado: (2023)
MAPSS: Manifold-based Assessment of Perceptual Source Separation
por: Ivry, Amir, et al.
Publicado: (2025)
por: Ivry, Amir, et al.
Publicado: (2025)
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
por: Tu, Zehai, et al.
Publicado: (2024)
por: Tu, Zehai, et al.
Publicado: (2024)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
por: Shi, Jiatong, et al.
Publicado: (2025)
por: Shi, Jiatong, et al.
Publicado: (2025)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
por: Maekaku, Takashi, et al.
Publicado: (2025)
por: Maekaku, Takashi, et al.
Publicado: (2025)
Uni-VERSA: Versatile Speech Assessment with a Unified Network
por: Shi, Jiatong, et al.
Publicado: (2025)
por: Shi, Jiatong, et al.
Publicado: (2025)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
por: Higuchi, Yosuke, et al.
Publicado: (2023)
por: Higuchi, Yosuke, et al.
Publicado: (2023)
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation
por: Luong, Hieu-Thi, et al.
Publicado: (2025)
por: Luong, Hieu-Thi, et al.
Publicado: (2025)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
por: Tsunoo, Emiru, et al.
Publicado: (2025)
por: Tsunoo, Emiru, et al.
Publicado: (2025)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
por: Chen, Li-Wei, et al.
Publicado: (2024)
por: Chen, Li-Wei, et al.
Publicado: (2024)
EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
por: Srivastava, Tejes, et al.
Publicado: (2023)
por: Srivastava, Tejes, et al.
Publicado: (2023)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
por: Udupa, Sathvik, et al.
Publicado: (2025)
por: Udupa, Sathvik, et al.
Publicado: (2025)
Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
por: Li, Chenda, et al.
Publicado: (2024)
por: Li, Chenda, et al.
Publicado: (2024)
LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
por: Luong, Hieu-Thi, et al.
Publicado: (2024)
por: Luong, Hieu-Thi, et al.
Publicado: (2024)
Improving Design of Input Condition Invariant Speech Enhancement
por: Zhang, Wangyou, et al.
Publicado: (2024)
por: Zhang, Wangyou, et al.
Publicado: (2024)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
por: Shakeel, Muhammad, et al.
Publicado: (2026)
por: Shakeel, Muhammad, et al.
Publicado: (2026)
Multichannel Voice Trigger Detection Based on Transform-average-concatenate
por: Higuchi, Takuya, et al.
Publicado: (2023)
por: Higuchi, Takuya, et al.
Publicado: (2023)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
por: Shi, Jiatong, et al.
Publicado: (2024)
por: Shi, Jiatong, et al.
Publicado: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
por: Guo, Pengcheng, et al.
Publicado: (2024)
por: Guo, Pengcheng, et al.
Publicado: (2024)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
por: Saeki, Takaaki, et al.
Publicado: (2024)
por: Saeki, Takaaki, et al.
Publicado: (2024)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
por: Yu, Yifeng, et al.
Publicado: (2024)
por: Yu, Yifeng, et al.
Publicado: (2024)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
por: Plaquet, Alexis, et al.
Publicado: (2025)
por: Plaquet, Alexis, et al.
Publicado: (2025)
Cross-Talk Reduction
por: Wang, Zhong-Qiu, et al.
Publicado: (2024)
por: Wang, Zhong-Qiu, et al.
Publicado: (2024)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
por: Wang, Wei, et al.
Publicado: (2025)
por: Wang, Wei, et al.
Publicado: (2025)
Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor
por: Lee, Younglo, et al.
Publicado: (2024)
por: Lee, Younglo, et al.
Publicado: (2024)
Autoregressive Speech Synthesis without Vector Quantization
por: Meng, Lingwei, et al.
Publicado: (2024)
por: Meng, Lingwei, et al.
Publicado: (2024)
EffectiveASR: A Single-Step Non-Autoregressive Mandarin Speech Recognition Architecture with High Accuracy and Inference Speed
por: Zhuang, Ziyang, et al.
Publicado: (2024)
por: Zhuang, Ziyang, et al.
Publicado: (2024)
Ejemplares similares
-
On-device Streaming Discrete Speech Units
por: Choi, Kwanghee, et al.
Publicado: (2025) -
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
por: Sudo, Yui, et al.
Publicado: (2024) -
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
por: Sudo, Yui, et al.
Publicado: (2024) -
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
por: Tsunoo, Emiru, et al.
Publicado: (2023) -
Differentiable K-means for Fully-optimized Discrete Token-based ASR
por: Onda, Kentaro, et al.
Publicado: (2025)