Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
Fuente:
arXiv
Guardado en:
| Autores principales: | Tian, Jinchuan, Wang, Haoran, Su, Bo-Hao, Huang, Chien-yu, Wang, Qingzheng, Shi, Jiatong, Chen, William, Gong, Xun, Arora, Siddhant, Li, Chin-Jou, Someki, Masao, Maekaku, Takashi, Goto, Keita, Shinohara, Yusuke, Sakuma, Jin, Yang, Chao-Han Huck, Watanabe, Shinji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
por: Goto, Keita, et al.
Publicado: (2026)
por: Goto, Keita, et al.
Publicado: (2026)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
por: Maekaku, Takashi, et al.
Publicado: (2025)
por: Maekaku, Takashi, et al.
Publicado: (2025)
OpusLM: A Family of Open Unified Speech Language Models
por: Tian, Jinchuan, et al.
Publicado: (2025)
por: Tian, Jinchuan, et al.
Publicado: (2025)
ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
por: Someki, Masao, et al.
Publicado: (2024)
por: Someki, Masao, et al.
Publicado: (2024)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
por: Arora, Siddhant, et al.
Publicado: (2025)
por: Arora, Siddhant, et al.
Publicado: (2025)
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
por: Arora, Siddhant, et al.
Publicado: (2026)
por: Arora, Siddhant, et al.
Publicado: (2026)
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
por: Someki, Masao, et al.
Publicado: (2023)
por: Someki, Masao, et al.
Publicado: (2023)
On-device Streaming Discrete Speech Units
por: Choi, Kwanghee, et al.
Publicado: (2025)
por: Choi, Kwanghee, et al.
Publicado: (2025)
BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
por: Wang, Haoran, et al.
Publicado: (2025)
por: Wang, Haoran, et al.
Publicado: (2025)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
por: Arora, Siddhant, et al.
Publicado: (2025)
por: Arora, Siddhant, et al.
Publicado: (2025)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
por: Tian, Jinchuan, et al.
Publicado: (2025)
por: Tian, Jinchuan, et al.
Publicado: (2025)
SingingSDS: A Singing-Capable Spoken Dialogue System for Conversational Roleplay Applications
por: Han, Jionghao, et al.
Publicado: (2025)
por: Han, Jionghao, et al.
Publicado: (2025)
Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR
por: Magoshi, Ryo, et al.
Publicado: (2026)
por: Magoshi, Ryo, et al.
Publicado: (2026)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
por: Wang, Shih-heng, et al.
Publicado: (2024)
por: Wang, Shih-heng, et al.
Publicado: (2024)
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
por: Shi, Jiatong, et al.
Publicado: (2025)
por: Shi, Jiatong, et al.
Publicado: (2025)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
por: Shi, Jiatong, et al.
Publicado: (2025)
por: Shi, Jiatong, et al.
Publicado: (2025)
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
por: Wang, Shih-Heng, et al.
Publicado: (2026)
por: Wang, Shih-Heng, et al.
Publicado: (2026)
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
por: Fujita, Yuya, et al.
Publicado: (2024)
por: Fujita, Yuya, et al.
Publicado: (2024)
Context-Driven Dynamic Pruning for Large Speech Foundation Models
por: Someki, Masao, et al.
Publicado: (2025)
por: Someki, Masao, et al.
Publicado: (2025)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
por: Arora, Siddhant, et al.
Publicado: (2025)
por: Arora, Siddhant, et al.
Publicado: (2025)
Semi-Autoregressive Streaming ASR With Label Context
por: Arora, Siddhant, et al.
Publicado: (2023)
por: Arora, Siddhant, et al.
Publicado: (2023)
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
por: Hirota, Yusuke, et al.
Publicado: (2024)
por: Hirota, Yusuke, et al.
Publicado: (2024)
On the Berkovich double residue fields and birational models
por: Goto, Keita
Publicado: (2020)
por: Goto, Keita
Publicado: (2020)
On The Two Types Of Affine Structures For Degenerating Kummer Surfaces -Non-Archimedean VS Gromov-Hausdorff Limits-
por: Goto, Keita
Publicado: (2022)
por: Goto, Keita
Publicado: (2022)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
por: Wang, Qingzheng, et al.
Publicado: (2025)
por: Wang, Qingzheng, et al.
Publicado: (2025)
Structural and Instrumental Power of Real Estate Companies in China and Government Regulation: The Case of Evergrande
por: Qingzheng Wang
Publicado: (2025)
por: Qingzheng Wang
Publicado: (2025)
Geolocation-Aware Robust Spoken Language Identification
por: Wang, Qingzheng, et al.
Publicado: (2025)
por: Wang, Qingzheng, et al.
Publicado: (2025)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
por: Chen, William, et al.
Publicado: (2025)
por: Chen, William, et al.
Publicado: (2025)
Preference Alignment Improves Language Model-Based TTS
por: Tian, Jinchuan, et al.
Publicado: (2024)
por: Tian, Jinchuan, et al.
Publicado: (2024)
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM
por: Shi, Jiatong, et al.
Publicado: (2025)
por: Shi, Jiatong, et al.
Publicado: (2025)
Adapting Speech Language Model to Singing Voice Synthesis
por: Zhao, Yiwen, et al.
Publicado: (2025)
por: Zhao, Yiwen, et al.
Publicado: (2025)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
por: Tsunoo, Emiru, et al.
Publicado: (2025)
por: Tsunoo, Emiru, et al.
Publicado: (2025)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
por: Futami, Hayato, et al.
Publicado: (2024)
por: Futami, Hayato, et al.
Publicado: (2024)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
por: Tsunoo, Emiru, et al.
Publicado: (2023)
por: Tsunoo, Emiru, et al.
Publicado: (2023)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
por: Tsunoo, Emiru, et al.
Publicado: (2024)
por: Tsunoo, Emiru, et al.
Publicado: (2024)
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
por: Lin, Guan-Ting, et al.
Publicado: (2025)
por: Lin, Guan-Ting, et al.
Publicado: (2025)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
por: Li, Chin-Jou, et al.
Publicado: (2025)
por: Li, Chin-Jou, et al.
Publicado: (2025)
Ejemplares similares
-
Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
por: Goto, Keita, et al.
Publicado: (2026) -
Evaluating Self-Supervised Speech Models via Text-Based LLMS
por: Maekaku, Takashi, et al.
Publicado: (2025) -
OpusLM: A Family of Open Unified Speech Language Models
por: Tian, Jinchuan, et al.
Publicado: (2025) -
ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
por: Someki, Masao, et al.
Publicado: (2024) -
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
por: Arora, Siddhant, et al.
Publicado: (2025)