Zipper: A Multi-Tower Decoder Architecture for Fusing Modalities
Fuente:
arXiv
Saved in:
| Main Authors: | Zayats, Vicky, Chen, Peter, Ferrari, Melissa, Padfield, Dirk |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Label-Looping: Highly Efficient Decoding for Transducers
by: Bataev, Vladimir, et al.
Published: (2024)
by: Bataev, Vladimir, et al.
Published: (2024)
Multimodal Sentiment Analysis with Missing Modality: A Knowledge-Transfer Approach
by: Liu, Weide, et al.
Published: (2023)
by: Liu, Weide, et al.
Published: (2023)
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
by: Grigoryan, Lilit, et al.
Published: (2025)
by: Grigoryan, Lilit, et al.
Published: (2025)
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
by: Grigoryan, Lilit, et al.
Published: (2025)
by: Grigoryan, Lilit, et al.
Published: (2025)
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
by: Li, Zehua Kcriss, et al.
Published: (2024)
by: Li, Zehua Kcriss, et al.
Published: (2024)
Decoding Poultry Vocalizations -- Natural Language Processing and Transformer Models for Semantic and Emotional Analysis
by: Manikandan, Venkatraman, et al.
Published: (2024)
by: Manikandan, Venkatraman, et al.
Published: (2024)
Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap
by: Nam, KiHyun, et al.
Published: (2025)
by: Nam, KiHyun, et al.
Published: (2025)
Modality-Invariant Bidirectional Temporal Representation Distillation Network for Missing Multimodal Sentiment Analysis
by: Wang, Xincheng, et al.
Published: (2025)
by: Wang, Xincheng, et al.
Published: (2025)
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings
by: Zhu, Yonggang, et al.
Published: (2026)
by: Zhu, Yonggang, et al.
Published: (2026)
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
by: Bataev, Vladimir, et al.
Published: (2025)
by: Bataev, Vladimir, et al.
Published: (2025)
Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
by: Noroozi, Vahid, et al.
Published: (2024)
by: Noroozi, Vahid, et al.
Published: (2024)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
by: Chen, William, et al.
Published: (2025)
by: Chen, William, et al.
Published: (2025)
TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
by: Zheng, Haolong, et al.
Published: (2025)
by: Zheng, Haolong, et al.
Published: (2025)
The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach
by: Ghazal, Nizar El, et al.
Published: (2025)
by: Ghazal, Nizar El, et al.
Published: (2025)
TIMIT Speaker Profiling: A Comparison of Multi-task learning and Single-task learning Approaches
by: Wang, Rong, et al.
Published: (2024)
by: Wang, Rong, et al.
Published: (2024)
Multi-Convformer: Extending Conformer with Multiple Convolution Kernels
by: Prabhu, Darshan, et al.
Published: (2024)
by: Prabhu, Darshan, et al.
Published: (2024)
NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
by: Lin, Yen-Ting, et al.
Published: (2024)
by: Lin, Yen-Ting, et al.
Published: (2024)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
by: Storey, Edward, et al.
Published: (2025)
by: Storey, Edward, et al.
Published: (2025)
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
by: Lee, Jaeyoung, et al.
Published: (2026)
by: Lee, Jaeyoung, et al.
Published: (2026)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
by: Yang, Chao-Han Huck, et al.
Published: (2024)
by: Yang, Chao-Han Huck, et al.
Published: (2024)
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
AMPS: ASR with Multimodal Paraphrase Supervision
by: Gupta, Abhishek, et al.
Published: (2024)
by: Gupta, Abhishek, et al.
Published: (2024)
An Automated End-to-End Open-Source Software for High-Quality Text-to-Speech Dataset Generation
by: Gunduz, Ahmet, et al.
Published: (2024)
by: Gunduz, Ahmet, et al.
Published: (2024)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
by: Hono, Yukiya, et al.
Published: (2023)
by: Hono, Yukiya, et al.
Published: (2023)
Effective internal language model training and fusion for factorized transducer model
by: Guo, Jinxi, et al.
Published: (2024)
by: Guo, Jinxi, et al.
Published: (2024)
Semantically Corrected Amharic Automatic Speech Recognition
by: Adnew, Samuael, et al.
Published: (2024)
by: Adnew, Samuael, et al.
Published: (2024)
LASER: An LLM-based ASR Scoring and Evaluation Rubric
by: Parulekar, Amruta, et al.
Published: (2025)
by: Parulekar, Amruta, et al.
Published: (2025)
Controllable Prosody Generation With Partial Inputs
by: Iliescu, Dan Andrei, et al.
Published: (2023)
by: Iliescu, Dan Andrei, et al.
Published: (2023)
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
by: Huang, Ruizhe, et al.
Published: (2024)
by: Huang, Ruizhe, et al.
Published: (2024)
AG-LSEC: Audio Grounded Lexical Speaker Error Correction
by: Paturi, Rohit, et al.
Published: (2024)
by: Paturi, Rohit, et al.
Published: (2024)
Latent Speech-Text Transformer
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs
by: Fedorchenko, Artem, et al.
Published: (2025)
by: Fedorchenko, Artem, et al.
Published: (2025)
Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR
by: Gupta, Abhishek, et al.
Published: (2024)
by: Gupta, Abhishek, et al.
Published: (2024)
Improving Acoustic Side-Channel Attacks on Keyboards Using Transformers and Large Language Models
by: Park, Jin Hyun, et al.
Published: (2025)
by: Park, Jin Hyun, et al.
Published: (2025)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
by: Zhou, Qinhao, et al.
Published: (2025)
by: Zhou, Qinhao, et al.
Published: (2025)
Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR
by: Kumar, Shashi, et al.
Published: (2026)
by: Kumar, Shashi, et al.
Published: (2026)
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
by: Gulzar, Kashaf, et al.
Published: (2025)
by: Gulzar, Kashaf, et al.
Published: (2025)
SpeechPrompt: Prompting Speech Language Models for Speech Processing Tasks
by: Chang, Kai-Wei, et al.
Published: (2024)
by: Chang, Kai-Wei, et al.
Published: (2024)
Similar Items
-
Label-Looping: Highly Efficient Decoding for Transducers
by: Bataev, Vladimir, et al.
Published: (2024) -
Multimodal Sentiment Analysis with Missing Modality: A Knowledge-Transfer Approach
by: Liu, Weide, et al.
Published: (2023) -
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
by: Grigoryan, Lilit, et al.
Published: (2025) -
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
by: Grigoryan, Lilit, et al.
Published: (2025) -
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
by: Li, Zehua Kcriss, et al.
Published: (2024)