Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Nachmani, Eliya, Levkovitch, Alon, Hirsch, Roy, Salazar, Julian, Asawaroengchai, Chulayuth, Mariooryad, Soroosh, Rivlin, Ehud, Skerry-Ryan, RJ, Ramanovich, Michelle Tadmor |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Translatotron 3: Speech to Speech Translation with Monolingual Data
by: Nachmani, Eliya, et al.
Published: (2023)
by: Nachmani, Eliya, et al.
Published: (2023)
Zero-Shot Mono-to-Binaural Speech Synthesis
by: Levkovitch, Alon, et al.
Published: (2024)
by: Levkovitch, Alon, et al.
Published: (2024)
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
by: Battenberg, Eric, et al.
Published: (2024)
by: Battenberg, Eric, et al.
Published: (2024)
SimulTron: On-Device Simultaneous Speech to Speech Translation
by: Agranovich, Alex, et al.
Published: (2024)
by: Agranovich, Alex, et al.
Published: (2024)
Long-Form Speech Generation with Spoken Language Models
by: Park, Se Jin, et al.
Published: (2024)
by: Park, Se Jin, et al.
Published: (2024)
Score Based Error Correcting Code Decoder
by: Helvits, Alon, et al.
Published: (2026)
by: Helvits, Alon, et al.
Published: (2026)
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
by: Yaish, Ofir, et al.
Published: (2025)
by: Yaish, Ofir, et al.
Published: (2025)
SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
by: Skerry-Ryan, RJ, et al.
Published: (2025)
by: Skerry-Ryan, RJ, et al.
Published: (2025)
Deep Active Speech Cancellation with Mamba-Masking Network
by: Mishaly, Yehuda, et al.
Published: (2025)
by: Mishaly, Yehuda, et al.
Published: (2025)
ILRR: Inference-Time Steering Method for Masked Diffusion Language Models
by: Avrahami, Eden, et al.
Published: (2026)
by: Avrahami, Eden, et al.
Published: (2026)
SAQ: Stabilizer-Aware Quantum Error Correction Decoder
by: Zenati, David, et al.
Published: (2025)
by: Zenati, David, et al.
Published: (2025)
Two-Dimensional Quantization for Geometry-Aware Audio Coding
by: Shuster, Tal, et al.
Published: (2025)
by: Shuster, Tal, et al.
Published: (2025)
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
by: Varshavsky-Hassid, Miri, et al.
Published: (2024)
by: Varshavsky-Hassid, Miri, et al.
Published: (2024)
Differential Mamba
by: Schneider, Nadav, et al.
Published: (2025)
by: Schneider, Nadav, et al.
Published: (2025)
Neural Minimum Weight Perfect Matching for Quantum Error Codes
by: Peled, Yotam, et al.
Published: (2026)
by: Peled, Yotam, et al.
Published: (2026)
Toward Optimal ANC: Establishing Mutual Information Lower Bound
by: Derrida, François, et al.
Published: (2025)
by: Derrida, François, et al.
Published: (2025)
STAB: Speech Tokenizer Assessment Benchmark
by: Vashishth, Shikhar, et al.
Published: (2024)
by: Vashishth, Shikhar, et al.
Published: (2024)
Neural Brain Fields: A NeRF-Inspired Approach for Generating Nonexistent EEG Electrodes
by: Kedem, Shahar Ain, et al.
Published: (2025)
by: Kedem, Shahar Ain, et al.
Published: (2025)
Hybrid Mamba-Transformer Decoder for Error-Correcting Codes
by: Cohen, Shy-el, et al.
Published: (2025)
by: Cohen, Shy-el, et al.
Published: (2025)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
The Role of Prosody in Spoken Question Answering
by: Chi, Jie, et al.
Published: (2025)
by: Chi, Jie, et al.
Published: (2025)
Neural Descriptors: Self-Supervised Learning of Robust Local Surface Descriptors Using Polynomial Patches
by: Yona, Gal, et al.
Published: (2025)
by: Yona, Gal, et al.
Published: (2025)
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
by: Hu, Jiliang, et al.
Published: (2025)
by: Hu, Jiliang, et al.
Published: (2025)
Attention-guided Evidence Grounding for Spoken Question Answering
by: Yang, Ke, et al.
Published: (2026)
by: Yang, Ke, et al.
Published: (2026)
Semantic Parsing of Colonoscopy Videos with Multi-Label Temporal Networks
by: Kelner, Ori, et al.
Published: (2023)
by: Kelner, Ori, et al.
Published: (2023)
HeySQuAD: A Spoken Question Answering Dataset
by: Wu, Yijing, et al.
Published: (2023)
by: Wu, Yijing, et al.
Published: (2023)
GSQA: An End-to-End Model for Generative Spoken Question Answering
by: Shih, Min-Han, et al.
Published: (2023)
by: Shih, Min-Han, et al.
Published: (2023)
Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings
by: Shor, Joel, et al.
Published: (2023)
by: Shor, Joel, et al.
Published: (2023)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
by: Labrak, Yanis, et al.
Published: (2024)
by: Labrak, Yanis, et al.
Published: (2024)
Self-Supervised Learning for Endoscopic Video Analysis
by: Hirsch, Roy, et al.
Published: (2023)
by: Hirsch, Roy, et al.
Published: (2023)
Molecular Diffusion Models with Virtual Receptors
by: Halfon, Matan, et al.
Published: (2024)
by: Halfon, Matan, et al.
Published: (2024)
MESSI: A Multi-Elevation Semantic Segmentation Image Dataset of an Urban Environment
by: Pinkovich, Barak, et al.
Published: (2025)
by: Pinkovich, Barak, et al.
Published: (2025)
Looks Too Good To Be True: An Information-Theoretic Analysis of Hallucinations in Generative Restoration Models
by: Cohen, Regev, et al.
Published: (2024)
by: Cohen, Regev, et al.
Published: (2024)
Turkey facing a new millenium: Coping with intertwined conflicts
by: Nachmani, Amikam
Published: (2010)
by: Nachmani, Amikam
Published: (2010)
Token-Based Audio Inpainting via Discrete Diffusion
by: Dror, Tali, et al.
Published: (2025)
by: Dror, Tali, et al.
Published: (2025)
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering
by: Rackauckas, Zackary, et al.
Published: (2025)
by: Rackauckas, Zackary, et al.
Published: (2025)
LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models
by: Zhao, Zihan, et al.
Published: (2023)
by: Zhao, Zihan, et al.
Published: (2023)
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
by: Xie, Jingran, et al.
Published: (2025)
by: Xie, Jingran, et al.
Published: (2025)
Principal Uncertainty Quantification with Spatial Correlation for Image Restoration Problems
by: Belhasin, Omer, et al.
Published: (2023)
by: Belhasin, Omer, et al.
Published: (2023)
Self-Supervised Polyp Re-Identification in Colonoscopy
by: Intrator, Yotam, et al.
Published: (2023)
by: Intrator, Yotam, et al.
Published: (2023)
Similar Items
-
Translatotron 3: Speech to Speech Translation with Monolingual Data
by: Nachmani, Eliya, et al.
Published: (2023) -
Zero-Shot Mono-to-Binaural Speech Synthesis
by: Levkovitch, Alon, et al.
Published: (2024) -
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
by: Battenberg, Eric, et al.
Published: (2024) -
SimulTron: On-Device Simultaneous Speech to Speech Translation
by: Agranovich, Alex, et al.
Published: (2024) -
Long-Form Speech Generation with Spoken Language Models
by: Park, Se Jin, et al.
Published: (2024)