SpeechVerse: A Large-scale Generalizable Audio Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Das, Nilaksh, Dingliwal, Saket, Ronanki, Srikanth, Paturi, Rohit, Huang, Zhaocheng, Mathur, Prashant, Yuan, Jie, Bekal, Dhanush, Niu, Xing, Jayanthi, Sai Muralidhar, Li, Xilai, Mundnich, Karel, Sunkara, Monica, Bodapati, Sravan, Srinivasan, Sundararajan, Han, Kyu J, Kirchhoff, Katrin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2023)
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2023)
Accelerated Test-Time Scaling with Model-Free Speculative Sampling
von: Song, Woomin, et al.
Veröffentlicht: (2025)
von: Song, Woomin, et al.
Veröffentlicht: (2025)
Zero-resource Speech Translation and Recognition with LLMs
von: Mundnich, Karel, et al.
Veröffentlicht: (2024)
von: Mundnich, Karel, et al.
Veröffentlicht: (2024)
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2025)
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2025)
Sequential Editing for Lifelong Training of Speech Recognition Models
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
von: Song, Woomin, et al.
Veröffentlicht: (2025)
von: Song, Woomin, et al.
Veröffentlicht: (2025)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
von: Min, Do June, et al.
Veröffentlicht: (2024)
von: Min, Do June, et al.
Veröffentlicht: (2024)
Think Clearly: Improving Reasoning via Redundant Token Pruning
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
AG-LSEC: Audio Grounded Lexical Speaker Error Correction
von: Paturi, Rohit, et al.
Veröffentlicht: (2024)
von: Paturi, Rohit, et al.
Veröffentlicht: (2024)
IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents
von: Choi, Daewon, et al.
Veröffentlicht: (2026)
von: Choi, Daewon, et al.
Veröffentlicht: (2026)
ExComm: Exploration-Stage Communication for Error-Resilient Agentic Test-Time Scaling
von: Song, Woomin, et al.
Veröffentlicht: (2026)
von: Song, Woomin, et al.
Veröffentlicht: (2026)
LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models
von: Kumar, Anurag, et al.
Veröffentlicht: (2025)
von: Kumar, Anurag, et al.
Veröffentlicht: (2025)
Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
von: Du, Yufeng, et al.
Veröffentlicht: (2025)
von: Du, Yufeng, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications
von: Shu, Raphael, et al.
Veröffentlicht: (2024)
von: Shu, Raphael, et al.
Veröffentlicht: (2024)
Salient Information Prompting to Steer Content in Prompt-based Abstractive Summarization
von: Xu, Lei, et al.
Veröffentlicht: (2024)
von: Xu, Lei, et al.
Veröffentlicht: (2024)
One‐Sided Schmitt Trigger‐Based 11T Carbon Nanotube Field Effect Transistor—Based Static Random‐Access Memory Cell for Modern IoT Embedded Devices at 32 nm Technology
von: Srinivasan Jayanthi, et al.
Veröffentlicht: (2025)
von: Srinivasan Jayanthi, et al.
Veröffentlicht: (2025)
Recommending Machine Translation Output to Translators by Estimating Translation Effort: A Case Study
von: Prashant Mathur
Veröffentlicht: (2013)
von: Prashant Mathur
Veröffentlicht: (2013)
Adaptive Video Understanding Agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning
von: Jeoung, Sullam, et al.
Veröffentlicht: (2024)
von: Jeoung, Sullam, et al.
Veröffentlicht: (2024)
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt Optimization for Text Generation
von: He, Han, et al.
Veröffentlicht: (2024)
von: He, Han, et al.
Veröffentlicht: (2024)
Facilitating Trustworthy Human-Agent Collaboration in LLM-based Multi-Agent System oriented Software Engineering
von: Ronanki, Krishna
Veröffentlicht: (2025)
von: Ronanki, Krishna
Veröffentlicht: (2025)
Reversibility, balance and expansivity of non-uniform cellular automata
von: Paturi, Katariina
Veröffentlicht: (2025)
von: Paturi, Katariina
Veröffentlicht: (2025)
Barrier Functions Inspired Reward Shaping for Reinforcement Learning
von: Nilaksh, Nilaksh, et al.
Veröffentlicht: (2024)
von: Nilaksh, Nilaksh, et al.
Veröffentlicht: (2024)
Limits to positional information in boundary-driven systems
von: Singh, Prashant, et al.
Veröffentlicht: (2025)
von: Singh, Prashant, et al.
Veröffentlicht: (2025)
Limits to positional information in boundary-driven systems
von: Singh, Prashant, et al.
Veröffentlicht: (2024)
von: Singh, Prashant, et al.
Veröffentlicht: (2024)
Inferring entropy production from time-dependent moments
von: Singh, Prashant, et al.
Veröffentlicht: (2023)
von: Singh, Prashant, et al.
Veröffentlicht: (2023)
Continual Reinforcement Learning for HVAC Systems Control: Integrating Hypernetworks and Transfer Learning
von: Bekal, Gautham Udayakumar, et al.
Veröffentlicht: (2025)
von: Bekal, Gautham Udayakumar, et al.
Veröffentlicht: (2025)
Continual Learning with Query-Only Attention
von: Bekal, Gautham, et al.
Veröffentlicht: (2025)
von: Bekal, Gautham, et al.
Veröffentlicht: (2025)
Regulating Cryptocurrency and Decentralized Finance for an Inclusive Economy
von: Muralidhar, Amrutha, et al.
Veröffentlicht: (2024)
von: Muralidhar, Amrutha, et al.
Veröffentlicht: (2024)
Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
TILES-2018 Sleep Benchmark Dataset: A Longitudinal Wearable Sleep Data Set of Hospital Workers for Modeling and Understanding Sleep Behaviors
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Fast Lifelong Adaptive Inverse Reinforcement Learning from Demonstrations
von: Chen, Letian, et al.
Veröffentlicht: (2022)
von: Chen, Letian, et al.
Veröffentlicht: (2022)
Reinforcement Learning as a Parsimonious Alternative to Prediction Cascades: A Case Study on Image Segmentation
von: Srikishan, Bharat, et al.
Veröffentlicht: (2024)
von: Srikishan, Bharat, et al.
Veröffentlicht: (2024)
On the surjunctivity and the Garden of Eden theorem for non-uniform cellular automata
von: Paturi, Katariina, et al.
Veröffentlicht: (2025)
von: Paturi, Katariina, et al.
Veröffentlicht: (2025)
PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024) -
DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2023) -
Accelerated Test-Time Scaling with Model-Free Speculative Sampling
von: Song, Woomin, et al.
Veröffentlicht: (2025) -
Zero-resource Speech Translation and Recognition with LLMs
von: Mundnich, Karel, et al.
Veröffentlicht: (2024) -
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
von: Huybrechts, Goeric, et al.
Veröffentlicht: (2025)