Listen, Think, and Understand
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Gong, Yuan, Luo, Hongyin, Liu, Alexander H., Karlinsky, Leonid, Glass, James |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
par: Rouditchenko, Andrew, et autres
Publié: (2024)
par: Rouditchenko, Andrew, et autres
Publié: (2024)
DIFFA: Large Language Diffusion Models Can Listen and Understand
par: Zhou, Jiaming, et autres
Publié: (2025)
par: Zhou, Jiaming, et autres
Publié: (2025)
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
par: Liu, Alexander H., et autres
Publié: (2024)
par: Liu, Alexander H., et autres
Publié: (2024)
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
par: Liu, Alexander H., et autres
Publié: (2024)
par: Liu, Alexander H., et autres
Publié: (2024)
Listen to Extract: Onset-Prompted Target Speaker Extraction
par: Shen, Pengjie, et autres
Publié: (2025)
par: Shen, Pengjie, et autres
Publié: (2025)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
par: Bhati, Saurabhchand, et autres
Publié: (2024)
par: Bhati, Saurabhchand, et autres
Publié: (2024)
USAD: Universal Speech and Audio Representation via Distillation
par: Chang, Heng-Jui, et autres
Publié: (2025)
par: Chang, Heng-Jui, et autres
Publié: (2025)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
par: Araujo, Edson, et autres
Publié: (2025)
par: Araujo, Edson, et autres
Publié: (2025)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
par: Jeong, Jihoon, et autres
Publié: (2026)
par: Jeong, Jihoon, et autres
Publié: (2026)
Evaluating Speech Enhancement Systems Through Listening Effort
par: Gelderblom, Femke B., et autres
Publié: (2024)
par: Gelderblom, Femke B., et autres
Publié: (2024)
RF-GML: Reference-Free Generative Machine Listener
par: Biswas, Arijit, et autres
Publié: (2024)
par: Biswas, Arijit, et autres
Publié: (2024)
Requirements for Mass Adoption of Assistive Listening Technology by the General Public
par: Kaufmann, Thomas B., et autres
Publié: (2023)
par: Kaufmann, Thomas B., et autres
Publié: (2023)
Joint Minimum Processing Beamforming and Near-end Listening Enhancement
par: Fuglsig, Andreas J., et autres
Publié: (2023)
par: Fuglsig, Andreas J., et autres
Publié: (2023)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
par: Raj, Desh
Publié: (2024)
par: Raj, Desh
Publié: (2024)
Reproducing the Acoustic Velocity Vectors in a Circular Listening Area
par: Wang, Jiarui, et autres
Publié: (2024)
par: Wang, Jiarui, et autres
Publié: (2024)
Listening broadband physical model for microphones: a first step
par: Millot, Laurent, et autres
Publié: (2024)
par: Millot, Laurent, et autres
Publié: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
par: Chen, Zhengyang, et autres
Publié: (2024)
par: Chen, Zhengyang, et autres
Publié: (2024)
Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
par: Ghosh, Bishal, et autres
Publié: (2024)
par: Ghosh, Bishal, et autres
Publié: (2024)
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
par: Rouditchenko, Andrew, et autres
Publié: (2025)
par: Rouditchenko, Andrew, et autres
Publié: (2025)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
par: Chung, Soo-Whan, et autres
Publié: (2025)
par: Chung, Soo-Whan, et autres
Publié: (2025)
R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces
par: Chang, Heng-Jui, et autres
Publié: (2023)
par: Chang, Heng-Jui, et autres
Publié: (2023)
Learning How to Listen: A Temporal-Frequential Attention Model for Sound Event Detection
par: Shen, Yu-Han, et autres
Publié: (2018)
par: Shen, Yu-Han, et autres
Publié: (2018)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
par: Yamamoto, Katsuhiko, et autres
Publié: (2025)
par: Yamamoto, Katsuhiko, et autres
Publié: (2025)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
par: Liu, Alexander H., et autres
Publié: (2025)
par: Liu, Alexander H., et autres
Publié: (2025)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
par: Zhang, Yiru, et autres
Publié: (2025)
par: Zhang, Yiru, et autres
Publié: (2025)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
par: Wu, Haibin, et autres
Publié: (2024)
par: Wu, Haibin, et autres
Publié: (2024)
What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
par: Kawamura, Takao, et autres
Publié: (2026)
par: Kawamura, Takao, et autres
Publié: (2026)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
par: Hu, Cheng-Hung, et autres
Publié: (2025)
par: Hu, Cheng-Hung, et autres
Publié: (2025)
State-Space Large Audio Language Models
par: Bhati, Saurabhchand, et autres
Publié: (2024)
par: Bhati, Saurabhchand, et autres
Publié: (2024)
Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
par: Wang, Liming, et autres
Publié: (2024)
par: Wang, Liming, et autres
Publié: (2024)
Listen, Analyze, and Adapt to Learn New Attacks: An Exemplar-Free Class Incremental Learning Method for Audio Deepfake Source Tracing
par: Xiao, Yang, et autres
Publié: (2025)
par: Xiao, Yang, et autres
Publié: (2025)
Development of the Listening in Spatialized Noise-Sentences (LiSN-S) Test in Brazilian Portuguese: Presentation Software, Speech Stimuli, and Sentence Equivalence
par: Masiero, Bruno S., et autres
Publié: (2024)
par: Masiero, Bruno S., et autres
Publié: (2024)
Spatial Analysis and Synthesis Methods: Subjective and Objective Evaluations Using Various Microphone Arrays in the Auralization of a Critical Listening Room
par: Pawlak, Alan, et autres
Publié: (2024)
par: Pawlak, Alan, et autres
Publié: (2024)
A Multi-loudspeaker Binaural Room Impulse Response Dataset with High-Resolution Translational and Rotational Head Coordinates in a Listening Room
par: Qiao, Yue, et autres
Publié: (2024)
par: Qiao, Yue, et autres
Publié: (2024)
SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing
par: Hu, Jinbo, et autres
Publié: (2025)
par: Hu, Jinbo, et autres
Publié: (2025)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
par: Rahimi, Akam, et autres
Publié: (2025)
par: Rahimi, Akam, et autres
Publié: (2025)
LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
par: Zou, Wenhao, et autres
Publié: (2026)
par: Zou, Wenhao, et autres
Publié: (2026)
Listening Between the Lines: Synthetic Speech Detection Disregarding Verbal Content
par: Salvi, Davide, et autres
Publié: (2024)
par: Salvi, Davide, et autres
Publié: (2024)
Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
par: Redondo, Rafael
Publié: (2024)
par: Redondo, Rafael
Publié: (2024)
Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
par: Wu, Shangda, et autres
Publié: (2026)
par: Wu, Shangda, et autres
Publié: (2026)
Documents similaires
-
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
par: Rouditchenko, Andrew, et autres
Publié: (2024) -
DIFFA: Large Language Diffusion Models Can Listen and Understand
par: Zhou, Jiaming, et autres
Publié: (2025) -
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
par: Liu, Alexander H., et autres
Publié: (2024) -
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
par: Liu, Alexander H., et autres
Publié: (2024) -
Listen to Extract: Onset-Prompted Target Speaker Extraction
par: Shen, Pengjie, et autres
Publié: (2025)