CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Xi, Guo, Ying, Zhen, Cheng, Li, Tong, Ao, Yingying, Yan, Pengfei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
di: Ghosh, Bishal, et al.
Pubblicazione: (2024)
di: Ghosh, Bishal, et al.
Pubblicazione: (2024)
Listen, Think, and Understand
di: Gong, Yuan, et al.
Pubblicazione: (2023)
di: Gong, Yuan, et al.
Pubblicazione: (2023)
RF-GML: Reference-Free Generative Machine Listener
di: Biswas, Arijit, et al.
Pubblicazione: (2024)
di: Biswas, Arijit, et al.
Pubblicazione: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
A Multi-loudspeaker Binaural Room Impulse Response Dataset with High-Resolution Translational and Rotational Head Coordinates in a Listening Room
di: Qiao, Yue, et al.
Pubblicazione: (2024)
di: Qiao, Yue, et al.
Pubblicazione: (2024)
Requirements for Mass Adoption of Assistive Listening Technology by the General Public
di: Kaufmann, Thomas B., et al.
Pubblicazione: (2023)
di: Kaufmann, Thomas B., et al.
Pubblicazione: (2023)
DIFFA: Large Language Diffusion Models Can Listen and Understand
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
di: Chung, Soo-Whan, et al.
Pubblicazione: (2025)
di: Chung, Soo-Whan, et al.
Pubblicazione: (2025)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
di: Jeong, Jihoon, et al.
Pubblicazione: (2026)
di: Jeong, Jihoon, et al.
Pubblicazione: (2026)
Evaluating Speech Enhancement Systems Through Listening Effort
di: Gelderblom, Femke B., et al.
Pubblicazione: (2024)
di: Gelderblom, Femke B., et al.
Pubblicazione: (2024)
Listen to Extract: Onset-Prompted Target Speaker Extraction
di: Shen, Pengjie, et al.
Pubblicazione: (2025)
di: Shen, Pengjie, et al.
Pubblicazione: (2025)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
di: Raj, Desh
Pubblicazione: (2024)
di: Raj, Desh
Pubblicazione: (2024)
Reproducing the Acoustic Velocity Vectors in a Circular Listening Area
di: Wang, Jiarui, et al.
Pubblicazione: (2024)
di: Wang, Jiarui, et al.
Pubblicazione: (2024)
Joint Minimum Processing Beamforming and Near-end Listening Enhancement
di: Fuglsig, Andreas J., et al.
Pubblicazione: (2023)
di: Fuglsig, Andreas J., et al.
Pubblicazione: (2023)
Listening broadband physical model for microphones: a first step
di: Millot, Laurent, et al.
Pubblicazione: (2024)
di: Millot, Laurent, et al.
Pubblicazione: (2024)
Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
di: Redondo, Rafael
Pubblicazione: (2024)
di: Redondo, Rafael
Pubblicazione: (2024)
What Do Neurons Listen To? A Neuron-level Dissection of a General-purpose Audio Model
di: Kawamura, Takao, et al.
Pubblicazione: (2026)
di: Kawamura, Takao, et al.
Pubblicazione: (2026)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
di: Hu, Cheng-Hung, et al.
Pubblicazione: (2025)
di: Hu, Cheng-Hung, et al.
Pubblicazione: (2025)
Learning How to Listen: A Temporal-Frequential Attention Model for Sound Event Detection
di: Shen, Yu-Han, et al.
Pubblicazione: (2018)
di: Shen, Yu-Han, et al.
Pubblicazione: (2018)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
di: Yamamoto, Katsuhiko, et al.
Pubblicazione: (2025)
di: Yamamoto, Katsuhiko, et al.
Pubblicazione: (2025)
Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience
di: Jiang, Xilin, et al.
Pubblicazione: (2024)
di: Jiang, Xilin, et al.
Pubblicazione: (2024)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
di: Rahimi, Akam, et al.
Pubblicazione: (2025)
di: Rahimi, Akam, et al.
Pubblicazione: (2025)
Listenable Maps for Audio Classifiers
di: Paissan, Francesco, et al.
Pubblicazione: (2024)
di: Paissan, Francesco, et al.
Pubblicazione: (2024)
Listening Between the Lines: Synthetic Speech Detection Disregarding Verbal Content
di: Salvi, Davide, et al.
Pubblicazione: (2024)
di: Salvi, Davide, et al.
Pubblicazione: (2024)
Can Masked Autoencoders Also Listen to Birds?
di: Rauch, Lukas, et al.
Pubblicazione: (2025)
di: Rauch, Lukas, et al.
Pubblicazione: (2025)
Deep CLAS: Deep Contextual Listen, Attend and Spell
di: Wang, Mengzhi, et al.
Pubblicazione: (2024)
di: Wang, Mengzhi, et al.
Pubblicazione: (2024)
Listen, Analyze, and Adapt to Learn New Attacks: An Exemplar-Free Class Incremental Learning Method for Audio Deepfake Source Tracing
di: Xiao, Yang, et al.
Pubblicazione: (2025)
di: Xiao, Yang, et al.
Pubblicazione: (2025)
Development of the Listening in Spatialized Noise-Sentences (LiSN-S) Test in Brazilian Portuguese: Presentation Software, Speech Stimuli, and Sentence Equivalence
di: Masiero, Bruno S., et al.
Pubblicazione: (2024)
di: Masiero, Bruno S., et al.
Pubblicazione: (2024)
Spatial Analysis and Synthesis Methods: Subjective and Objective Evaluations Using Various Microphone Arrays in the Auralization of a Critical Listening Room
di: Pawlak, Alan, et al.
Pubblicazione: (2024)
di: Pawlak, Alan, et al.
Pubblicazione: (2024)
A Convolutional Framework for Mapping Imagined Auditory MEG into Listened Brain Responses
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2025)
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2025)
OCR-Enhanced Multimodal ASR Can Read While Listening
di: Chen, Junli, et al.
Pubblicazione: (2026)
di: Chen, Junli, et al.
Pubblicazione: (2026)
Listenable Maps for Zero-Shot Audio Classifiers
di: Paissan, Francesco, et al.
Pubblicazione: (2024)
di: Paissan, Francesco, et al.
Pubblicazione: (2024)
Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
di: Kato, Kazushi, et al.
Pubblicazione: (2025)
di: Kato, Kazushi, et al.
Pubblicazione: (2025)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
di: Liu, Rui, et al.
Pubblicazione: (2025)
di: Liu, Rui, et al.
Pubblicazione: (2025)
Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners
di: Cao, Boxuan, et al.
Pubblicazione: (2025)
di: Cao, Boxuan, et al.
Pubblicazione: (2025)
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication
di: Huang, Jinhe, et al.
Pubblicazione: (2025)
di: Huang, Jinhe, et al.
Pubblicazione: (2025)
LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
di: Zou, Wenhao, et al.
Pubblicazione: (2026)
di: Zou, Wenhao, et al.
Pubblicazione: (2026)
ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
Listening for Expert Identified Linguistic Features: Assessment of Audio Deepfake Discernment among Undergraduate Students
di: Bhalli, Noshaba N., et al.
Pubblicazione: (2024)
di: Bhalli, Noshaba N., et al.
Pubblicazione: (2024)
Documenti analoghi
-
Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
di: Ghosh, Bishal, et al.
Pubblicazione: (2024) -
Listen, Think, and Understand
di: Gong, Yuan, et al.
Pubblicazione: (2023) -
RF-GML: Reference-Free Generative Machine Listener
di: Biswas, Arijit, et al.
Pubblicazione: (2024) -
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
di: Chen, Zhengyang, et al.
Pubblicazione: (2024) -
A Multi-loudspeaker Binaural Room Impulse Response Dataset with High-Resolution Translational and Rotational Head Coordinates in a Listening Room
di: Qiao, Yue, et al.
Pubblicazione: (2024)