An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
Fuente:
arXiv
Salvato in:
| Autori principali: | Inoue, Koji, Lala, Divesh, Elmers, Mikey, Ochi, Keiko, Kawahara, Tatsuya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prompt-Guided Turn-Taking Prediction
di: Inoue, Koji, et al.
Pubblicazione: (2025)
di: Inoue, Koji, et al.
Pubblicazione: (2025)
Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
di: Kato, Kazushi, et al.
Pubblicazione: (2025)
di: Kato, Kazushi, et al.
Pubblicazione: (2025)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
di: Inoue, Koji, et al.
Pubblicazione: (2024)
di: Inoue, Koji, et al.
Pubblicazione: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
di: Shi, Hao, et al.
Pubblicazione: (2024)
di: Shi, Hao, et al.
Pubblicazione: (2024)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
di: Ko, Yuka, et al.
Pubblicazione: (2024)
di: Ko, Yuka, et al.
Pubblicazione: (2024)
Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems
di: Elmers, Mikey, et al.
Pubblicazione: (2025)
di: Elmers, Mikey, et al.
Pubblicazione: (2025)
Exploration of Adapter for Noise Robust Automatic Speech Recognition
di: Shi, Hao, et al.
Pubblicazione: (2024)
di: Shi, Hao, et al.
Pubblicazione: (2024)
Why Do We Laugh? Annotation and Taxonomy Generation for Laughable Contexts in Spontaneous Text Conversation
di: Inoue, Koji, et al.
Pubblicazione: (2025)
di: Inoue, Koji, et al.
Pubblicazione: (2025)
Multilingual Turn-taking Prediction Using Voice Activity Projection
di: Inoue, Koji, et al.
Pubblicazione: (2024)
di: Inoue, Koji, et al.
Pubblicazione: (2024)
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
di: Deng, Yayue, et al.
Pubblicazione: (2025)
di: Deng, Yayue, et al.
Pubblicazione: (2025)
KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
di: Sharma, Rohan, et al.
Pubblicazione: (2025)
di: Sharma, Rohan, et al.
Pubblicazione: (2025)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
di: Gedeon, Máté, et al.
Pubblicazione: (2025)
di: Gedeon, Máté, et al.
Pubblicazione: (2025)
Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
di: Thebaud, Thomas, et al.
Pubblicazione: (2025)
di: Thebaud, Thomas, et al.
Pubblicazione: (2025)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
Efficient Streaming LLM for Speech Recognition
di: Jia, Junteng, et al.
Pubblicazione: (2024)
di: Jia, Junteng, et al.
Pubblicazione: (2024)
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
di: Zhao, Qiuming, et al.
Pubblicazione: (2025)
di: Zhao, Qiuming, et al.
Pubblicazione: (2025)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
di: He, Xinlu, et al.
Pubblicazione: (2025)
di: He, Xinlu, et al.
Pubblicazione: (2025)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
di: Sakshi, S, et al.
Pubblicazione: (2024)
di: Sakshi, S, et al.
Pubblicazione: (2024)
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
di: Wang, Xiong, et al.
Pubblicazione: (2024)
di: Wang, Xiong, et al.
Pubblicazione: (2024)
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
di: Shi, Hao, et al.
Pubblicazione: (2025)
di: Shi, Hao, et al.
Pubblicazione: (2025)
Benchmarking Automatic Speech Recognition for Indian Languages in Agricultural Contexts
di: S, Chandrashekar M, et al.
Pubblicazione: (2026)
di: S, Chandrashekar M, et al.
Pubblicazione: (2026)
Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
di: Khaertdinov, Bulat, et al.
Pubblicazione: (2024)
di: Khaertdinov, Bulat, et al.
Pubblicazione: (2024)
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
di: Niu, Yadong, et al.
Pubblicazione: (2025)
di: Niu, Yadong, et al.
Pubblicazione: (2025)
CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges
di: Li, Hui, et al.
Pubblicazione: (2025)
di: Li, Hui, et al.
Pubblicazione: (2025)
CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models
di: He, Jiajun, et al.
Pubblicazione: (2025)
di: He, Jiajun, et al.
Pubblicazione: (2025)
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
di: Honda, Tomoki, et al.
Pubblicazione: (2024)
di: Honda, Tomoki, et al.
Pubblicazione: (2024)
VoiceBench: Benchmarking LLM-Based Voice Assistants
di: Chen, Yiming, et al.
Pubblicazione: (2024)
di: Chen, Yiming, et al.
Pubblicazione: (2024)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
Real-Time Textless Dialogue Generation
di: Mai, Long, et al.
Pubblicazione: (2025)
di: Mai, Long, et al.
Pubblicazione: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
Analysis and Detection of Differences in Spoken User Behaviors between Autonomous and Wizard-of-Oz Systems
di: Elmers, Mikey, et al.
Pubblicazione: (2024)
di: Elmers, Mikey, et al.
Pubblicazione: (2024)
HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
di: Dutta, Soumya, et al.
Pubblicazione: (2023)
di: Dutta, Soumya, et al.
Pubblicazione: (2023)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
di: Inoue, Koji, et al.
Pubblicazione: (2024)
di: Inoue, Koji, et al.
Pubblicazione: (2024)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
di: Seide, Frank, et al.
Pubblicazione: (2024)
di: Seide, Frank, et al.
Pubblicazione: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2025)
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Prompt-Guided Turn-Taking Prediction
di: Inoue, Koji, et al.
Pubblicazione: (2025) -
Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
di: Kato, Kazushi, et al.
Pubblicazione: (2025) -
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
di: Inoue, Koji, et al.
Pubblicazione: (2024) -
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
di: Shi, Hao, et al.
Pubblicazione: (2024) -
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
di: Ko, Yuka, et al.
Pubblicazione: (2024)