Optimizing the role of human evaluation in LLM-based spoken document summarization systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kroll, Margaret, Kraus, Kelsey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Out-of-distribution generalisation in spoken language understanding
von: Porjazovski, Dejan, et al.
Veröffentlicht: (2024)
von: Porjazovski, Dejan, et al.
Veröffentlicht: (2024)
Iterative LLM-based improvement for French Clinical Interview Transcription and Speaker Diarization
von: Marie, Ambre, et al.
Veröffentlicht: (2026)
von: Marie, Ambre, et al.
Veröffentlicht: (2026)
Neural networks for Text-to-Speech evaluation
von: Trofimenko, Ilya, et al.
Veröffentlicht: (2026)
von: Trofimenko, Ilya, et al.
Veröffentlicht: (2026)
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
von: Song, Zheshu, et al.
Veröffentlicht: (2024)
von: Song, Zheshu, et al.
Veröffentlicht: (2024)
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
von: Fang, Qingkai, et al.
Veröffentlicht: (2025)
von: Fang, Qingkai, et al.
Veröffentlicht: (2025)
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
von: Majumder, Navonil, et al.
Veröffentlicht: (2024)
von: Majumder, Navonil, et al.
Veröffentlicht: (2024)
Differentiable Reward Optimization for LLM based TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
von: Kumar, Shashi, et al.
Veröffentlicht: (2024)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
LLM-Driven Multimodal Opinion Expression Identification
von: Jia, Bonian, et al.
Veröffentlicht: (2024)
von: Jia, Bonian, et al.
Veröffentlicht: (2024)
VoiceBench: Benchmarking LLM-Based Voice Assistants
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
Speech Prefix-Tuning with RNNT Loss for Improving LLM Predictions
von: Baskar, Murali Karthick, et al.
Veröffentlicht: (2024)
von: Baskar, Murali Karthick, et al.
Veröffentlicht: (2024)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
von: Inoue, Koji, et al.
Veröffentlicht: (2025)
von: Inoue, Koji, et al.
Veröffentlicht: (2025)
Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
von: Thebaud, Thomas, et al.
Veröffentlicht: (2025)
von: Thebaud, Thomas, et al.
Veröffentlicht: (2025)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models
von: He, Jiajun, et al.
Veröffentlicht: (2025)
von: He, Jiajun, et al.
Veröffentlicht: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
von: Yu, Wenyi, et al.
Veröffentlicht: (2024)
von: Yu, Wenyi, et al.
Veröffentlicht: (2024)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
von: Sharma, Rohan, et al.
Veröffentlicht: (2025)
von: Sharma, Rohan, et al.
Veröffentlicht: (2025)
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
von: Wang, Xiong, et al.
Veröffentlicht: (2024)
von: Wang, Xiong, et al.
Veröffentlicht: (2024)
Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
von: Manco, Ilaria, et al.
Veröffentlicht: (2024)
von: Manco, Ilaria, et al.
Veröffentlicht: (2024)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
von: Seide, Frank, et al.
Veröffentlicht: (2024)
von: Seide, Frank, et al.
Veröffentlicht: (2024)
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
Optimizing the Songwriting Process: Genre-Based Lyric Generation Using Deep Learning Models
von: Cai, Tracy, et al.
Veröffentlicht: (2024)
von: Cai, Tracy, et al.
Veröffentlicht: (2024)
Using LLM for Real-Time Transcription and Summarization of Doctor-Patient Interactions into ePuskesmas in Indonesia: A Proof-of-Concept Study
von: Khatim, Nur Ahmad, et al.
Veröffentlicht: (2024)
von: Khatim, Nur Ahmad, et al.
Veröffentlicht: (2024)
Optimizing Automatic Speech Assessment: W-RankSim Regularization and Hybrid Feature Fusion Strategies
von: Wu, Chung-Wen, et al.
Veröffentlicht: (2024)
von: Wu, Chung-Wen, et al.
Veröffentlicht: (2024)
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
Audio Contrastive-based Fine-tuning: Decoupling Representation Learning and Classification
von: Wang, Yang, et al.
Veröffentlicht: (2023)
von: Wang, Yang, et al.
Veröffentlicht: (2023)
Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Improving Low-Resource Dialect Classification Using Retrieval-based Voice Conversion
von: Fischbach, Lea, et al.
Veröffentlicht: (2025)
von: Fischbach, Lea, et al.
Veröffentlicht: (2025)
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
von: Papi, Sara, et al.
Veröffentlicht: (2024)
von: Papi, Sara, et al.
Veröffentlicht: (2024)
Swin-BERT: A Feature Fusion System designed for Speech-based Alzheimer's Dementia Detection
von: Pan, Yilin, et al.
Veröffentlicht: (2024)
von: Pan, Yilin, et al.
Veröffentlicht: (2024)
Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR
von: Lu, Xugang, et al.
Veröffentlicht: (2024)
von: Lu, Xugang, et al.
Veröffentlicht: (2024)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Out-of-distribution generalisation in spoken language understanding
von: Porjazovski, Dejan, et al.
Veröffentlicht: (2024) -
Iterative LLM-based improvement for French Clinical Interview Transcription and Speaker Diarization
von: Marie, Ambre, et al.
Veröffentlicht: (2026) -
Neural networks for Text-to-Speech evaluation
von: Trofimenko, Ilya, et al.
Veröffentlicht: (2026) -
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
von: Song, Zheshu, et al.
Veröffentlicht: (2024) -
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
von: Fang, Qingkai, et al.
Veröffentlicht: (2025)