Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Heeseung, Seo, Soonshin, Jeong, Kyeongseok, Kwon, Ohsung, Kim, Soyoon, Kim, Jungwhan, Lee, Jaehong, Song, Eunwoo, Oh, Myungwoo, Ha, Jung-Woo, Yoon, Sungroh, Yoo, Kang Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
von: Lee, Dongwook, et al.
Veröffentlicht: (2026)
von: Lee, Dongwook, et al.
Veröffentlicht: (2026)
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
von: Shin, Chaehun, et al.
Veröffentlicht: (2024)
von: Shin, Chaehun, et al.
Veröffentlicht: (2024)
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
EdiText: Controllable Coarse-to-Fine Text Editing with Diffusion Language Models
von: Lee, Che Hyun, et al.
Veröffentlicht: (2025)
von: Lee, Che Hyun, et al.
Veröffentlicht: (2025)
Style-Friendly SNR Sampler for Style-Driven Generation
von: Choi, Jooyoung, et al.
Veröffentlicht: (2024)
von: Choi, Jooyoung, et al.
Veröffentlicht: (2024)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
von: Yeom, Jiheum, et al.
Veröffentlicht: (2024)
von: Yeom, Jiheum, et al.
Veröffentlicht: (2024)
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2025)
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
von: Kim, Heeseung, et al.
Veröffentlicht: (2025)
von: Kim, Heeseung, et al.
Veröffentlicht: (2025)
Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
Learning to Assemble the Soma Cube with Legal-Action Masked DQN and Safe ZYZ Regrasp on a Doosan M0609
von: Oh, Jaehong, et al.
Veröffentlicht: (2025)
von: Oh, Jaehong, et al.
Veröffentlicht: (2025)
SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
Lightweight Audio Segmentation for Long-form Speech Translation
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
Learning Dexterous Bimanual Catch Skills through Adversarial-Cooperative Heterogeneous-Agent Reinforcement Learning
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
RT-HDIST: Ray-Tracing Core-based Hausdorff Distance Computation
von: Kim, YoungWoo, et al.
Veröffentlicht: (2025)
von: Kim, YoungWoo, et al.
Veröffentlicht: (2025)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
von: Yoon, Sion, et al.
Veröffentlicht: (2024)
von: Yoon, Sion, et al.
Veröffentlicht: (2024)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
What Linear Probes Miss: Multi-View Probing for Weight-Space Learning
von: Heo, Eunwoo, et al.
Veröffentlicht: (2026)
von: Heo, Eunwoo, et al.
Veröffentlicht: (2026)
Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
von: Jang, Jinhyeok, et al.
Veröffentlicht: (2025)
von: Jang, Jinhyeok, et al.
Veröffentlicht: (2025)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
von: Park, Sangha, et al.
Veröffentlicht: (2025)
von: Park, Sangha, et al.
Veröffentlicht: (2025)
Selectivity of Mannitol–Yolk–Polymyxin B Agar When Supplemented With Cefsulodin for Enumeration of Bacillus cereus in Food
von: Kun‐Ho Seo, et al.
Veröffentlicht: (2025)
von: Kun‐Ho Seo, et al.
Veröffentlicht: (2025)
Battling the Non-stationarity in Time Series Forecasting via Test-time Adaptation
von: Kim, HyunGi, et al.
Veröffentlicht: (2025)
von: Kim, HyunGi, et al.
Veröffentlicht: (2025)
Language as Cost: Proactive Hazard Mapping using VLM for Robot Navigation
von: Oh, Mintaek, et al.
Veröffentlicht: (2025)
von: Oh, Mintaek, et al.
Veröffentlicht: (2025)
DISPATCH: Distilling Selective Patches for Speech Enhancement
von: Kim, Dohwan, et al.
Veröffentlicht: (2025)
von: Kim, Dohwan, et al.
Veröffentlicht: (2025)
Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning
von: Choi, Jungwon, et al.
Veröffentlicht: (2026)
von: Choi, Jungwon, et al.
Veröffentlicht: (2026)
Holographic Tests for Giant Graviton Expansion
von: Kim, Seunggyu, et al.
Veröffentlicht: (2024)
von: Kim, Seunggyu, et al.
Veröffentlicht: (2024)
Continual Learning for Multiple Modalities
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
Stake the Points: Structure-Faithful Instance Unlearning
von: Hong, Kiseong, et al.
Veröffentlicht: (2026)
von: Hong, Kiseong, et al.
Veröffentlicht: (2026)
Generative Modeling of Class Probability for Multi-Modal Representation Learning
von: Shin, Jungkyoo, et al.
Veröffentlicht: (2025)
von: Shin, Jungkyoo, et al.
Veröffentlicht: (2025)
Empowering Personalized Learning through a Conversation-based Tutoring System with Student Modeling
von: Park, Minju, et al.
Veröffentlicht: (2024)
von: Park, Minju, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Deep Learning for Time Series Forecasting: Architectural Diversity and Open Challenges
von: Kim, Jongseon, et al.
Veröffentlicht: (2024)
von: Kim, Jongseon, et al.
Veröffentlicht: (2024)
ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning
von: Choi, Jae-Woo, et al.
Veröffentlicht: (2025)
von: Choi, Jae-Woo, et al.
Veröffentlicht: (2025)
A Large-Depth-Range Layer-Based Hologram Dataset for Machine Learning-Based 3D Computer-Generated Holography
von: Lee, Jaehong, et al.
Veröffentlicht: (2025)
von: Lee, Jaehong, et al.
Veröffentlicht: (2025)
ViSAGe: Video-to-Spatial Audio Generation
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
Exploring High-Order Self-Similarity for Video Understanding
von: Kim, Manjin, et al.
Veröffentlicht: (2026)
von: Kim, Manjin, et al.
Veröffentlicht: (2026)
ARES: Auxiliary Range Expansion for Outlier Synthesis
von: Jung, Eui-Soo, et al.
Veröffentlicht: (2025)
von: Jung, Eui-Soo, et al.
Veröffentlicht: (2025)
Exact QFT duals of AdS black holes
von: Choi, Sunjin, et al.
Veröffentlicht: (2021)
von: Choi, Sunjin, et al.
Veröffentlicht: (2021)
Maternal wellbeing amidst English fever: An integrative framework of vicarious pride, empathy, agency and hope (M–VEAH)
von: Yeji Han, et al.
Veröffentlicht: (2026)
von: Yeji Han, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
von: Lee, Dongwook, et al.
Veröffentlicht: (2026) -
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026) -
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
von: Shin, Chaehun, et al.
Veröffentlicht: (2024) -
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
von: Kim, Heeseung, et al.
Veröffentlicht: (2024) -
EdiText: Controllable Coarse-to-Fine Text Editing with Diffusion Language Models
von: Lee, Che Hyun, et al.
Veröffentlicht: (2025)