Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Hori, Chiori, Masuyama, Yoshiki, Jain, Siddarth, Corcodel, Radu, Jha, Devesh, Romeres, Diego, Roux, Jonathan Le |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
by: Masuyama, Yoshiki, et al.
Published: (2026)
by: Masuyama, Yoshiki, et al.
Published: (2026)
NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization
by: Masuyama, Yoshiki, et al.
Published: (2024)
by: Masuyama, Yoshiki, et al.
Published: (2024)
Factorized RVQ-GAN For Disentangled Speech Tokenization
by: Khurana, Sameer, et al.
Published: (2025)
by: Khurana, Sameer, et al.
Published: (2025)
Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization
by: Masuyama, Yoshiki, et al.
Published: (2025)
by: Masuyama, Yoshiki, et al.
Published: (2025)
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
by: Masuyama, Yoshiki, et al.
Published: (2025)
by: Masuyama, Yoshiki, et al.
Published: (2025)
Exploring the Capability of Mamba in Speech Applications
by: Miyazaki, Koichi, et al.
Published: (2024)
by: Miyazaki, Koichi, et al.
Published: (2024)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
by: Masuyama, Yoshiki, et al.
Published: (2024)
by: Masuyama, Yoshiki, et al.
Published: (2024)
Physics-Informed Direction-Aware Neural Acoustic Fields
by: Masuyama, Yoshiki, et al.
Published: (2025)
by: Masuyama, Yoshiki, et al.
Published: (2025)
Direction-Aware Neural Acoustic Fields for Few-Shot Interpolation of Ambisonic Impulse Responses
by: Ick, Christopher, et al.
Published: (2025)
by: Ick, Christopher, et al.
Published: (2025)
Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training
by: Ick, Christopher, et al.
Published: (2025)
by: Ick, Christopher, et al.
Published: (2025)
Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
by: Cornell, Samuele, et al.
Published: (2025)
by: Cornell, Samuele, et al.
Published: (2025)
FasTUSS: Faster Task-Aware Unified Source Separation
by: Paissan, Francesco, et al.
Published: (2025)
by: Paissan, Francesco, et al.
Published: (2025)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
by: Tian, Jinchuan, et al.
Published: (2025)
by: Tian, Jinchuan, et al.
Published: (2025)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
by: Yang, Yiqing, et al.
Published: (2025)
by: Yang, Yiqing, et al.
Published: (2025)
RoDia: A New Dataset for Romanian Dialect Identification from Speech
by: Rotaru, Codrut, et al.
Published: (2023)
by: Rotaru, Codrut, et al.
Published: (2023)
Can Whisper perform speech-based in-context learning?
by: Wang, Siyin, et al.
Published: (2023)
by: Wang, Siyin, et al.
Published: (2023)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
by: Le, Khanh, et al.
Published: (2025)
by: Le, Khanh, et al.
Published: (2025)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
by: Masuyama, Yoshiki, et al.
Published: (2025)
by: Masuyama, Yoshiki, et al.
Published: (2025)
Self-consistent context aware conformer transducer for speech recognition
by: Kolokolov, Konstantin, et al.
Published: (2024)
by: Kolokolov, Konstantin, et al.
Published: (2024)
Borderless Long Speech Synthesis
by: Song, Xingchen, et al.
Published: (2026)
by: Song, Xingchen, et al.
Published: (2026)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
by: Liu, Oli Danyi, et al.
Published: (2024)
by: Liu, Oli Danyi, et al.
Published: (2024)
Mmm whatcha say? Uncovering distal and proximal context effects in first and second-language word perception using psychophysical reverse correlation
by: Tuttösí, Paige, et al.
Published: (2024)
by: Tuttösí, Paige, et al.
Published: (2024)
EMMeTT: Efficient Multimodal Machine Translation Training
by: Żelasko, Piotr, et al.
Published: (2024)
by: Żelasko, Piotr, et al.
Published: (2024)
CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
by: Munakata, Hokuto, et al.
Published: (2025)
by: Munakata, Hokuto, et al.
Published: (2025)
Lightweight Audio Segmentation for Long-form Speech Translation
by: Lee, Jaesong, et al.
Published: (2024)
by: Lee, Jaesong, et al.
Published: (2024)
Long-Form Speech Generation with Spoken Language Models
by: Park, Se Jin, et al.
Published: (2024)
by: Park, Se Jin, et al.
Published: (2024)
Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
by: Hu, Jiliang, et al.
Published: (2024)
by: Hu, Jiliang, et al.
Published: (2024)
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
OCR-Enhanced Multimodal ASR Can Read While Listening
by: Chen, Junli, et al.
Published: (2026)
by: Chen, Junli, et al.
Published: (2026)
Multimodal Input Aids a Bayesian Model of Phonetic Learning
by: Zhi, Sophia, et al.
Published: (2024)
by: Zhi, Sophia, et al.
Published: (2024)
Guided by the Plan: Enhancing Faithful Autoregressive Text-to-Audio Generation with Guided Decoding
by: Wang, Juncheng, et al.
Published: (2026)
by: Wang, Juncheng, et al.
Published: (2026)
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
by: Peri, Raghuveer, et al.
Published: (2024)
by: Peri, Raghuveer, et al.
Published: (2024)
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation
by: Lei, Ligong, et al.
Published: (2026)
by: Lei, Ligong, et al.
Published: (2026)
GatedxLSTM: A Multimodal Affective Computing Approach for Emotion Recognition in Conversations
by: Li, Yupei, et al.
Published: (2025)
by: Li, Yupei, et al.
Published: (2025)
Multimodal Magic Elevating Depression Detection with a Fusion of Text and Audio Intelligence
by: Gan, Lindy, et al.
Published: (2025)
by: Gan, Lindy, et al.
Published: (2025)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
by: Wang, Xuechen, et al.
Published: (2024)
by: Wang, Xuechen, et al.
Published: (2024)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
by: Mousavi, Pooneh, et al.
Published: (2025)
by: Mousavi, Pooneh, et al.
Published: (2025)
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Similar Items
-
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
by: Masuyama, Yoshiki, et al.
Published: (2026) -
NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization
by: Masuyama, Yoshiki, et al.
Published: (2024) -
Factorized RVQ-GAN For Disentangled Speech Tokenization
by: Khurana, Sameer, et al.
Published: (2025) -
Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization
by: Masuyama, Yoshiki, et al.
Published: (2025) -
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
by: Masuyama, Yoshiki, et al.
Published: (2025)