Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
Fuente:
arXiv
Saved in:
| Main Authors: | Ahn, Jaewoo, Yun, Heeseung, Ko, Dayoon, Kim, Gunhee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?
by: Ko, Dayoon, et al.
Published: (2024)
by: Ko, Dayoon, et al.
Published: (2024)
Can Language Models Laugh at YouTube Short-form Videos?
by: Ko, Dayoon, et al.
Published: (2023)
by: Ko, Dayoon, et al.
Published: (2023)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
DynamicER: Resolving Emerging Mentions to Dynamic Entities for RAG
by: Kim, Jinyoung, et al.
Published: (2024)
by: Kim, Jinyoung, et al.
Published: (2024)
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
by: Yuhang, Yang, et al.
Published: (2024)
by: Yuhang, Yang, et al.
Published: (2024)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
by: Zhang, Pei, et al.
Published: (2025)
by: Zhang, Pei, et al.
Published: (2025)
Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
by: Yoo, HaeJun, et al.
Published: (2026)
by: Yoo, HaeJun, et al.
Published: (2026)
Spoken Language Identification with Pre-trained Models and Margin Loss
by: Fang, Zhihua, et al.
Published: (2026)
by: Fang, Zhihua, et al.
Published: (2026)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
by: Zeng, Aohan, et al.
Published: (2024)
by: Zeng, Aohan, et al.
Published: (2024)
Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
by: Tang, Yun, et al.
Published: (2025)
by: Tang, Yun, et al.
Published: (2025)
On the Cross-lingual Transferability of Pre-trained wav2vec2-based Models
by: Grosman, Jonatas, et al.
Published: (2025)
by: Grosman, Jonatas, et al.
Published: (2025)
ViSAGe: Video-to-Spatial Audio Generation
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
by: Lee, Dongwook, et al.
Published: (2026)
by: Lee, Dongwook, et al.
Published: (2026)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
by: Liu, Alexander H., et al.
Published: (2025)
by: Liu, Alexander H., et al.
Published: (2025)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
by: Lin, Tzu-Quan, et al.
Published: (2025)
by: Lin, Tzu-Quan, et al.
Published: (2025)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
by: Lim, Junyoung, et al.
Published: (2025)
by: Lim, Junyoung, et al.
Published: (2025)
Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark
by: Turetzky, Arnon, et al.
Published: (2026)
by: Turetzky, Arnon, et al.
Published: (2026)
PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech
by: Menta, Venkata Pushpak Teja
Published: (2026)
by: Menta, Venkata Pushpak Teja
Published: (2026)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
by: Pareras, Oriol, et al.
Published: (2025)
by: Pareras, Oriol, et al.
Published: (2025)
When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIR
by: Ko, Dayoon, et al.
Published: (2025)
by: Ko, Dayoon, et al.
Published: (2025)
Is a Peeled Apple Still Red? Evaluating LLMs' Ability for Conceptual Combination with Property Type
by: Song, Seokwon, et al.
Published: (2025)
by: Song, Seokwon, et al.
Published: (2025)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
by: Mousavi, Pooneh, et al.
Published: (2025)
by: Mousavi, Pooneh, et al.
Published: (2025)
BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
by: Wang, Chen, et al.
Published: (2023)
by: Wang, Chen, et al.
Published: (2023)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
by: Wang, Hsuan-Fu, et al.
Published: (2024)
by: Wang, Hsuan-Fu, et al.
Published: (2024)
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
by: Kim, Taesoo, et al.
Published: (2025)
by: Kim, Taesoo, et al.
Published: (2025)
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
by: Tang, Yun, et al.
Published: (2025)
by: Tang, Yun, et al.
Published: (2025)
Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss
by: Liu, Meizhu, et al.
Published: (2026)
by: Liu, Meizhu, et al.
Published: (2026)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
by: Zhao, Mengjie, et al.
Published: (2026)
by: Zhao, Mengjie, et al.
Published: (2026)
Soundwave: Less is More for Speech-Text Alignment in LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
by: Zhu, Yongxin, et al.
Published: (2024)
by: Zhu, Yongxin, et al.
Published: (2024)
Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
by: Dong, Lukuan, et al.
Published: (2024)
by: Dong, Lukuan, et al.
Published: (2024)
PRiSM: Benchmarking Phone Realization in Speech Models
by: Bharadwaj, Shikhar, et al.
Published: (2026)
by: Bharadwaj, Shikhar, et al.
Published: (2026)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
by: An, Keyu, et al.
Published: (2024)
by: An, Keyu, et al.
Published: (2024)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
by: Ma, Rao, et al.
Published: (2025)
by: Ma, Rao, et al.
Published: (2025)
Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR
by: Lee, Minsik, et al.
Published: (2026)
by: Lee, Minsik, et al.
Published: (2026)
GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
by: Gao, Yingying, et al.
Published: (2024)
by: Gao, Yingying, et al.
Published: (2024)
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
by: Peri, Raghuveer, et al.
Published: (2024)
by: Peri, Raghuveer, et al.
Published: (2024)
Massive Sound Embedding Benchmark (MSEB)
by: Heigold, Georg, et al.
Published: (2026)
by: Heigold, Georg, et al.
Published: (2026)
MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
by: Wu, Shih-Lun, et al.
Published: (2025)
by: Wu, Shih-Lun, et al.
Published: (2025)
Similar Items
-
GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?
by: Ko, Dayoon, et al.
Published: (2024) -
Can Language Models Laugh at YouTube Short-form Videos?
by: Ko, Dayoon, et al.
Published: (2023) -
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
by: Ahn, Jaewoo, et al.
Published: (2025) -
DynamicER: Resolving Emerging Mentions to Dynamic Entities for RAG
by: Kim, Jinyoung, et al.
Published: (2024) -
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
by: Yuhang, Yang, et al.
Published: (2024)