Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ryan, Yuriel, Ip, Hei Man, Kuek, Adriel, Liang, Paul Pu, Lee, Roy Ka-Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
von: Shan, Xiaojun, et al.
Veröffentlicht: (2025)
von: Shan, Xiaojun, et al.
Veröffentlicht: (2025)
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
An Ethical Literary Criticism of Han Suyin’s Autobiography
von: Kuek, Florence
Veröffentlicht: (2025)
von: Kuek, Florence
Veröffentlicht: (2025)
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs
von: Zun, Lee Qi, et al.
Veröffentlicht: (2025)
von: Zun, Lee Qi, et al.
Veröffentlicht: (2025)
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
von: Roy, Shuvendu, et al.
Veröffentlicht: (2025)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2025)
On the Adversarial Robustness of Instruction-Tuned Large Language Models for Code
von: Hossen, Md Imran, et al.
Veröffentlicht: (2024)
von: Hossen, Md Imran, et al.
Veröffentlicht: (2024)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
von: Tan, Rui Yang, et al.
Veröffentlicht: (2026)
von: Tan, Rui Yang, et al.
Veröffentlicht: (2026)
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
von: Tan, Bryan Chen Zhengyu, et al.
Veröffentlicht: (2025)
von: Tan, Bryan Chen Zhengyu, et al.
Veröffentlicht: (2025)
Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving
von: Chen, Steven-Shine, et al.
Veröffentlicht: (2025)
von: Chen, Steven-Shine, et al.
Veröffentlicht: (2025)
Imperfect World Models are Exploitable
von: Bhamidipaty, Logan Mondal, et al.
Veröffentlicht: (2026)
von: Bhamidipaty, Logan Mondal, et al.
Veröffentlicht: (2026)
Users’ quality expectations and their correspondence with the realistic features of translation applications
von: Jing Rou Kuek
Veröffentlicht: (2024)
von: Jing Rou Kuek
Veröffentlicht: (2024)
Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions
von: Hee, Ming Shan, et al.
Veröffentlicht: (2025)
von: Hee, Ming Shan, et al.
Veröffentlicht: (2025)
Interpreting Bias in Large Language Models: A Feature-Based Approach
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2024)
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2024)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
A Vision for Multisensory Intelligence: Sensing, Science, and Synergy
von: Liang, Paul Pu
Veröffentlicht: (2026)
von: Liang, Paul Pu
Veröffentlicht: (2026)
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
AromaGen: Interactive Generation of Rich Olfactory Experiences with Multimodal Language Models
von: Wen, Yunge, et al.
Veröffentlicht: (2026)
von: Wen, Yunge, et al.
Veröffentlicht: (2026)
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
von: Xiao, Yunze, et al.
Veröffentlicht: (2024)
von: Xiao, Yunze, et al.
Veröffentlicht: (2024)
Cell cycle re‐entry primes neuronal senescence in brain aging and dementia
von: Hei‐Man Chow
Veröffentlicht: (2024)
von: Hei‐Man Chow
Veröffentlicht: (2024)
Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models
von: Nakka, Kalyan, et al.
Veröffentlicht: (2024)
von: Nakka, Kalyan, et al.
Veröffentlicht: (2024)
Towards Robust Instruction Tuning on Multimodal Large Language Models
von: Han, Wei, et al.
Veröffentlicht: (2024)
von: Han, Wei, et al.
Veröffentlicht: (2024)
Table Comprehension in Building Codes using Vision Language Models and Domain-Specific Fine-Tuning
von: Aqib, Mohammad, et al.
Veröffentlicht: (2025)
von: Aqib, Mohammad, et al.
Veröffentlicht: (2025)
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
von: Singha, Mainak, et al.
Veröffentlicht: (2025)
Generalized Robust Fundus Photography-based Vision Loss Estimation for High Myopia
von: Yan, Zipei, et al.
Veröffentlicht: (2024)
von: Yan, Zipei, et al.
Veröffentlicht: (2024)
Concealed Interference: Stealth Device–Device Interaction With a Leadless Pacemaker System
von: James E. Ip
Veröffentlicht: (2025)
von: James E. Ip
Veröffentlicht: (2025)
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
von: Yu, Haofei, et al.
Veröffentlicht: (2023)
von: Yu, Haofei, et al.
Veröffentlicht: (2023)
MultiMed: Massively Multimodal and Multitask Medical Understanding
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
von: Hu, Yujia, et al.
Veröffentlicht: (2026)
Inference in Coarsened Time Series via Generalized Method of Moments
von: Man Fai Ip, et al.
Veröffentlicht: (2024)
von: Man Fai Ip, et al.
Veröffentlicht: (2024)
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
von: Bergerault, Antoine, et al.
Veröffentlicht: (2026)
von: Bergerault, Antoine, et al.
Veröffentlicht: (2026)
A cholesterol‐coupled N ‐acetyl‐aspartyl‐glutamate metabolic network facilitates the neuroprotective impact of progesterone‐estradiol crosstalk in neurons
von: Kim Hei‐Man Chow
Veröffentlicht: (2025)
von: Kim Hei‐Man Chow
Veröffentlicht: (2025)
Profilaxis Pre-Exposición en América Latina (Argentina, Brasil y México)
von: Adriel Maroni
Veröffentlicht: (2022)
von: Adriel Maroni
Veröffentlicht: (2022)
What Shapes Participant Data Quality? A Scoping Review and Case Study of Crowdsourced Webcam Eye Tracking in AI Interviews
von: Lau, Ka Hei Carrie, et al.
Veröffentlicht: (2026)
von: Lau, Ka Hei Carrie, et al.
Veröffentlicht: (2026)
Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression
von: Jennings, Roy H., et al.
Veröffentlicht: (2025)
von: Jennings, Roy H., et al.
Veröffentlicht: (2025)
On the Adversarial Robustness of 3D Large Vision-Language Models
von: Liu, Chao, et al.
Veröffentlicht: (2026)
von: Liu, Chao, et al.
Veröffentlicht: (2026)
Resolving space-time structures of quantum impurities with a numerically exact few-body algorithm
von: Núñez-Fernández, Yuriel, et al.
Veröffentlicht: (2025)
von: Núñez-Fernández, Yuriel, et al.
Veröffentlicht: (2025)
Tensorized orbitals for computational chemistry
von: Jolly, Nicolas, et al.
Veröffentlicht: (2023)
von: Jolly, Nicolas, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025) -
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
von: Shan, Xiaojun, et al.
Veröffentlicht: (2025) -
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
von: Cui, Xuanming, et al.
Veröffentlicht: (2025) -
An Ethical Literary Criticism of Han Suyin’s Autobiography
von: Kuek, Florence
Veröffentlicht: (2025) -
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
von: Wang, Jiale, et al.
Veröffentlicht: (2026)