Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Ryan, Yuriel, Ip, Hei Man, Kuek, Adriel, Liang, Paul Pu, Lee, Roy Ka-Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
di: Ryan, Yuriel, et al.
Pubblicazione: (2025)
di: Ryan, Yuriel, et al.
Pubblicazione: (2025)
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
di: Shan, Xiaojun, et al.
Pubblicazione: (2025)
di: Shan, Xiaojun, et al.
Pubblicazione: (2025)
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
di: Cui, Xuanming, et al.
Pubblicazione: (2025)
di: Cui, Xuanming, et al.
Pubblicazione: (2025)
An Ethical Literary Criticism of Han Suyin’s Autobiography
di: Kuek, Florence
Pubblicazione: (2025)
di: Kuek, Florence
Pubblicazione: (2025)
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
di: Wang, Jiale, et al.
Pubblicazione: (2026)
di: Wang, Jiale, et al.
Pubblicazione: (2026)
Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs
di: Zun, Lee Qi, et al.
Pubblicazione: (2025)
di: Zun, Lee Qi, et al.
Pubblicazione: (2025)
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
di: Roy, Shuvendu, et al.
Pubblicazione: (2025)
di: Roy, Shuvendu, et al.
Pubblicazione: (2025)
On the Adversarial Robustness of Instruction-Tuned Large Language Models for Code
di: Hossen, Md Imran, et al.
Pubblicazione: (2024)
di: Hossen, Md Imran, et al.
Pubblicazione: (2024)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
di: Tan, Rui Yang, et al.
Pubblicazione: (2026)
di: Tan, Rui Yang, et al.
Pubblicazione: (2026)
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
di: Tan, Bryan Chen Zhengyu, et al.
Pubblicazione: (2025)
di: Tan, Bryan Chen Zhengyu, et al.
Pubblicazione: (2025)
Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving
di: Chen, Steven-Shine, et al.
Pubblicazione: (2025)
di: Chen, Steven-Shine, et al.
Pubblicazione: (2025)
Imperfect World Models are Exploitable
di: Bhamidipaty, Logan Mondal, et al.
Pubblicazione: (2026)
di: Bhamidipaty, Logan Mondal, et al.
Pubblicazione: (2026)
Users’ quality expectations and their correspondence with the realistic features of translation applications
di: Jing Rou Kuek
Pubblicazione: (2024)
di: Jing Rou Kuek
Pubblicazione: (2024)
Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions
di: Hee, Ming Shan, et al.
Pubblicazione: (2025)
di: Hee, Ming Shan, et al.
Pubblicazione: (2025)
Interpreting Bias in Large Language Models: A Feature-Based Approach
di: Prakash, Nirmalendu, et al.
Pubblicazione: (2024)
di: Prakash, Nirmalendu, et al.
Pubblicazione: (2024)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
A Vision for Multisensory Intelligence: Sensing, Science, and Synergy
di: Liang, Paul Pu
Pubblicazione: (2026)
di: Liang, Paul Pu
Pubblicazione: (2026)
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
di: Wang, Han, et al.
Pubblicazione: (2026)
di: Wang, Han, et al.
Pubblicazione: (2026)
AromaGen: Interactive Generation of Rich Olfactory Experiences with Multimodal Language Models
di: Wen, Yunge, et al.
Pubblicazione: (2026)
di: Wen, Yunge, et al.
Pubblicazione: (2026)
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
di: Xiao, Yunze, et al.
Pubblicazione: (2024)
di: Xiao, Yunze, et al.
Pubblicazione: (2024)
Cell cycle re‐entry primes neuronal senescence in brain aging and dementia
di: Hei‐Man Chow
Pubblicazione: (2024)
di: Hei‐Man Chow
Pubblicazione: (2024)
Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models
di: Nakka, Kalyan, et al.
Pubblicazione: (2024)
di: Nakka, Kalyan, et al.
Pubblicazione: (2024)
Towards Robust Instruction Tuning on Multimodal Large Language Models
di: Han, Wei, et al.
Pubblicazione: (2024)
di: Han, Wei, et al.
Pubblicazione: (2024)
Table Comprehension in Building Codes using Vision Language Models and Domain-Specific Fine-Tuning
di: Aqib, Mohammad, et al.
Pubblicazione: (2025)
di: Aqib, Mohammad, et al.
Pubblicazione: (2025)
FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
di: Singha, Mainak, et al.
Pubblicazione: (2025)
di: Singha, Mainak, et al.
Pubblicazione: (2025)
Generalized Robust Fundus Photography-based Vision Loss Estimation for High Myopia
di: Yan, Zipei, et al.
Pubblicazione: (2024)
di: Yan, Zipei, et al.
Pubblicazione: (2024)
Concealed Interference: Stealth Device–Device Interaction With a Leadless Pacemaker System
di: James E. Ip
Pubblicazione: (2025)
di: James E. Ip
Pubblicazione: (2025)
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
di: Yu, Haofei, et al.
Pubblicazione: (2023)
di: Yu, Haofei, et al.
Pubblicazione: (2023)
MultiMed: Massively Multimodal and Multitask Medical Understanding
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
di: Hu, Yujia, et al.
Pubblicazione: (2026)
di: Hu, Yujia, et al.
Pubblicazione: (2026)
Inference in Coarsened Time Series via Generalized Method of Moments
di: Man Fai Ip, et al.
Pubblicazione: (2024)
di: Man Fai Ip, et al.
Pubblicazione: (2024)
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
di: Bergerault, Antoine, et al.
Pubblicazione: (2026)
di: Bergerault, Antoine, et al.
Pubblicazione: (2026)
A cholesterol‐coupled N ‐acetyl‐aspartyl‐glutamate metabolic network facilitates the neuroprotective impact of progesterone‐estradiol crosstalk in neurons
di: Kim Hei‐Man Chow
Pubblicazione: (2025)
di: Kim Hei‐Man Chow
Pubblicazione: (2025)
Profilaxis Pre-Exposición en América Latina (Argentina, Brasil y México)
di: Adriel Maroni
Pubblicazione: (2022)
di: Adriel Maroni
Pubblicazione: (2022)
What Shapes Participant Data Quality? A Scoping Review and Case Study of Crowdsourced Webcam Eye Tracking in AI Interviews
di: Lau, Ka Hei Carrie, et al.
Pubblicazione: (2026)
di: Lau, Ka Hei Carrie, et al.
Pubblicazione: (2026)
Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression
di: Jennings, Roy H., et al.
Pubblicazione: (2025)
di: Jennings, Roy H., et al.
Pubblicazione: (2025)
On the Adversarial Robustness of 3D Large Vision-Language Models
di: Liu, Chao, et al.
Pubblicazione: (2026)
di: Liu, Chao, et al.
Pubblicazione: (2026)
Resolving space-time structures of quantum impurities with a numerically exact few-body algorithm
di: Núñez-Fernández, Yuriel, et al.
Pubblicazione: (2025)
di: Núñez-Fernández, Yuriel, et al.
Pubblicazione: (2025)
Tensorized orbitals for computational chemistry
di: Jolly, Nicolas, et al.
Pubblicazione: (2023)
di: Jolly, Nicolas, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
di: Ryan, Yuriel, et al.
Pubblicazione: (2025) -
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
di: Shan, Xiaojun, et al.
Pubblicazione: (2025) -
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
di: Cui, Xuanming, et al.
Pubblicazione: (2025) -
An Ethical Literary Criticism of Han Suyin’s Autobiography
di: Kuek, Florence
Pubblicazione: (2025) -
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
di: Wang, Jiale, et al.
Pubblicazione: (2026)