Benchmarking Robustness of Multimodal Image-Text Models under Distribution Shift
Fuente:
arXiv
Salvato in:
| Autori principali: | Qiu, Jielin, Zhu, Yi, Shi, Xingjian, Wenzel, Florian, Tang, Zhiqiang, Zhao, Ding, Li, Bo, Li, Mu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Durability: Benchmark Insights into Multimodal Watermarking
di: Qiu, Jielin, et al.
Pubblicazione: (2024)
di: Qiu, Jielin, et al.
Pubblicazione: (2024)
COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
di: Li, Jiansheng, et al.
Pubblicazione: (2025)
di: Li, Jiansheng, et al.
Pubblicazione: (2025)
OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution Shift
di: Li, Lin, et al.
Pubblicazione: (2023)
di: Li, Lin, et al.
Pubblicazione: (2023)
On the Robustness of Human-Object Interaction Detection against Distribution Shift
di: Xie, Chi, et al.
Pubblicazione: (2025)
di: Xie, Chi, et al.
Pubblicazione: (2025)
FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion
di: Ruan, Jiacheng, et al.
Pubblicazione: (2024)
di: Ruan, Jiacheng, et al.
Pubblicazione: (2024)
Robust Distribution Alignment for Industrial Anomaly Detection under Distribution Shift
di: Liao, Jingyi, et al.
Pubblicazione: (2025)
di: Liao, Jingyi, et al.
Pubblicazione: (2025)
Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts
di: Zhong, Guowei, et al.
Pubblicazione: (2025)
di: Zhong, Guowei, et al.
Pubblicazione: (2025)
MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
di: Shan, Bin, et al.
Pubblicazione: (2024)
di: Shan, Bin, et al.
Pubblicazione: (2024)
Bootstrap Segmentation Foundation Model under Distribution Shift via Object-Centric Learning
di: Tang, Luyao, et al.
Pubblicazione: (2024)
di: Tang, Luyao, et al.
Pubblicazione: (2024)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
di: Li, Yuheng, et al.
Pubblicazione: (2024)
di: Li, Yuheng, et al.
Pubblicazione: (2024)
CRISP: Rank-Guided Iterative Squeezing for Robust Medical Image Segmentation under Domain Shift
di: Fang, Yizhou, et al.
Pubblicazione: (2026)
di: Fang, Yizhou, et al.
Pubblicazione: (2026)
Robust Visual Representation Learning with Multi-modal Prior Knowledge for Image Classification Under Distribution Shift
di: Zhou, Hongkuan, et al.
Pubblicazione: (2024)
di: Zhou, Hongkuan, et al.
Pubblicazione: (2024)
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
di: Ma, Lichen, et al.
Pubblicazione: (2026)
di: Ma, Lichen, et al.
Pubblicazione: (2026)
Open-Vocabulary Object Detectors: Robustness Challenges under Distribution Shifts
di: Chhipa, Prakash Chandra, et al.
Pubblicazione: (2024)
di: Chhipa, Prakash Chandra, et al.
Pubblicazione: (2024)
Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
di: Wang, Zhikai, et al.
Pubblicazione: (2025)
di: Wang, Zhikai, et al.
Pubblicazione: (2025)
Benchmarking Object Detectors under Real-World Distribution Shifts in Satellite Imagery
di: Al-Emadi, Sara, et al.
Pubblicazione: (2025)
di: Al-Emadi, Sara, et al.
Pubblicazione: (2025)
DomainVerse: A Benchmark Towards Real-World Distribution Shifts For Tuning-Free Adaptive Domain Generalization
di: Hou, Feng, et al.
Pubblicazione: (2024)
di: Hou, Feng, et al.
Pubblicazione: (2024)
CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance Shifts
di: Dünkel, Olaf, et al.
Pubblicazione: (2025)
di: Dünkel, Olaf, et al.
Pubblicazione: (2025)
Face4FairShifts: A Large Image Benchmark for Fairness and Robust Learning across Visual Domains
di: Lin, Yumeng, et al.
Pubblicazione: (2025)
di: Lin, Yumeng, et al.
Pubblicazione: (2025)
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
di: Wang, Bingli, et al.
Pubblicazione: (2026)
di: Wang, Bingli, et al.
Pubblicazione: (2026)
Robust Phase-Shifting Profilometry for Arbitrary Motion
di: Zhang, Geyou, et al.
Pubblicazione: (2025)
di: Zhang, Geyou, et al.
Pubblicazione: (2025)
MMRareBench: A Rare-Disease Multimodal and Multi-Image Medical Benchmark
di: Ning, Junzhi, et al.
Pubblicazione: (2026)
di: Ning, Junzhi, et al.
Pubblicazione: (2026)
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
di: Liang, Yiqing, et al.
Pubblicazione: (2025)
di: Liang, Yiqing, et al.
Pubblicazione: (2025)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
di: Yang, Shuyu, et al.
Pubblicazione: (2024)
di: Yang, Shuyu, et al.
Pubblicazione: (2024)
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
di: Zhang, Junzhe, et al.
Pubblicazione: (2025)
di: Zhang, Junzhe, et al.
Pubblicazione: (2025)
TextCAM: Explaining Class Activation Map with Text
di: Zhao, Qiming, et al.
Pubblicazione: (2025)
di: Zhao, Qiming, et al.
Pubblicazione: (2025)
Generalize or Detect? Towards Robust Semantic Segmentation Under Multiple Distribution Shifts
di: Gao, Zhitong, et al.
Pubblicazione: (2024)
di: Gao, Zhitong, et al.
Pubblicazione: (2024)
DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models
di: Wang, Jiarui, et al.
Pubblicazione: (2025)
di: Wang, Jiarui, et al.
Pubblicazione: (2025)
What Color Is It? A Text-Interference Multimodal Hallucination Benchmark
di: Zhao, Jinkun, et al.
Pubblicazione: (2025)
di: Zhao, Jinkun, et al.
Pubblicazione: (2025)
FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
di: Yi, Junchao, et al.
Pubblicazione: (2026)
di: Yi, Junchao, et al.
Pubblicazione: (2026)
Affective Image Editing: Shaping Emotional Factors via Text Descriptions
di: Zhang, Peixuan, et al.
Pubblicazione: (2025)
di: Zhang, Peixuan, et al.
Pubblicazione: (2025)
ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection
di: Ma, Ruize, et al.
Pubblicazione: (2025)
di: Ma, Ruize, et al.
Pubblicazione: (2025)
SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
di: Li, Bohao, et al.
Pubblicazione: (2024)
di: Li, Bohao, et al.
Pubblicazione: (2024)
LESSViT: Robust Hyperspectral Representation Learning under Spectral Configuration Shift
di: Si, Haozhe, et al.
Pubblicazione: (2026)
di: Si, Haozhe, et al.
Pubblicazione: (2026)
15M Multimodal Facial Image-Text Dataset
di: Dai, Dawei, et al.
Pubblicazione: (2024)
di: Dai, Dawei, et al.
Pubblicazione: (2024)
Class Similarity-Based Multimodal Classification under Heterogeneous Category Sets
di: Zhu, Yangrui, et al.
Pubblicazione: (2025)
di: Zhu, Yangrui, et al.
Pubblicazione: (2025)
T2MBench: A Benchmark for Out-of-Distribution Text-to-Motion Generation
di: Yang, Bin, et al.
Pubblicazione: (2026)
di: Yang, Bin, et al.
Pubblicazione: (2026)
Uncertainty-Aware Subset Selection for Robust Visual Explainability under Distribution Shifts
di: Gupta, Madhav, et al.
Pubblicazione: (2025)
di: Gupta, Madhav, et al.
Pubblicazione: (2025)
Generating Multimodal Images with GAN: Integrating Text, Image, and Style
di: Tan, Chaoyi, et al.
Pubblicazione: (2025)
di: Tan, Chaoyi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating Durability: Benchmark Insights into Multimodal Watermarking
di: Qiu, Jielin, et al.
Pubblicazione: (2024) -
COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
di: Li, Jiansheng, et al.
Pubblicazione: (2025) -
OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution Shift
di: Li, Lin, et al.
Pubblicazione: (2023) -
On the Robustness of Human-Object Interaction Detection against Distribution Shift
di: Xie, Chi, et al.
Pubblicazione: (2025) -
FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion
di: Ruan, Jiacheng, et al.
Pubblicazione: (2024)