How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Seongyun, Kim, Geewook, Kim, Jiyeon, Lee, Hyunji, Chang, Hoyeon, Park, Sue Hyun, Seo, Minjoon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
di: Kim, Geewook, et al.
Pubblicazione: (2024)
di: Kim, Geewook, et al.
Pubblicazione: (2024)
Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
di: Lee, Seongyun, et al.
Pubblicazione: (2023)
di: Lee, Seongyun, et al.
Pubblicazione: (2023)
State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models
di: Kim, Geewook, et al.
Pubblicazione: (2025)
di: Kim, Geewook, et al.
Pubblicazione: (2025)
Aligning to Thousands of Preferences via System Message Generalization
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy
di: Kim, Geewook, et al.
Pubblicazione: (2025)
di: Kim, Geewook, et al.
Pubblicazione: (2025)
Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams
di: Kim, Jiyeon, et al.
Pubblicazione: (2026)
di: Kim, Jiyeon, et al.
Pubblicazione: (2026)
Evaluating Multimodal Generative AI with Korean Educational Standards
di: Park, Sanghee, et al.
Pubblicazione: (2025)
di: Park, Sanghee, et al.
Pubblicazione: (2025)
MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models
di: Paik, Gio, et al.
Pubblicazione: (2025)
di: Paik, Gio, et al.
Pubblicazione: (2025)
Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
di: Lee, Hyunji, et al.
Pubblicazione: (2025)
di: Lee, Hyunji, et al.
Pubblicazione: (2025)
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models
di: Ju, Jeongho, et al.
Pubblicazione: (2024)
di: Ju, Jeongho, et al.
Pubblicazione: (2024)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
Intriguing Properties of Large Language and Vision Models
di: Lee, Young-Jun, et al.
Pubblicazione: (2024)
di: Lee, Young-Jun, et al.
Pubblicazione: (2024)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
di: Lee, Seongyun, et al.
Pubblicazione: (2026)
di: Lee, Seongyun, et al.
Pubblicazione: (2026)
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
di: Kim, Dain, et al.
Pubblicazione: (2026)
di: Kim, Dain, et al.
Pubblicazione: (2026)
TroL: Traversal of Layers for Large Language and Vision Models
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
Conflict Adaptation in Vision-Language Models
di: Hu, Xiaoyang
Pubblicazione: (2025)
di: Hu, Xiaoyang
Pubblicazione: (2025)
How Well Do Large Language Models Truly Ground?
di: Lee, Hyunji, et al.
Pubblicazione: (2023)
di: Lee, Hyunji, et al.
Pubblicazione: (2023)
Bayesian Principles Improve Prompt Learning In Vision-Language Models
di: Kim, Mingyu, et al.
Pubblicazione: (2025)
di: Kim, Mingyu, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
Can Vision-Language Models Solve the Shell Game?
di: Liu, Tiedong, et al.
Pubblicazione: (2026)
di: Liu, Tiedong, et al.
Pubblicazione: (2026)
CANVAS: A Benchmark for Vision-Language Models on Tool-Based User Interface Design
di: Jeong, Daeheon, et al.
Pubblicazione: (2025)
di: Jeong, Daeheon, et al.
Pubblicazione: (2025)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models
di: Kim, Jiyeon, et al.
Pubblicazione: (2026)
di: Kim, Jiyeon, et al.
Pubblicazione: (2026)
Do Vision-Language Models Understand Visual Persuasiveness?
di: Park, Gyuwon
Pubblicazione: (2025)
di: Park, Gyuwon
Pubblicazione: (2025)
Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition
di: Kim, Jiyeon, et al.
Pubblicazione: (2024)
di: Kim, Jiyeon, et al.
Pubblicazione: (2024)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
di: Chang, Hoyeon, et al.
Pubblicazione: (2024)
di: Chang, Hoyeon, et al.
Pubblicazione: (2024)
Toward Interactive Regional Understanding in Vision-Large Language Models
di: Lee, Jungbeom, et al.
Pubblicazione: (2024)
di: Lee, Jungbeom, et al.
Pubblicazione: (2024)
Vision-Language Models Do Not Understand Negation
di: Alhamoud, Kumail, et al.
Pubblicazione: (2025)
di: Alhamoud, Kumail, et al.
Pubblicazione: (2025)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
di: Seo, Hoigi, et al.
Pubblicazione: (2025)
Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
di: Salmè, Marco, et al.
Pubblicazione: (2025)
di: Salmè, Marco, et al.
Pubblicazione: (2025)
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
di: Zhang, Zheyuan, et al.
Pubblicazione: (2024)
di: Zhang, Zheyuan, et al.
Pubblicazione: (2024)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
di: Lee, Youngwan, et al.
Pubblicazione: (2025)
di: Lee, Youngwan, et al.
Pubblicazione: (2025)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
di: Kim, Jeonghwan, et al.
Pubblicazione: (2024)
di: Kim, Jeonghwan, et al.
Pubblicazione: (2024)
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
di: Lee, Jewon, et al.
Pubblicazione: (2025)
di: Lee, Jewon, et al.
Pubblicazione: (2025)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
di: Geigle, Gregor, et al.
Pubblicazione: (2024)
di: Geigle, Gregor, et al.
Pubblicazione: (2024)
Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
di: Cahyawijaya, Samuel, et al.
Pubblicazione: (2026)
di: Cahyawijaya, Samuel, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
di: Lee, Seongyun, et al.
Pubblicazione: (2024) -
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
di: Kim, Geewook, et al.
Pubblicazione: (2024) -
Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
di: Lee, Seongyun, et al.
Pubblicazione: (2023) -
State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models
di: Kim, Geewook, et al.
Pubblicazione: (2025) -
Aligning to Thousands of Preferences via System Message Generalization
di: Lee, Seongyun, et al.
Pubblicazione: (2024)