RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Woo, Sangmin, Jang, Jaehyuk, Kim, Donguk, Choi, Yubin, Kim, Changick |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
Denoising Task Routing for Diffusion Models
by: Park, Byeongjun, et al.
Published: (2023)
by: Park, Byeongjun, et al.
Published: (2023)
Efficient Test-Time Optimization for Depth Completion via Low-Rank Decoder Adaptation
by: Seo, Minseok, et al.
Published: (2026)
by: Seo, Minseok, et al.
Published: (2026)
Diffusion Model Patching via Mixture-of-Prompts
by: Ham, Seokil, et al.
Published: (2024)
by: Ham, Seokil, et al.
Published: (2024)
Spatio-Temporal Proximity-Aware Dual-Path Model for Panoramic Activity Recognition
by: Lee, Sumin, et al.
Published: (2024)
by: Lee, Sumin, et al.
Published: (2024)
Bridging Implicit and Explicit Geometric Transformation for Single-Image View Synthesis
by: Park, Byeongjun, et al.
Published: (2022)
by: Park, Byeongjun, et al.
Published: (2022)
Switch Diffusion Transformer: Synergizing Denoising Tasks with Sparse Mixture-of-Experts
by: Park, Byeongjun, et al.
Published: (2024)
by: Park, Byeongjun, et al.
Published: (2024)
Flow-Assisted Motion Learning Network for Weakly-Supervised Group Activity Recognition
by: Nugroho, Muhammad Adi, et al.
Published: (2024)
by: Nugroho, Muhammad Adi, et al.
Published: (2024)
Doubly-Universal Adversarial Perturbations: Deceiving Vision-Language Models Across Both Images and Text with a Single Perturbation
by: Kim, Hee-Seon, et al.
Published: (2024)
by: Kim, Hee-Seon, et al.
Published: (2024)
Avoid Wasted Annotation Costs in Open-set Active Learning with Pre-trained Vision-Language Model
by: Heo, Jaehyuk, et al.
Published: (2024)
by: Heo, Jaehyuk, et al.
Published: (2024)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
by: Zhang, Xinsong, et al.
Published: (2025)
by: Zhang, Xinsong, et al.
Published: (2025)
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
by: Woo, Sangmin, et al.
Published: (2021)
by: Woo, Sangmin, et al.
Published: (2021)
APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing
by: Han, Sangmin, et al.
Published: (2025)
by: Han, Sangmin, et al.
Published: (2025)
Towards Efficient Vision State Space Models via Token Merging
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation
by: Cho, Yubin, et al.
Published: (2024)
by: Cho, Yubin, et al.
Published: (2024)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
by: Kim, Mingyeong, et al.
Published: (2026)
by: Kim, Mingyeong, et al.
Published: (2026)
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026)
by: Park, Wooje, et al.
Published: (2026)
Detecting AI-Generated Images via Contextual Anomaly Estimation in Masked AutoEncoders
by: Jang, Minsuk, et al.
Published: (2025)
by: Jang, Minsuk, et al.
Published: (2025)
Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
by: Kim, Hee-Seon, et al.
Published: (2025)
by: Kim, Hee-Seon, et al.
Published: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
by: Kim, Ju-Young, et al.
Published: (2025)
by: Kim, Ju-Young, et al.
Published: (2025)
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
by: Kim, Bum Jun, et al.
Published: (2024)
by: Kim, Bum Jun, et al.
Published: (2024)
Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
by: Yang, Xu, et al.
Published: (2023)
by: Yang, Xu, et al.
Published: (2023)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2025)
by: Woo, Sangmin, et al.
Published: (2025)
Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
by: Kim, Hayeon, et al.
Published: (2026)
by: Kim, Hayeon, et al.
Published: (2026)
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
by: Kim, Eunki, et al.
Published: (2025)
by: Kim, Eunki, et al.
Published: (2025)
Stop learning it all to mitigate visual hallucination, Focus on the hallucination target
by: Yoon, Dokyoon, et al.
Published: (2025)
by: Yoon, Dokyoon, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
by: An, Na Min, et al.
Published: (2025)
by: An, Na Min, et al.
Published: (2025)
SAFIRE: Segment Any Forged Image Region
by: Kwon, Myung-Joon, et al.
Published: (2024)
by: Kwon, Myung-Joon, et al.
Published: (2024)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
by: Kim, Sohee, et al.
Published: (2025)
by: Kim, Sohee, et al.
Published: (2025)
BEEP3D: Box-Supervised End-to-End Pseudo-Mask Generation for 3D Instance Segmentation
by: Yoo, Youngju, et al.
Published: (2025)
by: Yoo, Youngju, et al.
Published: (2025)
Station2Radar: query conditioned gaussian splatting for precipitation field
by: Kim, Doyi, et al.
Published: (2026)
by: Kim, Doyi, et al.
Published: (2026)
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
by: Lee, Dong-Jae, et al.
Published: (2026)
by: Lee, Dong-Jae, et al.
Published: (2026)
Intersectional Fairness in Vision-Language Models for Medical Image Disease Classification
by: Zhang, Yupeng, et al.
Published: (2025)
by: Zhang, Yupeng, et al.
Published: (2025)
Soft Segmented Randomization: Enhancing Domain Generalization in SAR-ATR for Synthetic-to-Measured
by: Kim, Minjun, et al.
Published: (2024)
by: Kim, Minjun, et al.
Published: (2024)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
SELFI: Selective Fusion of Identity for Generalizable Deepfake Detection
by: Kim, Younghun, et al.
Published: (2025)
by: Kim, Younghun, et al.
Published: (2025)
Similar Items
-
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024) -
Denoising Task Routing for Diffusion Models
by: Park, Byeongjun, et al.
Published: (2023) -
Efficient Test-Time Optimization for Depth Completion via Low-Rank Decoder Adaptation
by: Seo, Minseok, et al.
Published: (2026) -
Diffusion Model Patching via Mixture-of-Prompts
by: Ham, Seokil, et al.
Published: (2024) -
Spatio-Temporal Proximity-Aware Dual-Path Model for Panoramic Activity Recognition
by: Lee, Sumin, et al.
Published: (2024)