TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Kejia, Tao, Keda, Luo, Zhiming, Liu, Chang, Tang, Jiasheng, Wang, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
by: Zhang, Kejia, et al.
Published: (2025)
by: Zhang, Kejia, et al.
Published: (2025)
CHASD: Language Increment-Calibrated Contrastive Decoding against Hallucination in LVLMs
by: Huang, Xiaoyi, et al.
Published: (2026)
by: Huang, Xiaoyi, et al.
Published: (2026)
A Survey of Token Compression for Efficient Multimodal Large Language Models
by: Shao, Kele, et al.
Published: (2025)
by: Shao, Kele, et al.
Published: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
by: Chen, Xueyi, et al.
Published: (2025)
by: Chen, Xueyi, et al.
Published: (2025)
MANI-Pure: Magnitude-Adaptive Noise Injection for Adversarial Purification
by: Huang, Xiaoyi, et al.
Published: (2025)
by: Huang, Xiaoyi, et al.
Published: (2025)
HoliTom: Holistic Token Merging for Fast Video Large Language Models
by: Shao, Kele, et al.
Published: (2025)
by: Shao, Kele, et al.
Published: (2025)
Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
by: Li, Yuanshuai, et al.
Published: (2025)
by: Li, Yuanshuai, et al.
Published: (2025)
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
by: Tao, Keda, et al.
Published: (2024)
by: Tao, Keda, et al.
Published: (2024)
OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
by: Tao, Keda, et al.
Published: (2025)
by: Tao, Keda, et al.
Published: (2025)
Min-Max-Jump distance and its applications
by: Liu, Gangli
Published: (2023)
by: Liu, Gangli
Published: (2023)
One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs Hallucination
by: Fa, Zhan, et al.
Published: (2026)
by: Fa, Zhan, et al.
Published: (2026)
Towards Adversarial Robustness via Debiased High-Confidence Logit Alignment
by: Zhang, Kejia, et al.
Published: (2024)
by: Zhang, Kejia, et al.
Published: (2024)
Is Oracle Pruning the True Oracle?
by: Feng, Sicheng, et al.
Published: (2024)
by: Feng, Sicheng, et al.
Published: (2024)
Generative Dataset Distillation using Min-Max Diffusion Model
by: Fan, Junqiao, et al.
Published: (2025)
by: Fan, Junqiao, et al.
Published: (2025)
SparseDiT: Token Sparsification for Efficient Diffusion Transformer
by: Chang, Shuning, et al.
Published: (2024)
by: Chang, Shuning, et al.
Published: (2024)
TARS: Traffic-Aware Radar Scene Flow Estimation
by: Wu, Jialong, et al.
Published: (2025)
by: Wu, Jialong, et al.
Published: (2025)
Mitigating Low-Frequency Bias: Feature Recalibration and Frequency Attention Regularization for Adversarial Robustness
by: Zhang, Kejia, et al.
Published: (2024)
by: Zhang, Kejia, et al.
Published: (2024)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024)
by: Zhang, Mengxi, et al.
Published: (2024)
Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs
by: Si, Guangzong, et al.
Published: (2025)
by: Si, Guangzong, et al.
Published: (2025)
Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs
by: Lin, Yuhui, et al.
Published: (2026)
by: Lin, Yuhui, et al.
Published: (2026)
Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
by: Tao, Keda, et al.
Published: (2025)
by: Tao, Keda, et al.
Published: (2025)
Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
by: Tu, Chongjun, et al.
Published: (2025)
by: Tu, Chongjun, et al.
Published: (2025)
Explore the Hallucination on Low-level Perception for MLLMs
by: Sun, Yinan, et al.
Published: (2024)
by: Sun, Yinan, et al.
Published: (2024)
Together, Then Apart: Revisiting Multimodal Survival Analysis via a Min-Max Perspective
by: Liu, Wenjing, et al.
Published: (2025)
by: Liu, Wenjing, et al.
Published: (2025)
Active Perception Agent for Omnimodal Audio-Video Understanding
by: Tao, Keda, et al.
Published: (2025)
by: Tao, Keda, et al.
Published: (2025)
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
by: Gu, Jihao, et al.
Published: (2024)
by: Gu, Jihao, et al.
Published: (2024)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
Harmonizing Feature Maps: A Graph Convolutional Approach for Enhancing Adversarial Robustness
by: Zhang, Kejia, et al.
Published: (2024)
by: Zhang, Kejia, et al.
Published: (2024)
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective
by: Zhao, Qiyan, et al.
Published: (2026)
by: Zhao, Qiyan, et al.
Published: (2026)
MMARD: Improving the Min-Max Optimization Process in Adversarial Robustness Distillation
by: Wang, Yuzheng, et al.
Published: (2025)
by: Wang, Yuzheng, et al.
Published: (2025)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
by: Jin, Xinqi, et al.
Published: (2025)
by: Jin, Xinqi, et al.
Published: (2025)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
by: Yin, Zhihan, et al.
Published: (2026)
by: Yin, Zhihan, et al.
Published: (2026)
Interpreting and Mitigating Hallucination in MLLMs through Multi-agent Debate
by: Lin, Zheng, et al.
Published: (2024)
by: Lin, Zheng, et al.
Published: (2024)
Revisiting Min-Max Optimization Problem in Adversarial Training
by: Ahmadi, Sina Hajer, et al.
Published: (2024)
by: Ahmadi, Sina Hajer, et al.
Published: (2024)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
InpaintDPO: Mitigating Spatial Relationship Hallucinations in Foreground-conditioned Inpainting via Diverse Preference Optimization
by: Li, Qirui, et al.
Published: (2025)
by: Li, Qirui, et al.
Published: (2025)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025)
by: Tang, Feilong, et al.
Published: (2025)
Similar Items
-
Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
by: Zhang, Kejia, et al.
Published: (2025) -
CHASD: Language Increment-Calibrated Contrastive Decoding against Hallucination in LVLMs
by: Huang, Xiaoyi, et al.
Published: (2026) -
A Survey of Token Compression for Efficient Multimodal Large Language Models
by: Shao, Kele, et al.
Published: (2025) -
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
by: Chen, Xueyi, et al.
Published: (2025) -
MANI-Pure: Magnitude-Adaptive Noise Injection for Adversarial Purification
by: Huang, Xiaoyi, et al.
Published: (2025)