Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Shuo, Sun, Jiajun, Zheng, Guodong, Fan, Xiaoran, Shen, Yujiong, Lu, Yi, Xi, Zhiheng, Yang, Yuming, Tan, Wenming, Ji, Tao, Gui, Tao, Zhang, Qi, Huang, Xuanjing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
von: Li, Shuo, et al.
Veröffentlicht: (2024)
von: Li, Shuo, et al.
Veröffentlicht: (2024)
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
The Role of Entropy in Visual Grounding: Analysis and Optimization
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective
von: Zhang, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2025)
ChartE$^{3}$: A Comprehensive Benchmark for End-to-End Chart Editing
von: Li, Shuo, et al.
Veröffentlicht: (2026)
von: Li, Shuo, et al.
Veröffentlicht: (2026)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
von: Fan, Xiaoran, et al.
Veröffentlicht: (2026)
von: Fan, Xiaoran, et al.
Veröffentlicht: (2026)
RoCoIns: Enhancing Robustness of Large Language Models through Code-Style Instructions
von: Zhang, Yuansen, et al.
Veröffentlicht: (2024)
von: Zhang, Yuansen, et al.
Veröffentlicht: (2024)
Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
von: Ding, Yiwen, et al.
Veröffentlicht: (2024)
von: Ding, Yiwen, et al.
Veröffentlicht: (2024)
SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
von: Tu, Chongjun, et al.
Veröffentlicht: (2025)
von: Tu, Chongjun, et al.
Veröffentlicht: (2025)
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
von: Yin, Hao, et al.
Veröffentlicht: (2025)
von: Yin, Hao, et al.
Veröffentlicht: (2025)
Self-Polish: Enhance Reasoning in Large Language Models via Problem Refinement
von: Xi, Zhiheng, et al.
Veröffentlicht: (2023)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2023)
Self-Demos: Eliciting Out-of-Demonstration Generalizability in Large Language Models
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment
von: Sarkar, Pritam, et al.
Veröffentlicht: (2024)
von: Sarkar, Pritam, et al.
Veröffentlicht: (2024)
TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
von: Liu, Yufang, et al.
Veröffentlicht: (2024)
von: Liu, Yufang, et al.
Veröffentlicht: (2024)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
von: Zhang, Xi, et al.
Veröffentlicht: (2025)
von: Zhang, Xi, et al.
Veröffentlicht: (2025)
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
von: Zhou, Enyu, et al.
Veröffentlicht: (2024)
von: Zhou, Enyu, et al.
Veröffentlicht: (2024)
StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
von: Guo, Xin, et al.
Veröffentlicht: (2025)
von: Guo, Xin, et al.
Veröffentlicht: (2025)
60 Data Points are Sufficient to Fine-Tune LLMs for Question-Answering
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
von: Zhang, Kejia, et al.
Veröffentlicht: (2025)
von: Zhang, Kejia, et al.
Veröffentlicht: (2025)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
von: Lu, Yi, et al.
Veröffentlicht: (2024)
von: Lu, Yi, et al.
Veröffentlicht: (2024)
CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges
von: Li, Hui, et al.
Veröffentlicht: (2025)
von: Li, Hui, et al.
Veröffentlicht: (2025)
ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
LoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
von: Dou, Shihan, et al.
Veröffentlicht: (2023)
von: Dou, Shihan, et al.
Veröffentlicht: (2023)
Unveiling Linguistic Regions in Large Language Models
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
Length Generalization of Causal Transformers without Position Encoding
von: Wang, Jie, et al.
Veröffentlicht: (2024)
von: Wang, Jie, et al.
Veröffentlicht: (2024)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
von: He, Wei, et al.
Veröffentlicht: (2024) -
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
von: Zheng, Rui, et al.
Veröffentlicht: (2024) -
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
von: Li, Shuo, et al.
Veröffentlicht: (2024) -
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
von: Jiang, Changhao, et al.
Veröffentlicht: (2025) -
The Role of Entropy in Visual Grounding: Analysis and Optimization
von: Li, Shuo, et al.
Veröffentlicht: (2025)