Self-Prophetic Decoding to Unlock Visual Search in LVLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Zhendong, Dai, Qiyuan, Li, Guanbin, Lin, Liang, Yang, Sibei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation
von: Dai, Qiyuan, et al.
Veröffentlicht: (2024)
von: Dai, Qiyuan, et al.
Veröffentlicht: (2024)
Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM
von: Dai, Qiyuan, et al.
Veröffentlicht: (2025)
von: Dai, Qiyuan, et al.
Veröffentlicht: (2025)
Sim-DETR: Unlock DETR for Temporal Sentence Grounding
von: Tang, Jiajin, et al.
Veröffentlicht: (2025)
von: Tang, Jiajin, et al.
Veröffentlicht: (2025)
Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement
von: Dai, Qiyuan, et al.
Veröffentlicht: (2025)
von: Dai, Qiyuan, et al.
Veröffentlicht: (2025)
VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction
von: He, Zijian, et al.
Veröffentlicht: (2025)
von: He, Zijian, et al.
Veröffentlicht: (2025)
Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
von: Zheng, Ge, et al.
Veröffentlicht: (2025)
von: Zheng, Ge, et al.
Veröffentlicht: (2025)
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
von: Dai, Ziyun, et al.
Veröffentlicht: (2025)
von: Dai, Ziyun, et al.
Veröffentlicht: (2025)
Chart Deep Research in LVLMs via Parallel Relative Policy Optimization
von: Tang, Jiajin, et al.
Veröffentlicht: (2026)
von: Tang, Jiajin, et al.
Veröffentlicht: (2026)
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
von: Li, Qiming, et al.
Veröffentlicht: (2025)
von: Li, Qiming, et al.
Veröffentlicht: (2025)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
von: Zhang, Ruifei, et al.
Veröffentlicht: (2025)
von: Zhang, Ruifei, et al.
Veröffentlicht: (2025)
Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection
von: de Senneville, Adhemar, et al.
Veröffentlicht: (2026)
von: de Senneville, Adhemar, et al.
Veröffentlicht: (2026)
Rethinking Query-based Transformer for Continual Image Segmentation
von: Zhu, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhu, Yuchen, et al.
Veröffentlicht: (2025)
One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
von: Liu, Jinxi, et al.
Veröffentlicht: (2025)
von: Liu, Jinxi, et al.
Veröffentlicht: (2025)
Improving Alignment in LVLMs with Debiased Self-Judgment
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
Credible Teacher for Semi-Supervised Object Detection in Open Scene
von: Zhuang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhuang, Jingyu, et al.
Veröffentlicht: (2024)
The devil is in the object boundary: towards annotation-free instance segmentation using Foundation Models
von: Shi, Cheng, et al.
Veröffentlicht: (2024)
von: Shi, Cheng, et al.
Veröffentlicht: (2024)
Self-Improving Small Object Grounding in LVLMs
von: Yang, Tianze, et al.
Veröffentlicht: (2026)
von: Yang, Tianze, et al.
Veröffentlicht: (2026)
Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
von: Wu, Yuanchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuanchen, et al.
Veröffentlicht: (2025)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
von: Yuan, Fan, et al.
Veröffentlicht: (2024)
von: Yuan, Fan, et al.
Veröffentlicht: (2024)
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
von: Huang, Siyuan, et al.
Veröffentlicht: (2026)
von: Huang, Siyuan, et al.
Veröffentlicht: (2026)
CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
von: Kan, Zhehan, et al.
Veröffentlicht: (2024)
von: Kan, Zhehan, et al.
Veröffentlicht: (2024)
Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention
von: Zhang, Zhendong
Veröffentlicht: (2025)
von: Zhang, Zhendong
Veröffentlicht: (2025)
Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement
von: Qin, Zhenxin, et al.
Veröffentlicht: (2026)
von: Qin, Zhenxin, et al.
Veröffentlicht: (2026)
Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs
von: Lyu, Guangtao, et al.
Veröffentlicht: (2026)
von: Lyu, Guangtao, et al.
Veröffentlicht: (2026)
GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection
von: Li, Jiaming, et al.
Veröffentlicht: (2026)
von: Li, Jiaming, et al.
Veröffentlicht: (2026)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs
von: Chen, Huiyi, et al.
Veröffentlicht: (2025)
von: Chen, Huiyi, et al.
Veröffentlicht: (2025)
Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
von: Xu, Rui, et al.
Veröffentlicht: (2025)
von: Xu, Rui, et al.
Veröffentlicht: (2025)
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuan, et al.
Veröffentlicht: (2025)
CoMemo: LVLMs Need Image Context with Image Memory
von: Liu, Shi, et al.
Veröffentlicht: (2025)
von: Liu, Shi, et al.
Veröffentlicht: (2025)
Penalizing Boundary Activation for Object Completeness in Diffusion Models
von: Xu, Haoyang, et al.
Veröffentlicht: (2025)
von: Xu, Haoyang, et al.
Veröffentlicht: (2025)
PDC-Net: Pattern Divide-and-Conquer Network for Pelvic Radiation Injury Segmentation
von: Xiong, Xinyu, et al.
Veröffentlicht: (2025)
von: Xiong, Xinyu, et al.
Veröffentlicht: (2025)
See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
von: He, Zijian, et al.
Veröffentlicht: (2024)
von: He, Zijian, et al.
Veröffentlicht: (2024)
ODMixer: Fine-grained Spatial-temporal MLP for Metro Origin-Destination Prediction
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
von: Song, Xinshuai, et al.
Veröffentlicht: (2024)
von: Song, Xinshuai, et al.
Veröffentlicht: (2024)
Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
von: Hu, Qingguo, et al.
Veröffentlicht: (2025)
von: Hu, Qingguo, et al.
Veröffentlicht: (2025)
CHASD: Language Increment-Calibrated Contrastive Decoding against Hallucination in LVLMs
von: Huang, Xiaoyi, et al.
Veröffentlicht: (2026)
von: Huang, Xiaoyi, et al.
Veröffentlicht: (2026)
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation
von: Dai, Qiyuan, et al.
Veröffentlicht: (2024) -
Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM
von: Dai, Qiyuan, et al.
Veröffentlicht: (2025) -
Sim-DETR: Unlock DETR for Temporal Sentence Grounding
von: Tang, Jiajin, et al.
Veröffentlicht: (2025) -
Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement
von: Dai, Qiyuan, et al.
Veröffentlicht: (2025) -
VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction
von: He, Zijian, et al.
Veröffentlicht: (2025)