Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Zekai, Li, Qiming, Feng, Xiaocheng, Chen, Ruihan, Li, Ziming, Ren, Haoyu, Chen, Kun, Tu, Dandan, Qin, Bing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
von: Chen, Ruihan, et al.
Veröffentlicht: (2025)
von: Chen, Ruihan, et al.
Veröffentlicht: (2025)
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
von: Li, Qiming, et al.
Veröffentlicht: (2025)
von: Li, Qiming, et al.
Veröffentlicht: (2025)
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
von: Li, Qiming, et al.
Veröffentlicht: (2025)
von: Li, Qiming, et al.
Veröffentlicht: (2025)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
von: Li, Qiming, et al.
Veröffentlicht: (2026)
von: Li, Qiming, et al.
Veröffentlicht: (2026)
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
von: Li, Qiming, et al.
Veröffentlicht: (2025)
von: Li, Qiming, et al.
Veröffentlicht: (2025)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
von: Ye, Zekai, et al.
Veröffentlicht: (2025)
von: Ye, Zekai, et al.
Veröffentlicht: (2025)
Culture-Aware Machine Translation in Large Language Models: Benchmarking and Investigation
von: Yuan, Zekun, et al.
Veröffentlicht: (2026)
von: Yuan, Zekun, et al.
Veröffentlicht: (2026)
Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation
von: Wu, Ning, et al.
Veröffentlicht: (2026)
von: Wu, Ning, et al.
Veröffentlicht: (2026)
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
von: Tang, Lexiang, et al.
Veröffentlicht: (2025)
von: Tang, Lexiang, et al.
Veröffentlicht: (2025)
Learning Fine-Grained Grounded Citations for Attributed Large Language Models
von: Huang, Lei, et al.
Veröffentlicht: (2024)
von: Huang, Lei, et al.
Veröffentlicht: (2024)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
von: Zhong, Weihong, et al.
Veröffentlicht: (2024)
von: Zhong, Weihong, et al.
Veröffentlicht: (2024)
Not All Qubits are Utilized Equally
von: Pope, Jeremie, et al.
Veröffentlicht: (2025)
von: Pope, Jeremie, et al.
Veröffentlicht: (2025)
Not All Alumni Engage Equally
Veröffentlicht: (2024)
Veröffentlicht: (2024)
Not All Data Are Unlearned Equally
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective
von: Yu, Xinmiao, et al.
Veröffentlicht: (2024)
von: Yu, Xinmiao, et al.
Veröffentlicht: (2024)
Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization
von: Qi, Zipeng, et al.
Veröffentlicht: (2024)
von: Qi, Zipeng, et al.
Veröffentlicht: (2024)
CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs
von: Ye, Yangfan, et al.
Veröffentlicht: (2026)
von: Ye, Yangfan, et al.
Veröffentlicht: (2026)
Efficient NeRF Optimization -- Not All Samples Remain Equally Hard
von: Korhonen, Juuso, et al.
Veröffentlicht: (2024)
von: Korhonen, Juuso, et al.
Veröffentlicht: (2024)
LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning
von: Ye, Yangfan, et al.
Veröffentlicht: (2025)
von: Ye, Yangfan, et al.
Veröffentlicht: (2025)
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models
von: Lyu, Zesen, et al.
Veröffentlicht: (2025)
von: Lyu, Zesen, et al.
Veröffentlicht: (2025)
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
von: Tyagi, Utkarsh, et al.
Veröffentlicht: (2026)
von: Tyagi, Utkarsh, et al.
Veröffentlicht: (2026)
x1: Learning to Think Adaptively Across Languages and Cultures
von: Ye, Yangfan, et al.
Veröffentlicht: (2026)
von: Ye, Yangfan, et al.
Veröffentlicht: (2026)
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
von: Jeffares, Alan, et al.
Veröffentlicht: (2025)
von: Jeffares, Alan, et al.
Veröffentlicht: (2025)
Not All Samples Should Be Utilized Equally: Towards Understanding and Improving Dataset Distillation
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
von: Feng, Xiachong, et al.
Veröffentlicht: (2026)
von: Feng, Xiachong, et al.
Veröffentlicht: (2026)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
von: Peng, Cihang, et al.
Veröffentlicht: (2025)
von: Peng, Cihang, et al.
Veröffentlicht: (2025)
Enhancing Non-English Capabilities of English-Centric Large Language Models through Deep Supervision Fine-Tuning
von: Huo, Wenshuai, et al.
Veröffentlicht: (2025)
von: Huo, Wenshuai, et al.
Veröffentlicht: (2025)
Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
von: Li, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Li, Kaiyuan, et al.
Veröffentlicht: (2025)
Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect Clustering
von: Zhu, Kun, et al.
Veröffentlicht: (2025)
von: Zhu, Kun, et al.
Veröffentlicht: (2025)
NCT-CRC-HE: Not All Histopathological Datasets Are Equally Useful
von: Ignatov, Andrey, et al.
Veröffentlicht: (2024)
von: Ignatov, Andrey, et al.
Veröffentlicht: (2024)
Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
von: Li, Rong, et al.
Veröffentlicht: (2024)
von: Li, Rong, et al.
Veröffentlicht: (2024)
Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
von: Luo, Wenyang, et al.
Veröffentlicht: (2025)
von: Luo, Wenyang, et al.
Veröffentlicht: (2025)
SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization
von: Wang, Zhengcheng, et al.
Veröffentlicht: (2025)
von: Wang, Zhengcheng, et al.
Veröffentlicht: (2025)
Not All Instances Are Equally Valuable: Towards Influence-Weighted Dataset Distillation
von: Deng, Qiyan, et al.
Veröffentlicht: (2025)
von: Deng, Qiyan, et al.
Veröffentlicht: (2025)
Not All Factors Crowd Equally: Modeling, Measuring, and Trading on Alpha Decay
von: Lee, Chorok
Veröffentlicht: (2025)
von: Lee, Chorok
Veröffentlicht: (2025)
One for All: Update Parameterized Knowledge Across Multiple Models
von: Ma, Weitao, et al.
Veröffentlicht: (2025)
von: Ma, Weitao, et al.
Veröffentlicht: (2025)
Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth Routing
von: He, Jinmin, et al.
Veröffentlicht: (2023)
von: He, Jinmin, et al.
Veröffentlicht: (2023)
All-in-One: Transferring Vision Foundation Models into Stereo Matching
von: Zhou, Jingyi, et al.
Veröffentlicht: (2024)
von: Zhou, Jingyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
von: Chen, Ruihan, et al.
Veröffentlicht: (2025) -
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
von: Li, Qiming, et al.
Veröffentlicht: (2025) -
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
von: Li, Qiming, et al.
Veröffentlicht: (2025) -
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
von: Li, Qiming, et al.
Veröffentlicht: (2026) -
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
von: Li, Qiming, et al.
Veröffentlicht: (2025)