Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Qiming, Feng, Xiaocheng, Ma, Yixuan, Ye, Zekai, Chen, Ruihan, Feng, Xiachong, Qin, Bing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
di: Li, Qiming, et al.
Pubblicazione: (2025)
di: Li, Qiming, et al.
Pubblicazione: (2025)
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
di: Li, Qiming, et al.
Pubblicazione: (2025)
di: Li, Qiming, et al.
Pubblicazione: (2025)
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
di: Li, Qiming, et al.
Pubblicazione: (2026)
di: Li, Qiming, et al.
Pubblicazione: (2026)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
di: Ye, Zekai, et al.
Pubblicazione: (2025)
di: Ye, Zekai, et al.
Pubblicazione: (2025)
Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection
di: de Senneville, Adhemar, et al.
Pubblicazione: (2026)
di: de Senneville, Adhemar, et al.
Pubblicazione: (2026)
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective
di: Yu, Xinmiao, et al.
Pubblicazione: (2024)
di: Yu, Xinmiao, et al.
Pubblicazione: (2024)
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
di: He, Zhendong, et al.
Pubblicazione: (2026)
di: He, Zhendong, et al.
Pubblicazione: (2026)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
di: Leng, Jixuan, et al.
Pubblicazione: (2025)
di: Leng, Jixuan, et al.
Pubblicazione: (2025)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
di: Chen, Qiguang, et al.
Pubblicazione: (2026)
di: Chen, Qiguang, et al.
Pubblicazione: (2026)
Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
di: Wang, Chaoyang, et al.
Pubblicazione: (2025)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
Elevating Visual Question Answering through Implicitly Learned Reasoning Pathways in LVLMs
di: Jing, Liu, et al.
Pubblicazione: (2025)
di: Jing, Liu, et al.
Pubblicazione: (2025)
See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
di: Zhang, Yongchang, et al.
Pubblicazione: (2026)
di: Zhang, Yongchang, et al.
Pubblicazione: (2026)
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
di: Jeon, Yerim, et al.
Pubblicazione: (2025)
di: Jeon, Yerim, et al.
Pubblicazione: (2025)
Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection
di: Yu, Peipeng, et al.
Pubblicazione: (2025)
di: Yu, Peipeng, et al.
Pubblicazione: (2025)
Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement
di: Qin, Zhenxin, et al.
Pubblicazione: (2026)
di: Qin, Zhenxin, et al.
Pubblicazione: (2026)
Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
di: Li, Hengzhuang, et al.
Pubblicazione: (2025)
di: Li, Hengzhuang, et al.
Pubblicazione: (2025)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
di: Li, Hongyu, et al.
Pubblicazione: (2025)
di: Li, Hongyu, et al.
Pubblicazione: (2025)
SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification
di: Zhang, Jiacheng, et al.
Pubblicazione: (2026)
di: Zhang, Jiacheng, et al.
Pubblicazione: (2026)
Interpretable Oracle Bone Script Decipherment through Radical and Pictographic Analysis with LVLMs
di: Peng, Kaixin, et al.
Pubblicazione: (2025)
di: Peng, Kaixin, et al.
Pubblicazione: (2025)
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification
di: Sun, Han, et al.
Pubblicazione: (2026)
di: Sun, Han, et al.
Pubblicazione: (2026)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
di: Zhang, Xiaofeng, et al.
Pubblicazione: (2024)
di: Zhang, Xiaofeng, et al.
Pubblicazione: (2024)
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
di: Li, Wenhao, et al.
Pubblicazione: (2025)
di: Li, Wenhao, et al.
Pubblicazione: (2025)
DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
di: Chen, Qian, et al.
Pubblicazione: (2025)
di: Chen, Qian, et al.
Pubblicazione: (2025)
Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs
di: Yu, Liu, et al.
Pubblicazione: (2025)
di: Yu, Liu, et al.
Pubblicazione: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
di: Guo, Zichun, et al.
Pubblicazione: (2026)
di: Guo, Zichun, et al.
Pubblicazione: (2026)
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
di: Sun, Zhichao, et al.
Pubblicazione: (2026)
di: Sun, Zhichao, et al.
Pubblicazione: (2026)
Evidence Packing for Cross-Domain Image Deepfake Detection with LVLMs
di: Liu, Yuxin, et al.
Pubblicazione: (2026)
di: Liu, Yuxin, et al.
Pubblicazione: (2026)
The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment
di: Liu, Hongyuan, et al.
Pubblicazione: (2026)
di: Liu, Hongyuan, et al.
Pubblicazione: (2026)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
di: Zhang, Beichen, et al.
Pubblicazione: (2024)
di: Zhang, Beichen, et al.
Pubblicazione: (2024)
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
di: Chen, Tieyuan, et al.
Pubblicazione: (2024)
di: Chen, Tieyuan, et al.
Pubblicazione: (2024)
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
di: Chen, Yangneng, et al.
Pubblicazione: (2026)
di: Chen, Yangneng, et al.
Pubblicazione: (2026)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
di: Yuan, Fan, et al.
Pubblicazione: (2024)
di: Yuan, Fan, et al.
Pubblicazione: (2024)
VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
di: Liu, Tao, et al.
Pubblicazione: (2024)
di: Liu, Tao, et al.
Pubblicazione: (2024)
LVLMs as inspectors: an agentic framework for category-level structural defect annotation
di: Jiang, Sheng, et al.
Pubblicazione: (2025)
di: Jiang, Sheng, et al.
Pubblicazione: (2025)
Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
di: Sui, Xiangjie, et al.
Pubblicazione: (2025)
di: Sui, Xiangjie, et al.
Pubblicazione: (2025)
CoV: Chain-of-View Prompting for Spatial Reasoning
di: Zhao, Haoyu, et al.
Pubblicazione: (2026)
di: Zhao, Haoyu, et al.
Pubblicazione: (2026)
ContextGuard-LVLM: Enhancing News Veracity through Fine-grained Cross-modal Contextual Consistency Verification
di: Ma, Sihan, et al.
Pubblicazione: (2025)
di: Ma, Sihan, et al.
Pubblicazione: (2025)
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
di: Kan, Zhehan, et al.
Pubblicazione: (2025)
di: Kan, Zhehan, et al.
Pubblicazione: (2025)
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
di: Dai, Ziyun, et al.
Pubblicazione: (2025)
di: Dai, Ziyun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
di: Li, Qiming, et al.
Pubblicazione: (2025) -
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
di: Li, Qiming, et al.
Pubblicazione: (2025) -
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
di: Li, Qiming, et al.
Pubblicazione: (2026) -
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
di: Ye, Zekai, et al.
Pubblicazione: (2025) -
Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection
di: de Senneville, Adhemar, et al.
Pubblicazione: (2026)