Mitigating Coordinate Prediction Bias from Positional Encoding Failures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tao, Xingjian, Wang, Yiwei, Cai, Yujun, Luo, Yihong, Han, Kai, Tang, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
von: Wu, Yike, et al.
Veröffentlicht: (2025)
von: Wu, Yike, et al.
Veröffentlicht: (2025)
Energy-Calibrated VAE with Test Time Free Lunch
von: Luo, Yihong, et al.
Veröffentlicht: (2023)
von: Luo, Yihong, et al.
Veröffentlicht: (2023)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
SeqPE: Transformer with Sequential Position Encoding
von: Li, Huayang, et al.
Veröffentlicht: (2025)
von: Li, Huayang, et al.
Veröffentlicht: (2025)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
von: Lan, Jian, et al.
Veröffentlicht: (2025)
von: Lan, Jian, et al.
Veröffentlicht: (2025)
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions
von: Su, Junhao, et al.
Veröffentlicht: (2025)
von: Su, Junhao, et al.
Veröffentlicht: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
von: Luo, Yihong, et al.
Veröffentlicht: (2026)
von: Luo, Yihong, et al.
Veröffentlicht: (2026)
Finding Distributed Object-Centric Properties in Self-Supervised Transformers
von: Rawlekar, Samyak, et al.
Veröffentlicht: (2026)
von: Rawlekar, Samyak, et al.
Veröffentlicht: (2026)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
Mitigating Bias with Words: Inducing Demographic Ambiguity in Face Recognition Templates by Text Encoding
von: Chettaoui, Tahar, et al.
Veröffentlicht: (2025)
von: Chettaoui, Tahar, et al.
Veröffentlicht: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
von: Li, Sifan, et al.
Veröffentlicht: (2025)
von: Li, Sifan, et al.
Veröffentlicht: (2025)
Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
von: Kim, Minkyoung, et al.
Veröffentlicht: (2024)
von: Kim, Minkyoung, et al.
Veröffentlicht: (2024)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2024)
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2024)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection
von: La, Tuan-Vinh, et al.
Veröffentlicht: (2025)
von: La, Tuan-Vinh, et al.
Veröffentlicht: (2025)
Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
von: Omasa, Takamitsu, et al.
Veröffentlicht: (2025)
von: Omasa, Takamitsu, et al.
Veröffentlicht: (2025)
Mitigating Query Selection Bias in Referring Video Object Segmentation
von: Zhang, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhang, Dingwei, et al.
Veröffentlicht: (2025)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
von: Elobaid, Alaa
Veröffentlicht: (2026)
von: Elobaid, Alaa
Veröffentlicht: (2026)
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
von: Zhao, Dachuan, et al.
Veröffentlicht: (2025)
von: Zhao, Dachuan, et al.
Veröffentlicht: (2025)
Cameras as Relative Positional Encoding
von: Li, Ruilong, et al.
Veröffentlicht: (2025)
von: Li, Ruilong, et al.
Veröffentlicht: (2025)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
von: Ye, Zekai, et al.
Veröffentlicht: (2025)
von: Ye, Zekai, et al.
Veröffentlicht: (2025)
Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
DreamAvoid: Critical-Phase Test-Time Dreaming to Avoid Failures in VLA Policies
von: Fan, Xianzhe, et al.
Veröffentlicht: (2026)
von: Fan, Xianzhe, et al.
Veröffentlicht: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
von: Wu, Hang, et al.
Veröffentlicht: (2026)
von: Wu, Hang, et al.
Veröffentlicht: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
von: Huang, Yihong, et al.
Veröffentlicht: (2026)
von: Huang, Yihong, et al.
Veröffentlicht: (2026)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
von: Lim, Junyoung, et al.
Veröffentlicht: (2025)
Mitigating Multilingual Hallucination in Large Vision-Language Models
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
von: Singh, Anshul, et al.
Veröffentlicht: (2025)
von: Singh, Anshul, et al.
Veröffentlicht: (2025)
BAMI: Training-Free Bias Mitigation in GUI Grounding
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
Weierstrass Positional Encoding for Vision Transformers
von: Xin, Zhihang, et al.
Veröffentlicht: (2026)
von: Xin, Zhihang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026) -
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
von: Wu, Yike, et al.
Veröffentlicht: (2025) -
Energy-Calibrated VAE with Test Time Free Lunch
von: Luo, Yihong, et al.
Veröffentlicht: (2023) -
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
von: Sun, Bowen, et al.
Veröffentlicht: (2025) -
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
von: Ge, Haonan, et al.
Veröffentlicht: (2025)