Mitigating Coordinate Prediction Bias from Positional Encoding Failures
Fuente:
arXiv
Salvato in:
| Autori principali: | Tao, Xingjian, Wang, Yiwei, Cai, Yujun, Luo, Yihong, Han, Kai, Tang, Jing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
di: Tao, Xingjian, et al.
Pubblicazione: (2026)
di: Tao, Xingjian, et al.
Pubblicazione: (2026)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
di: Wu, Yike, et al.
Pubblicazione: (2025)
di: Wu, Yike, et al.
Pubblicazione: (2025)
Energy-Calibrated VAE with Test Time Free Lunch
di: Luo, Yihong, et al.
Pubblicazione: (2023)
di: Luo, Yihong, et al.
Pubblicazione: (2023)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
di: Sun, Bowen, et al.
Pubblicazione: (2025)
di: Sun, Bowen, et al.
Pubblicazione: (2025)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
di: Ge, Haonan, et al.
Pubblicazione: (2025)
di: Ge, Haonan, et al.
Pubblicazione: (2025)
SeqPE: Transformer with Sequential Position Encoding
di: Li, Huayang, et al.
Pubblicazione: (2025)
di: Li, Huayang, et al.
Pubblicazione: (2025)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
di: Wang, Zhaochen, et al.
Pubblicazione: (2025)
di: Wang, Zhaochen, et al.
Pubblicazione: (2025)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
di: Chen, Zhanpeng, et al.
Pubblicazione: (2025)
di: Chen, Zhanpeng, et al.
Pubblicazione: (2025)
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
di: Lan, Jian, et al.
Pubblicazione: (2025)
di: Lan, Jian, et al.
Pubblicazione: (2025)
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions
di: Su, Junhao, et al.
Pubblicazione: (2025)
di: Su, Junhao, et al.
Pubblicazione: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
di: Ge, Haonan, et al.
Pubblicazione: (2025)
di: Ge, Haonan, et al.
Pubblicazione: (2025)
TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
di: Luo, Yihong, et al.
Pubblicazione: (2026)
di: Luo, Yihong, et al.
Pubblicazione: (2026)
Finding Distributed Object-Centric Properties in Self-Supervised Transformers
di: Rawlekar, Samyak, et al.
Pubblicazione: (2026)
di: Rawlekar, Samyak, et al.
Pubblicazione: (2026)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
di: Li, Bingxuan, et al.
Pubblicazione: (2025)
di: Li, Bingxuan, et al.
Pubblicazione: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
di: Li, Shuo, et al.
Pubblicazione: (2025)
di: Li, Shuo, et al.
Pubblicazione: (2025)
Mitigating Bias with Words: Inducing Demographic Ambiguity in Face Recognition Templates by Text Encoding
di: Chettaoui, Tahar, et al.
Pubblicazione: (2025)
di: Chettaoui, Tahar, et al.
Pubblicazione: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
di: Li, Sifan, et al.
Pubblicazione: (2025)
di: Li, Sifan, et al.
Pubblicazione: (2025)
Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
di: Cheng, Jiali, et al.
Pubblicazione: (2025)
di: Cheng, Jiali, et al.
Pubblicazione: (2025)
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
di: Kim, Minkyoung, et al.
Pubblicazione: (2024)
di: Kim, Minkyoung, et al.
Pubblicazione: (2024)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
di: Zhang, Wanpeng, et al.
Pubblicazione: (2024)
di: Zhang, Wanpeng, et al.
Pubblicazione: (2024)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
di: Chen, Zeyu, et al.
Pubblicazione: (2026)
di: Chen, Zeyu, et al.
Pubblicazione: (2026)
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection
di: La, Tuan-Vinh, et al.
Pubblicazione: (2025)
di: La, Tuan-Vinh, et al.
Pubblicazione: (2025)
Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
di: Omasa, Takamitsu, et al.
Pubblicazione: (2025)
di: Omasa, Takamitsu, et al.
Pubblicazione: (2025)
Mitigating Query Selection Bias in Referring Video Object Segmentation
di: Zhang, Dingwei, et al.
Pubblicazione: (2025)
di: Zhang, Dingwei, et al.
Pubblicazione: (2025)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
di: Elobaid, Alaa
Pubblicazione: (2026)
di: Elobaid, Alaa
Pubblicazione: (2026)
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
di: Zhao, Dachuan, et al.
Pubblicazione: (2025)
di: Zhao, Dachuan, et al.
Pubblicazione: (2025)
Cameras as Relative Positional Encoding
di: Li, Ruilong, et al.
Pubblicazione: (2025)
di: Li, Ruilong, et al.
Pubblicazione: (2025)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
di: Ye, Zekai, et al.
Pubblicazione: (2025)
di: Ye, Zekai, et al.
Pubblicazione: (2025)
Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation
di: Tang, Tao, et al.
Pubblicazione: (2025)
di: Tang, Tao, et al.
Pubblicazione: (2025)
DreamAvoid: Critical-Phase Test-Time Dreaming to Avoid Failures in VLA Policies
di: Fan, Xianzhe, et al.
Pubblicazione: (2026)
di: Fan, Xianzhe, et al.
Pubblicazione: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
di: Fu, Honghao, et al.
Pubblicazione: (2026)
di: Fu, Honghao, et al.
Pubblicazione: (2026)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
di: Wu, Hang, et al.
Pubblicazione: (2026)
di: Wu, Hang, et al.
Pubblicazione: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
di: Wu, Qianhui, et al.
Pubblicazione: (2025)
di: Wu, Qianhui, et al.
Pubblicazione: (2025)
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
di: Huang, Yihong, et al.
Pubblicazione: (2026)
di: Huang, Yihong, et al.
Pubblicazione: (2026)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
di: Lim, Junyoung, et al.
Pubblicazione: (2025)
di: Lim, Junyoung, et al.
Pubblicazione: (2025)
Mitigating Multilingual Hallucination in Large Vision-Language Models
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
di: Singh, Anshul, et al.
Pubblicazione: (2025)
di: Singh, Anshul, et al.
Pubblicazione: (2025)
BAMI: Training-Free Bias Mitigation in GUI Grounding
di: Zhang, Borui, et al.
Pubblicazione: (2026)
di: Zhang, Borui, et al.
Pubblicazione: (2026)
Weierstrass Positional Encoding for Vision Transformers
di: Xin, Zhihang, et al.
Pubblicazione: (2026)
di: Xin, Zhihang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
di: Tao, Xingjian, et al.
Pubblicazione: (2026) -
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
di: Wu, Yike, et al.
Pubblicazione: (2025) -
Energy-Calibrated VAE with Test Time Free Lunch
di: Luo, Yihong, et al.
Pubblicazione: (2023) -
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
di: Sun, Bowen, et al.
Pubblicazione: (2025) -
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
di: Ge, Haonan, et al.
Pubblicazione: (2025)