Mitigating Coordinate Prediction Bias from Positional Encoding Failures
Fuente:
arXiv
Guardado en:
| Autores principales: | Tao, Xingjian, Wang, Yiwei, Cai, Yujun, Luo, Yihong, Han, Kai, Tang, Jing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
por: Tao, Xingjian, et al.
Publicado: (2026)
por: Tao, Xingjian, et al.
Publicado: (2026)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
por: Wu, Yike, et al.
Publicado: (2025)
por: Wu, Yike, et al.
Publicado: (2025)
Energy-Calibrated VAE with Test Time Free Lunch
por: Luo, Yihong, et al.
Publicado: (2023)
por: Luo, Yihong, et al.
Publicado: (2023)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
por: Sun, Bowen, et al.
Publicado: (2025)
por: Sun, Bowen, et al.
Publicado: (2025)
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
SeqPE: Transformer with Sequential Position Encoding
por: Li, Huayang, et al.
Publicado: (2025)
por: Li, Huayang, et al.
Publicado: (2025)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
por: Wang, Zhaochen, et al.
Publicado: (2025)
por: Wang, Zhaochen, et al.
Publicado: (2025)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
por: Chen, Zhanpeng, et al.
Publicado: (2025)
por: Chen, Zhanpeng, et al.
Publicado: (2025)
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
por: Lan, Jian, et al.
Publicado: (2025)
por: Lan, Jian, et al.
Publicado: (2025)
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions
por: Su, Junhao, et al.
Publicado: (2025)
por: Su, Junhao, et al.
Publicado: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
por: Luo, Yihong, et al.
Publicado: (2026)
por: Luo, Yihong, et al.
Publicado: (2026)
Finding Distributed Object-Centric Properties in Self-Supervised Transformers
por: Rawlekar, Samyak, et al.
Publicado: (2026)
por: Rawlekar, Samyak, et al.
Publicado: (2026)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
por: Wu, Chengyue, et al.
Publicado: (2024)
por: Wu, Chengyue, et al.
Publicado: (2024)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
por: Li, Bingxuan, et al.
Publicado: (2025)
por: Li, Bingxuan, et al.
Publicado: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
por: Li, Shuo, et al.
Publicado: (2025)
por: Li, Shuo, et al.
Publicado: (2025)
Mitigating Bias with Words: Inducing Demographic Ambiguity in Face Recognition Templates by Text Encoding
por: Chettaoui, Tahar, et al.
Publicado: (2025)
por: Chettaoui, Tahar, et al.
Publicado: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
por: Li, Sifan, et al.
Publicado: (2025)
por: Li, Sifan, et al.
Publicado: (2025)
Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
por: Cheng, Jiali, et al.
Publicado: (2025)
por: Cheng, Jiali, et al.
Publicado: (2025)
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
por: Kim, Minkyoung, et al.
Publicado: (2024)
por: Kim, Minkyoung, et al.
Publicado: (2024)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
por: Zhang, Wanpeng, et al.
Publicado: (2024)
por: Zhang, Wanpeng, et al.
Publicado: (2024)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
por: Chen, Zeyu, et al.
Publicado: (2026)
por: Chen, Zeyu, et al.
Publicado: (2026)
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection
por: La, Tuan-Vinh, et al.
Publicado: (2025)
por: La, Tuan-Vinh, et al.
Publicado: (2025)
Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
por: Omasa, Takamitsu, et al.
Publicado: (2025)
por: Omasa, Takamitsu, et al.
Publicado: (2025)
Mitigating Query Selection Bias in Referring Video Object Segmentation
por: Zhang, Dingwei, et al.
Publicado: (2025)
por: Zhang, Dingwei, et al.
Publicado: (2025)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
por: Elobaid, Alaa
Publicado: (2026)
por: Elobaid, Alaa
Publicado: (2026)
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
por: Zhao, Dachuan, et al.
Publicado: (2025)
por: Zhao, Dachuan, et al.
Publicado: (2025)
Cameras as Relative Positional Encoding
por: Li, Ruilong, et al.
Publicado: (2025)
por: Li, Ruilong, et al.
Publicado: (2025)
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
por: Ye, Zekai, et al.
Publicado: (2025)
por: Ye, Zekai, et al.
Publicado: (2025)
Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation
por: Tang, Tao, et al.
Publicado: (2025)
por: Tang, Tao, et al.
Publicado: (2025)
DreamAvoid: Critical-Phase Test-Time Dreaming to Avoid Failures in VLA Policies
por: Fan, Xianzhe, et al.
Publicado: (2026)
por: Fan, Xianzhe, et al.
Publicado: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
por: Fu, Honghao, et al.
Publicado: (2026)
por: Fu, Honghao, et al.
Publicado: (2026)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
por: Wu, Hang, et al.
Publicado: (2026)
por: Wu, Hang, et al.
Publicado: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
por: Wu, Qianhui, et al.
Publicado: (2025)
por: Wu, Qianhui, et al.
Publicado: (2025)
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
por: Huang, Yihong, et al.
Publicado: (2026)
por: Huang, Yihong, et al.
Publicado: (2026)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
por: Lim, Junyoung, et al.
Publicado: (2025)
por: Lim, Junyoung, et al.
Publicado: (2025)
Mitigating Multilingual Hallucination in Large Vision-Language Models
por: Qu, Xiaoye, et al.
Publicado: (2024)
por: Qu, Xiaoye, et al.
Publicado: (2024)
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
por: Singh, Anshul, et al.
Publicado: (2025)
por: Singh, Anshul, et al.
Publicado: (2025)
BAMI: Training-Free Bias Mitigation in GUI Grounding
por: Zhang, Borui, et al.
Publicado: (2026)
por: Zhang, Borui, et al.
Publicado: (2026)
Weierstrass Positional Encoding for Vision Transformers
por: Xin, Zhihang, et al.
Publicado: (2026)
por: Xin, Zhihang, et al.
Publicado: (2026)
Ejemplares similares
-
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
por: Tao, Xingjian, et al.
Publicado: (2026) -
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
por: Wu, Yike, et al.
Publicado: (2025) -
Energy-Calibrated VAE with Test Time Free Lunch
por: Luo, Yihong, et al.
Publicado: (2023) -
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
por: Sun, Bowen, et al.
Publicado: (2025) -
MRFD: Multi-Region Fusion Decoding with Self-Consistency for Mitigating Hallucinations in LVLMs
por: Ge, Haonan, et al.
Publicado: (2025)