TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zijian, Zheng, Xuhui, Wu, Xuecheng, Peng, Chong, Cao, Xuezhi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
por: Zhang, Zijian, et al.
Publicado: (2025)
por: Zhang, Zijian, et al.
Publicado: (2025)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
por: He, Haibin, et al.
Publicado: (2025)
por: He, Haibin, et al.
Publicado: (2025)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
por: Xue, Wei, et al.
Publicado: (2026)
por: Xue, Wei, et al.
Publicado: (2026)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
por: Li, Yanshu, et al.
Publicado: (2025)
por: Li, Yanshu, et al.
Publicado: (2025)
Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search
por: Tan, Lei, et al.
Publicado: (2024)
por: Tan, Lei, et al.
Publicado: (2024)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
por: Pan, Kaihang, et al.
Publicado: (2025)
por: Pan, Kaihang, et al.
Publicado: (2025)
Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma Grading
por: Pan, Li, et al.
Publicado: (2024)
por: Pan, Li, et al.
Publicado: (2024)
Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation
por: Xing, Xiaoying, et al.
Publicado: (2025)
por: Xing, Xiaoying, et al.
Publicado: (2025)
Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios
por: Tang, Mingwei, et al.
Publicado: (2025)
por: Tang, Mingwei, et al.
Publicado: (2025)
Focusing Image Generation to Mitigate Spurious Correlations
por: Li, Xuewei, et al.
Publicado: (2024)
por: Li, Xuewei, et al.
Publicado: (2024)
Generative Multi-Focus Image Fusion
por: Xie, Xinzhe, et al.
Publicado: (2025)
por: Xie, Xinzhe, et al.
Publicado: (2025)
Focus-Consistent Multi-Level Aggregation for Compositional Zero-Shot Learning
por: Dai, Fengyuan, et al.
Publicado: (2024)
por: Dai, Fengyuan, et al.
Publicado: (2024)
DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
por: Cao, Yuhang, et al.
Publicado: (2024)
por: Cao, Yuhang, et al.
Publicado: (2024)
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning
por: Chang, Aofei, et al.
Publicado: (2025)
por: Chang, Aofei, et al.
Publicado: (2025)
DRWKV: Focusing on Object Edges for Low-Light Image Enhancement
por: Bai, Xuecheng, et al.
Publicado: (2025)
por: Bai, Xuecheng, et al.
Publicado: (2025)
Dual-Camera All-in-Focus Neural Radiance Fields
por: Luo, Xianrui, et al.
Publicado: (2025)
por: Luo, Xianrui, et al.
Publicado: (2025)
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
por: Ouyang, Mingyu, et al.
Publicado: (2026)
por: Ouyang, Mingyu, et al.
Publicado: (2026)
Multi-Focused Video Group Activities Hashing
por: Qi, Zhongmiao, et al.
Publicado: (2025)
por: Qi, Zhongmiao, et al.
Publicado: (2025)
CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolution
por: Liu, Xin, et al.
Publicado: (2025)
por: Liu, Xin, et al.
Publicado: (2025)
Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
por: Zhao, Zaiying, et al.
Publicado: (2025)
por: Zhao, Zaiying, et al.
Publicado: (2025)
MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation
por: Zhang, Guohui, et al.
Publicado: (2025)
por: Zhang, Guohui, et al.
Publicado: (2025)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
por: Shen, Meng, et al.
Publicado: (2026)
por: Shen, Meng, et al.
Publicado: (2026)
SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning
por: Chen, Chen, et al.
Publicado: (2025)
por: Chen, Chen, et al.
Publicado: (2025)
FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus
por: Jin, Qiaoqiao, et al.
Publicado: (2025)
por: Jin, Qiaoqiao, et al.
Publicado: (2025)
Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage
por: Xie, Junfei, et al.
Publicado: (2026)
por: Xie, Junfei, et al.
Publicado: (2026)
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
por: Zhu, Yuke, et al.
Publicado: (2024)
por: Zhu, Yuke, et al.
Publicado: (2024)
Addressing the Depth-of-Field Constraint: A New Paradigm for High Resolution Multi-Focus Image Fusion
por: Piano, Luca, et al.
Publicado: (2025)
por: Piano, Luca, et al.
Publicado: (2025)
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
por: Cheng, Zheng, et al.
Publicado: (2024)
por: Cheng, Zheng, et al.
Publicado: (2024)
Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
por: Wu, Yuanchen, et al.
Publicado: (2025)
por: Wu, Yuanchen, et al.
Publicado: (2025)
LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment
por: Hu, Junyi, et al.
Publicado: (2026)
por: Hu, Junyi, et al.
Publicado: (2026)
Minority-Focused Text-to-Image Generation via Prompt Optimization
por: Um, Soobin, et al.
Publicado: (2024)
por: Um, Soobin, et al.
Publicado: (2024)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
por: Chen, Zhangquan, et al.
Publicado: (2025)
por: Chen, Zhangquan, et al.
Publicado: (2025)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
por: He, Haibin, et al.
Publicado: (2026)
por: He, Haibin, et al.
Publicado: (2026)
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
por: Cao, Jianjian, et al.
Publicado: (2024)
por: Cao, Jianjian, et al.
Publicado: (2024)
A Lightweight Sparse Focus Transformer for Remote Sensing Image Change Captioning
por: Sun, Dongwei, et al.
Publicado: (2024)
por: Sun, Dongwei, et al.
Publicado: (2024)
FocusTrack: One-Stage Focus-and-Suppress Framework for 3D Point Cloud Object Tracking
por: Zhou, Sifan, et al.
Publicado: (2026)
por: Zhou, Sifan, et al.
Publicado: (2026)
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
por: Zhang, Yan, et al.
Publicado: (2025)
por: Zhang, Yan, et al.
Publicado: (2025)
Reinforcing Video Reasoning with Focused Thinking
por: Dang, Jisheng, et al.
Publicado: (2025)
por: Dang, Jisheng, et al.
Publicado: (2025)
Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation
por: Liu, Tzu Ling, et al.
Publicado: (2026)
por: Liu, Tzu Ling, et al.
Publicado: (2026)
TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
por: Kim, Jongha, et al.
Publicado: (2025)
por: Kim, Jongha, et al.
Publicado: (2025)
Ejemplares similares
-
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
por: Zhang, Zijian, et al.
Publicado: (2025) -
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
por: He, Haibin, et al.
Publicado: (2025) -
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
por: Xue, Wei, et al.
Publicado: (2026) -
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
por: Li, Yanshu, et al.
Publicado: (2025) -
Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search
por: Tan, Lei, et al.
Publicado: (2024)