Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Zhecheng, Song, Guoxian, Cai, Yujun, Xiong, Zhen, Yuan, Junsong, Wang, Yiwei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
$A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
di: Li, Zhecheng, et al.
Pubblicazione: (2025)
di: Li, Zhecheng, et al.
Pubblicazione: (2025)
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
di: Li, Zhecheng, et al.
Pubblicazione: (2025)
di: Li, Zhecheng, et al.
Pubblicazione: (2025)
Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
di: Li, Sifan, et al.
Pubblicazione: (2025)
di: Li, Sifan, et al.
Pubblicazione: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
di: Li, Sifan, et al.
Pubblicazione: (2025)
di: Li, Sifan, et al.
Pubblicazione: (2025)
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
Lost in Embeddings: Information Loss in Vision-Language Models
di: Li, Wenyan, et al.
Pubblicazione: (2025)
di: Li, Wenyan, et al.
Pubblicazione: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
di: Tao, Xingjian, et al.
Pubblicazione: (2026)
di: Tao, Xingjian, et al.
Pubblicazione: (2026)
Unveiling the Potential of Diffusion Large Language Model in Controllable Generation
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
Towards Zero-Shot Annotation of the Built Environment with Vision-Language Models (Vision Paper)
di: Han, Bin, et al.
Pubblicazione: (2024)
di: Han, Bin, et al.
Pubblicazione: (2024)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
di: Wang, Zhaochen, et al.
Pubblicazione: (2025)
di: Wang, Zhaochen, et al.
Pubblicazione: (2025)
Lost in Edits? A $λ$-Compass for AIGC Provenance
di: You, Wenhao, et al.
Pubblicazione: (2025)
di: You, Wenhao, et al.
Pubblicazione: (2025)
Large Vision-Language Models Get Lost in Attention
di: Xi, Gongli, et al.
Pubblicazione: (2026)
di: Xi, Gongli, et al.
Pubblicazione: (2026)
Evaluating Vision-Language Models for Emotion Recognition
di: Bhattacharyya, Sree, et al.
Pubblicazione: (2025)
di: Bhattacharyya, Sree, et al.
Pubblicazione: (2025)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
di: Deng, Ken, et al.
Pubblicazione: (2026)
di: Deng, Ken, et al.
Pubblicazione: (2026)
Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
di: Qu, Tianyuan, et al.
Pubblicazione: (2025)
di: Qu, Tianyuan, et al.
Pubblicazione: (2025)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
di: Jiang, Lei, et al.
Pubblicazione: (2025)
di: Jiang, Lei, et al.
Pubblicazione: (2025)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
di: Lu, Jiaying, et al.
Pubblicazione: (2023)
di: Lu, Jiaying, et al.
Pubblicazione: (2023)
Shape and Texture Recognition in Large Vision-Language Models
di: Eppel, Sagi, et al.
Pubblicazione: (2025)
di: Eppel, Sagi, et al.
Pubblicazione: (2025)
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
di: Cai, Hengxing, et al.
Pubblicazione: (2025)
di: Cai, Hengxing, et al.
Pubblicazione: (2025)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
di: Wu, Ruijia, et al.
Pubblicazione: (2025)
di: Wu, Ruijia, et al.
Pubblicazione: (2025)
Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLM
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
AI Based Font Pair Suggestion Modelling For Graphic Design
di: Singh, Aryan, et al.
Pubblicazione: (2025)
di: Singh, Aryan, et al.
Pubblicazione: (2025)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
di: Kim, Jeonghwan, et al.
Pubblicazione: (2024)
di: Kim, Jeonghwan, et al.
Pubblicazione: (2024)
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
di: Castro, Santiago, et al.
Pubblicazione: (2024)
di: Castro, Santiago, et al.
Pubblicazione: (2024)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
di: He, Xingwei, et al.
Pubblicazione: (2024)
di: He, Xingwei, et al.
Pubblicazione: (2024)
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
di: Li, Yue, et al.
Pubblicazione: (2025)
di: Li, Yue, et al.
Pubblicazione: (2025)
Mitigating Coordinate Prediction Bias from Positional Encoding Failures
di: Tao, Xingjian, et al.
Pubblicazione: (2025)
di: Tao, Xingjian, et al.
Pubblicazione: (2025)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
di: Shahgir, Haz Sameen, et al.
Pubblicazione: (2026)
di: Shahgir, Haz Sameen, et al.
Pubblicazione: (2026)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
di: Li, Bin, et al.
Pubblicazione: (2025)
di: Li, Bin, et al.
Pubblicazione: (2025)
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
di: Wu, Hang, et al.
Pubblicazione: (2026)
di: Wu, Hang, et al.
Pubblicazione: (2026)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
di: Wang, Shengkang, et al.
Pubblicazione: (2024)
di: Wang, Shengkang, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
di: Lu, Yifan, et al.
Pubblicazione: (2025)
di: Lu, Yifan, et al.
Pubblicazione: (2025)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
di: Gu, Jihao, et al.
Pubblicazione: (2025)
di: Gu, Jihao, et al.
Pubblicazione: (2025)
Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images
di: Mahanta, Cristina, et al.
Pubblicazione: (2025)
di: Mahanta, Cristina, et al.
Pubblicazione: (2025)
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
di: Xu, Yige, et al.
Pubblicazione: (2026)
di: Xu, Yige, et al.
Pubblicazione: (2026)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
di: Liu, Zheng, et al.
Pubblicazione: (2024)
di: Liu, Zheng, et al.
Pubblicazione: (2024)
ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations
di: Jiang, Bowen, et al.
Pubblicazione: (2025)
di: Jiang, Bowen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
$A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
di: Li, Zhecheng, et al.
Pubblicazione: (2025) -
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
di: Li, Zhecheng, et al.
Pubblicazione: (2025) -
Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
di: Li, Sifan, et al.
Pubblicazione: (2025) -
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
di: Li, Sifan, et al.
Pubblicazione: (2025) -
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
di: Xiong, Zhen, et al.
Pubblicazione: (2025)