$A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhecheng, Song, Guoxian, Wang, Yiwei, Xiong, Zhen, Yuan, Junsong, Cai, Yujun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
by: Li, Zhecheng, et al.
Published: (2025)
by: Li, Zhecheng, et al.
Published: (2025)
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
by: Li, Zhecheng, et al.
Published: (2025)
by: Li, Zhecheng, et al.
Published: (2025)
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
by: Xiong, Zhen, et al.
Published: (2025)
by: Xiong, Zhen, et al.
Published: (2025)
Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLM
by: Xiong, Zhen, et al.
Published: (2025)
by: Xiong, Zhen, et al.
Published: (2025)
Unveiling the Potential of Diffusion Large Language Model in Controllable Generation
by: Xiong, Zhen, et al.
Published: (2025)
by: Xiong, Zhen, et al.
Published: (2025)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
by: Xiong, Zhen, et al.
Published: (2025)
by: Xiong, Zhen, et al.
Published: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Visualizing 2x2 Normal-Form Games: twoxtwogame LaTeX Package
by: Marris, Luke, et al.
Published: (2024)
by: Marris, Luke, et al.
Published: (2024)
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
Vulnerability of LLMs to Vertically Aligned Text Manipulations
by: Li, Zhecheng, et al.
Published: (2024)
by: Li, Zhecheng, et al.
Published: (2024)
LaTeX Compilation: Challenges in the Era of LLMs
by: Liu, Tianyou, et al.
Published: (2026)
by: Liu, Tianyou, et al.
Published: (2026)
TeXBLEU: Automatic Metric for Evaluate LaTeX Format
by: Jung, Kyudan, et al.
Published: (2024)
by: Jung, Kyudan, et al.
Published: (2024)
Image-to-LaTeX Converter for Mathematical Formulas and Text
by: Gurgurov, Daniil, et al.
Published: (2024)
by: Gurgurov, Daniil, et al.
Published: (2024)
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
by: Wang, Chengye, et al.
Published: (2026)
by: Wang, Chengye, et al.
Published: (2026)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
LaTeXTrans: Structured LaTeX Translation with Multi-Agent Coordination
by: Zhu, Ziming, et al.
Published: (2025)
by: Zhu, Ziming, et al.
Published: (2025)
Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs
by: Dan, Nifu, et al.
Published: (2025)
by: Dan, Nifu, et al.
Published: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
by: Tao, Xingjian, et al.
Published: (2026)
by: Tao, Xingjian, et al.
Published: (2026)
AI-Friendly LaTeX: Using LaTeX Code as a Knowledge Source for Retrieval-Augmented Generation
by: Verhoeff, Tom
Published: (2026)
by: Verhoeff, Tom
Published: (2026)
DRS: Deep Question Reformulation With Structured Output
by: Li, Zhecheng, et al.
Published: (2024)
by: Li, Zhecheng, et al.
Published: (2024)
Integration of LaTeX formula in computer-based test application for academic purposes
by: Onyenwe, Ikechukwu E., et al.
Published: (2024)
by: Onyenwe, Ikechukwu E., et al.
Published: (2024)
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation
by: Gkritzali, Evangelia, et al.
Published: (2024)
by: Gkritzali, Evangelia, et al.
Published: (2024)
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
by: Kale, Sahil, et al.
Published: (2025)
by: Kale, Sahil, et al.
Published: (2025)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
by: Wu, Yike, et al.
Published: (2025)
by: Wu, Yike, et al.
Published: (2025)
Automated LaTeX Code Generation from Handwritten Math Expressions Using Vision Transformer
by: Sundararaj, Jayaprakash, et al.
Published: (2024)
by: Sundararaj, Jayaprakash, et al.
Published: (2024)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
by: Wang, Zhaochen, et al.
Published: (2025)
by: Wang, Zhaochen, et al.
Published: (2025)
Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
by: Yang, Zhengyuan, et al.
Published: (2023)
by: Yang, Zhengyuan, et al.
Published: (2023)
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
by: Korzh, Dmitrii, et al.
Published: (2025)
by: Korzh, Dmitrii, et al.
Published: (2025)
Img2CADSeq: Image-to-CAD Generation via Sequence-Based Diffusion
by: Tan, Shiyu, et al.
Published: (2026)
by: Tan, Shiyu, et al.
Published: (2026)
Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
by: Sun, Bowen, et al.
Published: (2025)
by: Sun, Bowen, et al.
Published: (2025)
Uni-Neur2Img: Unified Neural Signal-Guided Image Generation, Editing, and Stylization via Diffusion Transformers
by: Bai, Xiyue, et al.
Published: (2025)
by: Bai, Xiyue, et al.
Published: (2025)
Rendering LaTeX in R
by: Murrell, Paul
Published: (2025)
by: Murrell, Paul
Published: (2025)
OverCite: Add citations in LaTeX without leaving the editor
by: Shariat, Cheyanne
Published: (2026)
by: Shariat, Cheyanne
Published: (2026)
Think Carefully and Check Again! Meta-Generation Unlocking LLMs for Low-Resource Cross-Lingual Summarization
by: Li, Zhecheng, et al.
Published: (2024)
by: Li, Zhecheng, et al.
Published: (2024)
Lost in Edits? A $λ$-Compass for AIGC Provenance
by: You, Wenhao, et al.
Published: (2025)
by: You, Wenhao, et al.
Published: (2025)
X-Portrait: Expressive Portrait Animation with Hierarchical Motion Attention
by: Xie, You, et al.
Published: (2024)
by: Xie, You, et al.
Published: (2024)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
by: Ge, Haonan, et al.
Published: (2025)
by: Ge, Haonan, et al.
Published: (2025)
AutoDrive-R$^2$: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
by: Yuan, Zhenlong, et al.
Published: (2025)
by: Yuan, Zhenlong, et al.
Published: (2025)
AM^2-EmoJE: Adaptive Missing-Modality Emotion Recognition in Conversation via Joint Embedding Learning
by: Devulapally, Naresh Kumar, et al.
Published: (2024)
by: Devulapally, Naresh Kumar, et al.
Published: (2024)
Similar Items
-
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
by: Li, Zhecheng, et al.
Published: (2025) -
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
by: Li, Zhecheng, et al.
Published: (2025) -
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
by: Xiong, Zhen, et al.
Published: (2025) -
Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLM
by: Xiong, Zhen, et al.
Published: (2025) -
Unveiling the Potential of Diffusion Large Language Model in Controllable Generation
by: Xiong, Zhen, et al.
Published: (2025)