EAGLE: Elevating Geometric Reasoning through LLM-empowered Visual Instruction Tuning
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Zhihao, Du, Yao, Liu, Yang, Zhang, Yan, Liu, Yufang, Zhang, Mengdi, Cai, Xunliang, Ling, Charles, Wang, Boyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
di: Liu, Yufang, et al.
Pubblicazione: (2025)
di: Liu, Yufang, et al.
Pubblicazione: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
di: Du, Yifan, et al.
Pubblicazione: (2023)
di: Du, Yifan, et al.
Pubblicazione: (2023)
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
di: Duan, Chengqi, et al.
Pubblicazione: (2025)
EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
di: Villa, Andrés, et al.
Pubblicazione: (2025)
di: Villa, Andrés, et al.
Pubblicazione: (2025)
Elevating Visual Question Answering through Implicitly Learned Reasoning Pathways in LVLMs
di: Jing, Liu, et al.
Pubblicazione: (2025)
di: Jing, Liu, et al.
Pubblicazione: (2025)
Personalized Visual Instruction Tuning
di: Pi, Renjie, et al.
Pubblicazione: (2024)
di: Pi, Renjie, et al.
Pubblicazione: (2024)
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
di: Jia, Yongju, et al.
Pubblicazione: (2025)
di: Jia, Yongju, et al.
Pubblicazione: (2025)
Prejudge-Before-Think: Enhancing Large Language Models at Test-Time by Process Prejudge Reasoning
di: Wang, Jianing, et al.
Pubblicazione: (2025)
di: Wang, Jianing, et al.
Pubblicazione: (2025)
Osprey: Pixel Understanding with Visual Instruction Tuning
di: Yuan, Yuqian, et al.
Pubblicazione: (2023)
di: Yuan, Yuqian, et al.
Pubblicazione: (2023)
Learning to Instruct for Visual Instruction Tuning
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
Reconstructive Visual Instruction Tuning
di: Wang, Haochen, et al.
Pubblicazione: (2024)
di: Wang, Haochen, et al.
Pubblicazione: (2024)
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
di: Zhang, Renrui, et al.
Pubblicazione: (2024)
EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models
di: Zhang, Jiacheng, et al.
Pubblicazione: (2024)
di: Zhang, Jiacheng, et al.
Pubblicazione: (2024)
Visual Instruction Tuning with Chain of Region-of-Interest
di: Chen, Yixin, et al.
Pubblicazione: (2025)
di: Chen, Yixin, et al.
Pubblicazione: (2025)
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
di: Vogel, Alexander, et al.
Pubblicazione: (2025)
di: Vogel, Alexander, et al.
Pubblicazione: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
di: Zhang, Xiaoman, et al.
Pubblicazione: (2023)
di: Zhang, Xiaoman, et al.
Pubblicazione: (2023)
LLaVA-c: Continual Improved Visual Instruction Tuning
di: Liu, Wenzhuo, et al.
Pubblicazione: (2025)
di: Liu, Wenzhuo, et al.
Pubblicazione: (2025)
Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning
di: Gou, Yunhao, et al.
Pubblicazione: (2025)
di: Gou, Yunhao, et al.
Pubblicazione: (2025)
EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision
di: Cao, Yifei, et al.
Pubblicazione: (2025)
di: Cao, Yifei, et al.
Pubblicazione: (2025)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
di: li, Bonan, et al.
Pubblicazione: (2025)
di: li, Bonan, et al.
Pubblicazione: (2025)
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
di: Chen, Zhangquan, et al.
Pubblicazione: (2025)
Improved Baselines with Visual Instruction Tuning
di: Liu, Haotian, et al.
Pubblicazione: (2023)
di: Liu, Haotian, et al.
Pubblicazione: (2023)
Generative Visual Instruction Tuning
di: Hernandez, Jefferson, et al.
Pubblicazione: (2024)
di: Hernandez, Jefferson, et al.
Pubblicazione: (2024)
Comparison Visual Instruction Tuning
di: Lin, Wei, et al.
Pubblicazione: (2024)
di: Lin, Wei, et al.
Pubblicazione: (2024)
GloFinder: AI-empowered QuPath Plugin for WSI-level Glomerular Detection, Visualization, and Curation
di: Yue, Jialin, et al.
Pubblicazione: (2024)
di: Yue, Jialin, et al.
Pubblicazione: (2024)
Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning
di: Yang, Zexian, et al.
Pubblicazione: (2025)
di: Yang, Zexian, et al.
Pubblicazione: (2025)
Parrot: Multilingual Visual Instruction Tuning
di: Sun, Hai-Long, et al.
Pubblicazione: (2024)
di: Sun, Hai-Long, et al.
Pubblicazione: (2024)
Visual Generation Tuning
di: Guo, Jiahao, et al.
Pubblicazione: (2025)
di: Guo, Jiahao, et al.
Pubblicazione: (2025)
AutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text Interaction
di: Yang, Jiashu, et al.
Pubblicazione: (2026)
di: Yang, Jiashu, et al.
Pubblicazione: (2026)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
di: Chang, Boyu, et al.
Pubblicazione: (2026)
di: Chang, Boyu, et al.
Pubblicazione: (2026)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
di: Wang, Dongsheng, et al.
Pubblicazione: (2024)
di: Wang, Dongsheng, et al.
Pubblicazione: (2024)
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
di: Tang, Jingqun, et al.
Pubblicazione: (2024)
di: Tang, Jingqun, et al.
Pubblicazione: (2024)
Class Overwhelms: Mutual Conditional Blended-Target Domain Adaptation
di: Xu, Pengcheng, et al.
Pubblicazione: (2023)
di: Xu, Pengcheng, et al.
Pubblicazione: (2023)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
di: Peng, Wujian, et al.
Pubblicazione: (2024)
di: Peng, Wujian, et al.
Pubblicazione: (2024)
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
di: Zeng, Fanhu, et al.
Pubblicazione: (2024)
di: Zeng, Fanhu, et al.
Pubblicazione: (2024)
GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning
di: Sun, Jiayin, et al.
Pubblicazione: (2026)
di: Sun, Jiayin, et al.
Pubblicazione: (2026)
Do Large Language Models Excel in Complex Logical Reasoning with Formal Language?
di: Jiang, Jin, et al.
Pubblicazione: (2025)
di: Jiang, Jin, et al.
Pubblicazione: (2025)
Improving Medical Visual Reinforcement Fine-Tuning via Perception and Reasoning Augmentation
di: Yang, Guangjing, et al.
Pubblicazione: (2026)
di: Yang, Guangjing, et al.
Pubblicazione: (2026)
Visual Autoregressive Modeling for Instruction-Guided Image Editing
di: Mao, Qingyang, et al.
Pubblicazione: (2025)
di: Mao, Qingyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
di: Liu, Yufang, et al.
Pubblicazione: (2025) -
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
di: Du, Yifan, et al.
Pubblicazione: (2023) -
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
di: Duan, Chengqi, et al.
Pubblicazione: (2025) -
EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
di: Villa, Andrés, et al.
Pubblicazione: (2025) -
Elevating Visual Question Answering through Implicitly Learned Reasoning Pathways in LVLMs
di: Jing, Liu, et al.
Pubblicazione: (2025)