Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Singh, Ayush, Gupta, Mansi, Garg, Shivank, Kumar, Abhinav, Agrawal, Vansh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Are VLMs Really Blind
di: Singh, Ayush, et al.
Pubblicazione: (2024)
di: Singh, Ayush, et al.
Pubblicazione: (2024)
Give me a hint: Can LLMs take a hint to solve math problems?
di: Agrawal, Vansh, et al.
Pubblicazione: (2024)
di: Agrawal, Vansh, et al.
Pubblicazione: (2024)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
di: Zhang, Di, et al.
Pubblicazione: (2024)
di: Zhang, Di, et al.
Pubblicazione: (2024)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
di: Siingh, Shikhhar, et al.
Pubblicazione: (2025)
di: Siingh, Shikhhar, et al.
Pubblicazione: (2025)
Re:Verse -- Can Your VLM Read a Manga?
di: Baranwal, Aaditya, et al.
Pubblicazione: (2025)
di: Baranwal, Aaditya, et al.
Pubblicazione: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
di: Liu, Zheng, et al.
Pubblicazione: (2024)
di: Liu, Zheng, et al.
Pubblicazione: (2024)
Unmasking the Veil: An Investigation into Concept Ablation for Privacy and Copyright Protection in Images
di: Garg, Shivank, et al.
Pubblicazione: (2024)
di: Garg, Shivank, et al.
Pubblicazione: (2024)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
di: Chen, Xiaofu, et al.
Pubblicazione: (2025)
di: Chen, Xiaofu, et al.
Pubblicazione: (2025)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
di: Mañas, Oscar, et al.
Pubblicazione: (2024)
di: Mañas, Oscar, et al.
Pubblicazione: (2024)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
di: Gwilliam, Matthew, et al.
Pubblicazione: (2023)
di: Gwilliam, Matthew, et al.
Pubblicazione: (2023)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
di: Merchant, Nicholas, et al.
Pubblicazione: (2025)
di: Merchant, Nicholas, et al.
Pubblicazione: (2025)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
di: Awal, Rabiul, et al.
Pubblicazione: (2023)
di: Awal, Rabiul, et al.
Pubblicazione: (2023)
Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
di: Berger, Uri, et al.
Pubblicazione: (2025)
di: Berger, Uri, et al.
Pubblicazione: (2025)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
di: Zhang, Wentao, et al.
Pubblicazione: (2025)
di: Zhang, Wentao, et al.
Pubblicazione: (2025)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
di: Ma, Ziyu, et al.
Pubblicazione: (2024)
di: Ma, Ziyu, et al.
Pubblicazione: (2024)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
di: Kumar, Divake, et al.
Pubblicazione: (2026)
di: Kumar, Divake, et al.
Pubblicazione: (2026)
OmniCaptioner: One Captioner to Rule Them All
di: Lu, Yiting, et al.
Pubblicazione: (2025)
di: Lu, Yiting, et al.
Pubblicazione: (2025)
Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models
di: Sharma, Pranav, et al.
Pubblicazione: (2025)
di: Sharma, Pranav, et al.
Pubblicazione: (2025)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
di: Ahmadi, Saba, et al.
Pubblicazione: (2023)
di: Ahmadi, Saba, et al.
Pubblicazione: (2023)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
di: Yang, Qian, et al.
Pubblicazione: (2024)
di: Yang, Qian, et al.
Pubblicazione: (2024)
Can We Predict Performance of Large Models across Vision-Language Tasks?
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
di: Sheta, Hala, et al.
Pubblicazione: (2025)
di: Sheta, Hala, et al.
Pubblicazione: (2025)
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
di: Dinh, Quang Minh, et al.
Pubblicazione: (2024)
di: Dinh, Quang Minh, et al.
Pubblicazione: (2024)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
di: Liu, Peng, et al.
Pubblicazione: (2025)
di: Liu, Peng, et al.
Pubblicazione: (2025)
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
di: Yang, Zhantao, et al.
Pubblicazione: (2024)
di: Yang, Zhantao, et al.
Pubblicazione: (2024)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
di: Chaffin, Antoine, et al.
Pubblicazione: (2024)
di: Chaffin, Antoine, et al.
Pubblicazione: (2024)
SegMASt3R: Geometry Grounded Segment Matching
di: Jayanti, Rohit, et al.
Pubblicazione: (2025)
di: Jayanti, Rohit, et al.
Pubblicazione: (2025)
Image-Caption Encoding for Improving Zero-Shot Generalization
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
di: Lokesh, K, et al.
Pubblicazione: (2026)
di: Lokesh, K, et al.
Pubblicazione: (2026)
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
di: Feinglass, Joshua, et al.
Pubblicazione: (2024)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
di: Li, Yuying, et al.
Pubblicazione: (2025)
di: Li, Yuying, et al.
Pubblicazione: (2025)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
di: Shi, Weikang, et al.
Pubblicazione: (2025)
di: Shi, Weikang, et al.
Pubblicazione: (2025)
From Image Captioning to Visual Storytelling
di: Passadakis, Admitos, et al.
Pubblicazione: (2025)
di: Passadakis, Admitos, et al.
Pubblicazione: (2025)
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
di: Yu, Yongcan, et al.
Pubblicazione: (2025)
di: Yu, Yongcan, et al.
Pubblicazione: (2025)
GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow Networks
di: Kang, Haoqiang, et al.
Pubblicazione: (2025)
di: Kang, Haoqiang, et al.
Pubblicazione: (2025)
MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
di: Goel, Arushi, et al.
Pubblicazione: (2026)
di: Goel, Arushi, et al.
Pubblicazione: (2026)
BabyVision: Visual Reasoning Beyond Language
di: Chen, Liang, et al.
Pubblicazione: (2026)
di: Chen, Liang, et al.
Pubblicazione: (2026)
From Pixels to Prose: A Large Dataset of Dense Image Captions
di: Singla, Vasu, et al.
Pubblicazione: (2024)
di: Singla, Vasu, et al.
Pubblicazione: (2024)
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
di: Zhang, Wenqi, et al.
Pubblicazione: (2025)
di: Zhang, Wenqi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Are VLMs Really Blind
di: Singh, Ayush, et al.
Pubblicazione: (2024) -
Give me a hint: Can LLMs take a hint to solve math problems?
di: Agrawal, Vansh, et al.
Pubblicazione: (2024) -
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
di: Zhang, Di, et al.
Pubblicazione: (2024) -
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
di: Siingh, Shikhhar, et al.
Pubblicazione: (2025) -
Re:Verse -- Can Your VLM Read a Manga?
di: Baranwal, Aaditya, et al.
Pubblicazione: (2025)