CapGeo: A Caption-Assisted Approach to Geometric Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Yuying, Qian, Siyi, Liang, Hao, Zheng, Leqi, An, Ruichuan, Guo, Yongzhen, Zhang, Wentao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ChartCap: Mitigating Hallucination of Dense Chart Captioning
por: Lim, Junyoung, et al.
Publicado: (2025)
por: Lim, Junyoung, et al.
Publicado: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
por: Xing, Long, et al.
Publicado: (2025)
por: Xing, Long, et al.
Publicado: (2025)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
por: Ng, Ho Yin 'Sam', et al.
Publicado: (2025)
por: Ng, Ho Yin 'Sam', et al.
Publicado: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
por: Liu, Zheng, et al.
Publicado: (2024)
por: Liu, Zheng, et al.
Publicado: (2024)
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
por: Hsu, Ting-Yao E., et al.
Publicado: (2025)
por: Hsu, Ting-Yao E., et al.
Publicado: (2025)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
por: Cheng, Kanzhi, et al.
Publicado: (2025)
por: Cheng, Kanzhi, et al.
Publicado: (2025)
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
por: Huang, Ting-Hao 'Kenneth', et al.
Publicado: (2025)
por: Huang, Ting-Hao 'Kenneth', et al.
Publicado: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
por: Guo, Ziyu, et al.
Publicado: (2025)
por: Guo, Ziyu, et al.
Publicado: (2025)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
por: Singh, Ayush, et al.
Publicado: (2024)
por: Singh, Ayush, et al.
Publicado: (2024)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
por: Xing, Long, et al.
Publicado: (2025)
por: Xing, Long, et al.
Publicado: (2025)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
por: Matsuda, Kazuki, et al.
Publicado: (2025)
por: Matsuda, Kazuki, et al.
Publicado: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
por: Du, Yang, et al.
Publicado: (2025)
por: Du, Yang, et al.
Publicado: (2025)
Imagine How To Change: Explicit Procedure Modeling for Change Captioning
por: Sun, Jiayang, et al.
Publicado: (2026)
por: Sun, Jiayang, et al.
Publicado: (2026)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
por: Tu, Yunbin, et al.
Publicado: (2024)
por: Tu, Yunbin, et al.
Publicado: (2024)
Agri-CPJ: A Training-Free Explainable Framework for Agricultural Pest Diagnosis Using Caption-Prompt-Judge and LLM-as-a-Judge
por: Zhang, Wentao, et al.
Publicado: (2026)
por: Zhang, Wentao, et al.
Publicado: (2026)
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
por: Sukhani, Siddhant, et al.
Publicado: (2025)
por: Sukhani, Siddhant, et al.
Publicado: (2025)
Unveiling the Invisible: Captioning Videos with Metaphors
por: Kalarani, Abisek Rajakumar, et al.
Publicado: (2024)
por: Kalarani, Abisek Rajakumar, et al.
Publicado: (2024)
Text-only Synthesis for Image Captioning
por: Zhou, Qing, et al.
Publicado: (2024)
por: Zhou, Qing, et al.
Publicado: (2024)
The Role of Data Curation in Image Captioning
por: Li, Wenyan, et al.
Publicado: (2023)
por: Li, Wenyan, et al.
Publicado: (2023)
WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
por: Chen, Pingyi, et al.
Publicado: (2023)
por: Chen, Pingyi, et al.
Publicado: (2023)
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
por: Ye, Shaokai, et al.
Publicado: (2026)
por: Ye, Shaokai, et al.
Publicado: (2026)
Updating CLIP to Prefer Descriptions Over Captions
por: Zur, Amir, et al.
Publicado: (2024)
por: Zur, Amir, et al.
Publicado: (2024)
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
por: Cai, Shihao, et al.
Publicado: (2024)
por: Cai, Shihao, et al.
Publicado: (2024)
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
por: Liu, Zhaochen, et al.
Publicado: (2026)
por: Liu, Zhaochen, et al.
Publicado: (2026)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
por: Bajpai, Divya Jyoti, et al.
Publicado: (2024)
por: Bajpai, Divya Jyoti, et al.
Publicado: (2024)
CIC: A Framework for Culturally-Aware Image Captioning
por: Yun, Youngsik, et al.
Publicado: (2024)
por: Yun, Youngsik, et al.
Publicado: (2024)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
por: Chen, Yuyan, et al.
Publicado: (2024)
por: Chen, Yuyan, et al.
Publicado: (2024)
Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning
por: Shen, Zijun, et al.
Publicado: (2026)
por: Shen, Zijun, et al.
Publicado: (2026)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
por: Kim, Hyunjong, et al.
Publicado: (2025)
por: Kim, Hyunjong, et al.
Publicado: (2025)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
por: Lu, Xingyu, et al.
Publicado: (2026)
por: Lu, Xingyu, et al.
Publicado: (2026)
Beyond Symbolic Solving: Multi Chain-of-Thought Voting for Geometric Reasoning in Large Language Models
por: Siddique, Md. Abu Bakor, et al.
Publicado: (2026)
por: Siddique, Md. Abu Bakor, et al.
Publicado: (2026)
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback
por: Singh, Ashish, et al.
Publicado: (2023)
por: Singh, Ashish, et al.
Publicado: (2023)
Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs
por: An, Ruichuan, et al.
Publicado: (2025)
por: An, Ruichuan, et al.
Publicado: (2025)
Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
por: Liang, Hao, et al.
Publicado: (2025)
por: Liang, Hao, et al.
Publicado: (2025)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
por: Zhang, Wentao, et al.
Publicado: (2025)
por: Zhang, Wentao, et al.
Publicado: (2025)
Unbiased Visual Reasoning with Controlled Visual Inputs
por: Li, Zhaonan, et al.
Publicado: (2025)
por: Li, Zhaonan, et al.
Publicado: (2025)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
por: Möller, Lucas, et al.
Publicado: (2024)
por: Möller, Lucas, et al.
Publicado: (2024)
MAMI: Multi-Attentional Mutual-Information for Long Sequence Neuron Captioning
por: Fauzulhaq, Alfirsa Damasyifa, et al.
Publicado: (2024)
por: Fauzulhaq, Alfirsa Damasyifa, et al.
Publicado: (2024)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
por: Sarto, Sara, et al.
Publicado: (2025)
por: Sarto, Sara, et al.
Publicado: (2025)
Ejemplares similares
-
ChartCap: Mitigating Hallucination of Dense Chart Captioning
por: Lim, Junyoung, et al.
Publicado: (2025) -
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
por: Xing, Long, et al.
Publicado: (2025) -
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
por: Ng, Ho Yin 'Sam', et al.
Publicado: (2025) -
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
por: Liu, Zheng, et al.
Publicado: (2024) -
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
por: Hsu, Ting-Yao E., et al.
Publicado: (2025)