VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Peng, Shen, Haozhan, Fang, Chunxin, Sun, Zhicheng, Liao, Jiajia, Zhao, Tiancheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
di: Shen, Haozhan, et al.
Pubblicazione: (2025)
di: Shen, Haozhan, et al.
Pubblicazione: (2025)
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research
di: Zhang, Qianqian, et al.
Pubblicazione: (2025)
di: Zhang, Qianqian, et al.
Pubblicazione: (2025)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
di: Shen, Yifan, et al.
Pubblicazione: (2025)
di: Shen, Yifan, et al.
Pubblicazione: (2025)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
di: Zhao, Tiancheng, et al.
Pubblicazione: (2024)
di: Zhao, Tiancheng, et al.
Pubblicazione: (2024)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
di: Bandraupalli, Srihari, et al.
Pubblicazione: (2025)
di: Bandraupalli, Srihari, et al.
Pubblicazione: (2025)
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
di: Chen, Guizhen, et al.
Pubblicazione: (2025)
di: Chen, Guizhen, et al.
Pubblicazione: (2025)
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
di: Carvalho, Miguel, et al.
Pubblicazione: (2025)
di: Carvalho, Miguel, et al.
Pubblicazione: (2025)
Talking to Yourself: Defying Forgetting in Large Language Models
di: Sun, Yutao, et al.
Pubblicazione: (2026)
di: Sun, Yutao, et al.
Pubblicazione: (2026)
Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs
di: Zhu, Mengdan, et al.
Pubblicazione: (2026)
di: Zhu, Mengdan, et al.
Pubblicazione: (2026)
ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents
di: Zhang, Xiaohui, et al.
Pubblicazione: (2026)
di: Zhang, Xiaohui, et al.
Pubblicazione: (2026)
Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation
di: Sun, Yirong, et al.
Pubblicazione: (2024)
di: Sun, Yirong, et al.
Pubblicazione: (2024)
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
di: Wang, Yeyuan, et al.
Pubblicazione: (2025)
di: Wang, Yeyuan, et al.
Pubblicazione: (2025)
Awakening LLMs' Reasoning Potential: A Fine-Grained Pipeline to Evaluate and Mitigate Vague Perception
di: Ling, Zipeng, et al.
Pubblicazione: (2025)
di: Ling, Zipeng, et al.
Pubblicazione: (2025)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
di: Avogaro, Niccolo, et al.
Pubblicazione: (2026)
di: Avogaro, Niccolo, et al.
Pubblicazione: (2026)
On the Perception Bottleneck of VLMs for Chart Understanding
di: Liu, Junteng, et al.
Pubblicazione: (2025)
di: Liu, Junteng, et al.
Pubblicazione: (2025)
S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
di: Sun, Shaoning, et al.
Pubblicazione: (2025)
di: Sun, Shaoning, et al.
Pubblicazione: (2025)
Bridging the Arithmetic Gap: The Cognitive Complexity Benchmark and Financial-PoT for Robust Financial Reasoning
di: Zhao, Boxiang, et al.
Pubblicazione: (2026)
di: Zhao, Boxiang, et al.
Pubblicazione: (2026)
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
di: Shen, Haozhan, et al.
Pubblicazione: (2026)
di: Shen, Haozhan, et al.
Pubblicazione: (2026)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
di: Wang, Shengao, et al.
Pubblicazione: (2025)
di: Wang, Shengao, et al.
Pubblicazione: (2025)
Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning
di: Yang, Zhaorui, et al.
Pubblicazione: (2024)
di: Yang, Zhaorui, et al.
Pubblicazione: (2024)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
di: Zhang, Di, et al.
Pubblicazione: (2024)
di: Zhang, Di, et al.
Pubblicazione: (2024)
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
di: Yu, Yongcan, et al.
Pubblicazione: (2025)
di: Yu, Yongcan, et al.
Pubblicazione: (2025)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025)
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025)
Verifiable Generation with Subsentence-Level Fine-Grained Citations
di: Cao, Shuyang, et al.
Pubblicazione: (2024)
di: Cao, Shuyang, et al.
Pubblicazione: (2024)
Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
di: Dong, Lingzhong, et al.
Pubblicazione: (2025)
di: Dong, Lingzhong, et al.
Pubblicazione: (2025)
Bridging the Editing Gap in LLMs: FineEdit for Precise and Targeted Text Modifications
di: Zeng, Yiming, et al.
Pubblicazione: (2025)
di: Zeng, Yiming, et al.
Pubblicazione: (2025)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
di: Wang, Ante, et al.
Pubblicazione: (2024)
di: Wang, Ante, et al.
Pubblicazione: (2024)
Medchain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence
di: Liu, Jie, et al.
Pubblicazione: (2024)
di: Liu, Jie, et al.
Pubblicazione: (2024)
ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews
di: Gao, Xian, et al.
Pubblicazione: (2025)
di: Gao, Xian, et al.
Pubblicazione: (2025)
Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
di: Jia, Ruipeng, et al.
Pubblicazione: (2025)
di: Jia, Ruipeng, et al.
Pubblicazione: (2025)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
di: He, Lehan, et al.
Pubblicazione: (2024)
di: He, Lehan, et al.
Pubblicazione: (2024)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
di: Zhao, Tiancheng, et al.
Pubblicazione: (2022)
di: Zhao, Tiancheng, et al.
Pubblicazione: (2022)
Understanding and Rectifying Safety Perception Distortion in VLMs
di: Zou, Xiaohan, et al.
Pubblicazione: (2025)
di: Zou, Xiaohan, et al.
Pubblicazione: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
di: Zhao, Zhixian, et al.
Pubblicazione: (2026)
di: Zhao, Zhixian, et al.
Pubblicazione: (2026)
[De|Re]constructing VLMs' Reasoning in Counting
di: Alghisi, Simone, et al.
Pubblicazione: (2025)
di: Alghisi, Simone, et al.
Pubblicazione: (2025)
Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
di: Xia, Runze, et al.
Pubblicazione: (2025)
di: Xia, Runze, et al.
Pubblicazione: (2025)
VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
di: Zhang, Jianshu, et al.
Pubblicazione: (2025)
di: Zhang, Jianshu, et al.
Pubblicazione: (2025)
ReFT: Reasoning with Reinforced Fine-Tuning
di: Luong, Trung Quoc, et al.
Pubblicazione: (2024)
di: Luong, Trung Quoc, et al.
Pubblicazione: (2024)
Sentiment Analysis Dataset in Moroccan Dialect: Bridging the Gap Between Arabic and Latin Scripted dialect
di: Jbel, Mouad, et al.
Pubblicazione: (2023)
di: Jbel, Mouad, et al.
Pubblicazione: (2023)
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
di: Larionov, Daniil, et al.
Pubblicazione: (2024)
di: Larionov, Daniil, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
di: Shen, Haozhan, et al.
Pubblicazione: (2025) -
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research
di: Zhang, Qianqian, et al.
Pubblicazione: (2025) -
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
di: Shen, Yifan, et al.
Pubblicazione: (2025) -
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
di: Zhao, Tiancheng, et al.
Pubblicazione: (2024) -
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
di: Bandraupalli, Srihari, et al.
Pubblicazione: (2025)