Reinforced Visual Perception with Tools
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Zetong, Chen, Dongping, Ma, Zixian, Hu, Zhihan, Fu, Mingyang, Wang, Sinan, Wan, Yao, Zhao, Zhou, Krishna, Ranjay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Seeking and Updating with Live Visual Knowledge
di: Fu, Mingyang, et al.
Pubblicazione: (2025)
di: Fu, Mingyang, et al.
Pubblicazione: (2025)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
di: Ma, Zixian, et al.
Pubblicazione: (2024)
di: Ma, Zixian, et al.
Pubblicazione: (2024)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
MultiRef: Controllable Image Generation with Multiple Visual References
di: Chen, Ruoxi, et al.
Pubblicazione: (2025)
di: Chen, Ruoxi, et al.
Pubblicazione: (2025)
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
di: Chen, Dongping, et al.
Pubblicazione: (2026)
di: Chen, Dongping, et al.
Pubblicazione: (2026)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
di: Song, Mingyang, et al.
Pubblicazione: (2026)
di: Song, Mingyang, et al.
Pubblicazione: (2026)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
di: Hu, Yushi, et al.
Pubblicazione: (2023)
di: Hu, Yushi, et al.
Pubblicazione: (2023)
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
di: Hu, Yushi, et al.
Pubblicazione: (2024)
di: Hu, Yushi, et al.
Pubblicazione: (2024)
Judge Anything: MLLM as a Judge Across Any Modality
di: Pu, Shu, et al.
Pubblicazione: (2025)
di: Pu, Shu, et al.
Pubblicazione: (2025)
Mitigating Object Hallucination via Robust Local Perception Search
di: Gao, Zixian, et al.
Pubblicazione: (2025)
di: Gao, Zixian, et al.
Pubblicazione: (2025)
Visual Representations inside the Language Model
di: Liu, Benlin, et al.
Pubblicazione: (2025)
di: Liu, Benlin, et al.
Pubblicazione: (2025)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
di: Li, Baiqi, et al.
Pubblicazione: (2024)
di: Li, Baiqi, et al.
Pubblicazione: (2024)
Paper2Web: Let's Make Your Paper Alive!
di: Chen, Yuhang, et al.
Pubblicazione: (2025)
di: Chen, Yuhang, et al.
Pubblicazione: (2025)
The Hard Positive Truth about Vision-Language Compositionality
di: Kamath, Amita, et al.
Pubblicazione: (2024)
di: Kamath, Amita, et al.
Pubblicazione: (2024)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
di: Li, Linjie, et al.
Pubblicazione: (2025)
di: Li, Linjie, et al.
Pubblicazione: (2025)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
di: Bigverdi, Mahtab, et al.
Pubblicazione: (2024)
di: Bigverdi, Mahtab, et al.
Pubblicazione: (2024)
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
di: Li, Shaoxuan, et al.
Pubblicazione: (2026)
di: Li, Shaoxuan, et al.
Pubblicazione: (2026)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
di: Yang, Yinuo, et al.
Pubblicazione: (2026)
di: Yang, Yinuo, et al.
Pubblicazione: (2026)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
di: Wu, Daiqing, et al.
Pubblicazione: (2025)
Perception-R1: Pioneering Perception Policy with Reinforcement Learning
di: Yu, En, et al.
Pubblicazione: (2025)
di: Yu, En, et al.
Pubblicazione: (2025)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
di: Zhou, Yucheng, et al.
Pubblicazione: (2024)
di: Zhou, Yucheng, et al.
Pubblicazione: (2024)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
di: Kamath, Amita, et al.
Pubblicazione: (2026)
di: Kamath, Amita, et al.
Pubblicazione: (2026)
Video-Based Reward Modeling for Computer-Use Agents
di: Song, Linxin, et al.
Pubblicazione: (2026)
di: Song, Linxin, et al.
Pubblicazione: (2026)
Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
di: Shen, Wenxuan, et al.
Pubblicazione: (2025)
di: Shen, Wenxuan, et al.
Pubblicazione: (2025)
REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations
di: Sushko, Peter, et al.
Pubblicazione: (2025)
di: Sushko, Peter, et al.
Pubblicazione: (2025)
One RL to See Them All: Visual Triple Unified Reinforcement Learning
di: Ma, Yan, et al.
Pubblicazione: (2025)
di: Ma, Yan, et al.
Pubblicazione: (2025)
BLINK: Multimodal Large Language Models Can See but Not Perceive
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
di: Fu, Xingyu, et al.
Pubblicazione: (2024)
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models
di: Shao, Zhenwei, et al.
Pubblicazione: (2025)
di: Shao, Zhenwei, et al.
Pubblicazione: (2025)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
Learning to Instruct for Visual Instruction Tuning
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
Visual Confused Deputy: Exploiting and Defending Perception Failures in Computer-Using Agents
di: Liu, Xunzhuo, et al.
Pubblicazione: (2026)
di: Liu, Xunzhuo, et al.
Pubblicazione: (2026)
Synthetic Visual Genome
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
di: Jiang, Dongfu, et al.
Pubblicazione: (2025)
di: Jiang, Dongfu, et al.
Pubblicazione: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
di: Du, Yifan, et al.
Pubblicazione: (2023)
di: Du, Yifan, et al.
Pubblicazione: (2023)
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Seeking and Updating with Live Visual Knowledge
di: Fu, Mingyang, et al.
Pubblicazione: (2025) -
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
di: Ma, Zixian, et al.
Pubblicazione: (2024) -
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
di: Chen, Dongping, et al.
Pubblicazione: (2024) -
MultiRef: Controllable Image Generation with Multiple Visual References
di: Chen, Ruoxi, et al.
Pubblicazione: (2025) -
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
di: Chen, Dongping, et al.
Pubblicazione: (2026)