Enregistré dans:
Détails bibliographiques
Auteurs principaux: Reddy, N Dinesh, Snyder, Dylan, Kiragu, Lona, Mohin, Mirajul, Amin, Shahrear Bin, Pillai, Sudeep
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2511.14210
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914164754087936
author Reddy, N Dinesh
Snyder, Dylan
Kiragu, Lona
Mohin, Mirajul
Amin, Shahrear Bin
Pillai, Sudeep
author_facet Reddy, N Dinesh
Snyder, Dylan
Kiragu, Lona
Mohin, Mirajul
Amin, Shahrear Bin
Pillai, Sudeep
contents We introduce Orion, a visual agent that integrates vision-based reasoning with tool-augmented execution to achieve powerful, precise, multi-step visual intelligence across images, video, and documents. Unlike traditional vision-language models that generate descriptive outputs, Orion orchestrates a suite of specialized computer vision tools, including object detection, keypoint localization, panoptic segmentation, Optical Character Recognition (OCR), and geometric analysis, to execute complex multi-step visual workflows. The system achieves competitive performance across MMMU, MMBench, DocVQA, and MMLongBench while extending monolithic VLM capabilities to production-grade visual intelligence. Through its agentic, tool-augmented approach, Orion enables autonomous visual reasoning that bridges neural perception with symbolic execution, marking the transition from passive visual understanding to active, tool-driven visual intelligence. Try Orion for free at: https://chat.vlm.run Learn more at: https://www.vlm.run/orion
format Preprint
id arxiv_https___arxiv_org_abs_2511_14210
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
Reddy, N Dinesh
Snyder, Dylan
Kiragu, Lona
Mohin, Mirajul
Amin, Shahrear Bin
Pillai, Sudeep
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We introduce Orion, a visual agent that integrates vision-based reasoning with tool-augmented execution to achieve powerful, precise, multi-step visual intelligence across images, video, and documents. Unlike traditional vision-language models that generate descriptive outputs, Orion orchestrates a suite of specialized computer vision tools, including object detection, keypoint localization, panoptic segmentation, Optical Character Recognition (OCR), and geometric analysis, to execute complex multi-step visual workflows. The system achieves competitive performance across MMMU, MMBench, DocVQA, and MMLongBench while extending monolithic VLM capabilities to production-grade visual intelligence. Through its agentic, tool-augmented approach, Orion enables autonomous visual reasoning that bridges neural perception with symbolic execution, marking the transition from passive visual understanding to active, tool-driven visual intelligence. Try Orion for free at: https://chat.vlm.run Learn more at: https://www.vlm.run/orion
title Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.14210