Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2511.14210 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866914164754087936 |
|---|---|
| author | Reddy, N Dinesh Snyder, Dylan Kiragu, Lona Mohin, Mirajul Amin, Shahrear Bin Pillai, Sudeep |
| author_facet | Reddy, N Dinesh Snyder, Dylan Kiragu, Lona Mohin, Mirajul Amin, Shahrear Bin Pillai, Sudeep |
| contents | We introduce Orion, a visual agent that integrates vision-based reasoning with tool-augmented execution to achieve powerful, precise, multi-step visual intelligence across images, video, and documents. Unlike traditional vision-language models that generate descriptive outputs, Orion orchestrates a suite of specialized computer vision tools, including object detection, keypoint localization, panoptic segmentation, Optical Character Recognition (OCR), and geometric analysis, to execute complex multi-step visual workflows. The system achieves competitive performance across MMMU, MMBench, DocVQA, and MMLongBench while extending monolithic VLM capabilities to production-grade visual intelligence. Through its agentic, tool-augmented approach, Orion enables autonomous visual reasoning that bridges neural perception with symbolic execution, marking the transition from passive visual understanding to active, tool-driven visual intelligence.
Try Orion for free at: https://chat.vlm.run
Learn more at: https://www.vlm.run/orion |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_14210 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution Reddy, N Dinesh Snyder, Dylan Kiragu, Lona Mohin, Mirajul Amin, Shahrear Bin Pillai, Sudeep Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning We introduce Orion, a visual agent that integrates vision-based reasoning with tool-augmented execution to achieve powerful, precise, multi-step visual intelligence across images, video, and documents. Unlike traditional vision-language models that generate descriptive outputs, Orion orchestrates a suite of specialized computer vision tools, including object detection, keypoint localization, panoptic segmentation, Optical Character Recognition (OCR), and geometric analysis, to execute complex multi-step visual workflows. The system achieves competitive performance across MMMU, MMBench, DocVQA, and MMLongBench while extending monolithic VLM capabilities to production-grade visual intelligence. Through its agentic, tool-augmented approach, Orion enables autonomous visual reasoning that bridges neural perception with symbolic execution, marking the transition from passive visual understanding to active, tool-driven visual intelligence. Try Orion for free at: https://chat.vlm.run Learn more at: https://www.vlm.run/orion |
| title | Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2511.14210 |