Artemis: Structured Visual Reasoning for Perception Policy Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Wei, Sun, Yanpeng, Zhang, Shan, Bo, Weihao, Li, Xiaofan, Koniusz, Piotr, Li, Wei, Zhao, Na, Li, Zechao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
by: Bo, Weihao, et al.
Published: (2025)
by: Bo, Weihao, et al.
Published: (2025)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Exploring Effective Factors for Improving Visual In-Context Learning
by: Sun, Yanpeng, et al.
Published: (2023)
by: Sun, Yanpeng, et al.
Published: (2023)
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
by: Tang, Wei, et al.
Published: (2026)
by: Tang, Wei, et al.
Published: (2026)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
by: Sun, Yanpeng, et al.
Published: (2025)
by: Sun, Yanpeng, et al.
Published: (2025)
VRP-SAM: SAM with Visual Reference Prompt
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
FVAR: Visual Autoregressive Modeling via Next Focus Prediction
by: Li, Xiaofan, et al.
Published: (2025)
by: Li, Xiaofan, et al.
Published: (2025)
PACE: Marrying generalization in PArameter-efficient fine-tuning with Consistency rEgularization
by: Ni, Yao, et al.
Published: (2024)
by: Ni, Yao, et al.
Published: (2024)
Video4Edit: Viewing Image Editing as a Degenerate Temporal Process
by: Li, Xiaofan, et al.
Published: (2025)
by: Li, Xiaofan, et al.
Published: (2025)
Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting
by: Bo, Yuntian, et al.
Published: (2026)
by: Bo, Yuntian, et al.
Published: (2026)
Uncertainty-DTW for Sequences and Visual Tokens
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Learning Gaussian Representation for Eye Fixation Prediction
by: Song, Peipei, et al.
Published: (2024)
by: Song, Peipei, et al.
Published: (2024)
CHAIN: Enhancing Generalization in Data-Efficient GANs via lipsCHitz continuity constrAIned Normalization
by: Ni, Yao, et al.
Published: (2024)
by: Ni, Yao, et al.
Published: (2024)
Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited Samples
by: Fang, Ziye, et al.
Published: (2023)
by: Fang, Ziye, et al.
Published: (2023)
OpenKD: Opening Prompt Diversity for Zero- and Few-shot Keypoint Detection
by: Lu, Changsheng, et al.
Published: (2024)
by: Lu, Changsheng, et al.
Published: (2024)
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
by: Liu, Junming, et al.
Published: (2026)
by: Liu, Junming, et al.
Published: (2026)
Feature Hallucination for Self-supervised Action Recognition
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Artemis: Articulated Neural Pets with Appearance and Motion synthesis
by: Luo, Haimin, et al.
Published: (2022)
by: Luo, Haimin, et al.
Published: (2022)
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
by: Xuan, Shiyu, et al.
Published: (2025)
by: Xuan, Shiyu, et al.
Published: (2025)
A Comprehensive Survey on Visual Concept Mining in Text-to-image Diffusion Models
by: Li, Ziqiang, et al.
Published: (2025)
by: Li, Ziqiang, et al.
Published: (2025)
Artemis: Towards Referential Understanding in Complex Videos
by: Qiu, Jihao, et al.
Published: (2024)
by: Qiu, Jihao, et al.
Published: (2024)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
by: Zhao, Shijie, et al.
Published: (2025)
by: Zhao, Shijie, et al.
Published: (2025)
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
by: Li, Zejun, et al.
Published: (2025)
by: Li, Zejun, et al.
Published: (2025)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
by: Xue, Wei, et al.
Published: (2026)
by: Xue, Wei, et al.
Published: (2026)
CSGO: Content-Style Composition in Text-to-Image Generation
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
URA-Net: Uncertainty-Integrated Anomaly Perception and Restoration Attention Network for Unsupervised Anomaly Detection
by: Luo, Wei, et al.
Published: (2026)
by: Luo, Wei, et al.
Published: (2026)
Perception-R1: Pioneering Perception Policy with Reinforcement Learning
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
Learning Time in Static Classifiers
by: Ding, Xi, et al.
Published: (2025)
by: Ding, Xi, et al.
Published: (2025)
Subspace Kernel Learning on Tensor Sequences
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Adaptive Multi-head Contrastive Learning
by: Wang, Lei, et al.
Published: (2023)
by: Wang, Lei, et al.
Published: (2023)
Spatial Structure Constraints for Weakly Supervised Semantic Segmentation
by: Chen, Tao, et al.
Published: (2024)
by: Chen, Tao, et al.
Published: (2024)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
by: Bai, Sule, et al.
Published: (2025)
by: Bai, Sule, et al.
Published: (2025)
Video Understanding by Design: How Datasets Shape Architectures and Insights
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
See the Text: From Tokenization to Visual Reading
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
by: Xuan, Shiyu, et al.
Published: (2026)
by: Xuan, Shiyu, et al.
Published: (2026)
Contrastive Graph Modeling for Cross-Domain Few-Shot Medical Image Segmentation
by: Bo, Yuntian, et al.
Published: (2025)
by: Bo, Yuntian, et al.
Published: (2025)
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Similar Items
-
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
by: Bo, Weihao, et al.
Published: (2025) -
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
by: Sun, Yanpeng, et al.
Published: (2024) -
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025) -
Exploring Effective Factors for Improving Visual In-Context Learning
by: Sun, Yanpeng, et al.
Published: (2023) -
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
by: Tang, Wei, et al.
Published: (2026)