Active Visual Perception: Opportunities and Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yian, Guo, Xiaoyu, Zhang, Hao, Li, Shuiwang, Dai, Xiaowei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Low-Light Object Tracking: A Benchmark
by: Zhong, Pengzhi, et al.
Published: (2024)
by: Zhong, Pengzhi, et al.
Published: (2024)
Vision-Based Anti Unmanned Aerial Technology: Opportunities and Challenges
by: Ding, Guanghai, et al.
Published: (2025)
by: Ding, Guanghai, et al.
Published: (2025)
Tracking Reflected Objects: A Benchmark
by: Guo, Xiaoyu, et al.
Published: (2024)
by: Guo, Xiaoyu, et al.
Published: (2024)
Camouflaged Object Tracking: A Benchmark
by: Guo, Xiaoyu, et al.
Published: (2024)
by: Guo, Xiaoyu, et al.
Published: (2024)
Silver medal Solution for Image Matching Challenge 2024
by: Wang, Yian
Published: (2024)
by: Wang, Yian
Published: (2024)
Act2See: Emergent Active Visual Perception for Video Reasoning
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
Forging Vision Foundation Models for Autonomous Driving: Challenges, Methodologies, and Opportunities
by: Yan, Xu, et al.
Published: (2024)
by: Yan, Xu, et al.
Published: (2024)
Adaptively Bypassing Vision Transformer Blocks for Efficient Visual Tracking
by: Yang, Xiangyang, et al.
Published: (2024)
by: Yang, Xiangyang, et al.
Published: (2024)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
ActFormer: Scalable Collaborative Perception via Active Queries
by: Huang, Suozhi, et al.
Published: (2024)
by: Huang, Suozhi, et al.
Published: (2024)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models
by: Yang, Xiaoyu, et al.
Published: (2023)
by: Yang, Xiaoyu, et al.
Published: (2023)
Visual Bridge: Universal Visual Perception Representations Generating
by: Gao, Yilin, et al.
Published: (2025)
by: Gao, Yilin, et al.
Published: (2025)
Artemis: Structured Visual Reasoning for Perception Policy Learning
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
by: Bao, Yanqi, et al.
Published: (2024)
by: Bao, Yanqi, et al.
Published: (2024)
Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
by: Luo, Zhengyi, et al.
Published: (2025)
by: Luo, Zhengyi, et al.
Published: (2025)
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
by: Zhu, Jiashun, et al.
Published: (2026)
by: Zhu, Jiashun, et al.
Published: (2026)
Collaborative Perception for Connected and Autonomous Driving: Challenges, Possible Solutions and Opportunities
by: Hu, Senkang, et al.
Published: (2024)
by: Hu, Senkang, et al.
Published: (2024)
SonoSelect: Efficient Ultrasound Perception via Active Probe Exploration
by: Zhang, Yixin, et al.
Published: (2026)
by: Zhang, Yixin, et al.
Published: (2026)
SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
by: Zheng, Hongpei, et al.
Published: (2025)
by: Zheng, Hongpei, et al.
Published: (2025)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
by: Wang, Zhaokai, et al.
Published: (2025)
by: Wang, Zhaokai, et al.
Published: (2025)
The Shape of Sight: A Homological Framework for Unifying Visual Perception
by: Li, Xin
Published: (2018)
by: Li, Xin
Published: (2018)
Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
by: Tang, Lv, et al.
Published: (2023)
by: Tang, Lv, et al.
Published: (2023)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
by: Li, Yian, et al.
Published: (2026)
by: Li, Yian, et al.
Published: (2026)
Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking
by: Zhong, Pengzhi, et al.
Published: (2026)
by: Zhong, Pengzhi, et al.
Published: (2026)
Breaking the Passive Learning Trap: An Active Perception Strategy for Human Motion Prediction
by: Hu, Juncheng, et al.
Published: (2025)
by: Hu, Juncheng, et al.
Published: (2025)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
by: Duan, Yuchen, et al.
Published: (2024)
by: Duan, Yuchen, et al.
Published: (2024)
GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
by: Liu, Ruixun, et al.
Published: (2025)
by: Liu, Ruixun, et al.
Published: (2025)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
by: Dai, Haocheng, et al.
Published: (2024)
by: Dai, Haocheng, et al.
Published: (2024)
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
by: Sun, Xiaoxiao, et al.
Published: (2026)
by: Sun, Xiaoxiao, et al.
Published: (2026)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
by: Guan, Tongkun, et al.
Published: (2026)
by: Guan, Tongkun, et al.
Published: (2026)
Vision Language Models for Spreadsheet Understanding: Challenges and Opportunities
by: Xia, Shiyu, et al.
Published: (2024)
by: Xia, Shiyu, et al.
Published: (2024)
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
by: Shi, Dai
Published: (2023)
by: Shi, Dai
Published: (2023)
The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning
by: Li, Yaohui, et al.
Published: (2024)
by: Li, Yaohui, et al.
Published: (2024)
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
by: Wang, Ziyue, et al.
Published: (2024)
by: Wang, Ziyue, et al.
Published: (2024)
Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving
by: Zhang, Bozhou, et al.
Published: (2025)
by: Zhang, Bozhou, et al.
Published: (2025)
Instance-level Visual Active Tracking with Occlusion-Aware Planning
by: Sun, Haowei, et al.
Published: (2026)
by: Sun, Haowei, et al.
Published: (2026)
Region-Adaptive Video Sharpening via Rate-Perception Optimization
by: Pang, Yingxue, et al.
Published: (2025)
by: Pang, Yingxue, et al.
Published: (2025)
Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach
by: Bi, Jing, et al.
Published: (2024)
by: Bi, Jing, et al.
Published: (2024)
Similar Items
-
Low-Light Object Tracking: A Benchmark
by: Zhong, Pengzhi, et al.
Published: (2024) -
Vision-Based Anti Unmanned Aerial Technology: Opportunities and Challenges
by: Ding, Guanghai, et al.
Published: (2025) -
Tracking Reflected Objects: A Benchmark
by: Guo, Xiaoyu, et al.
Published: (2024) -
Camouflaged Object Tracking: A Benchmark
by: Guo, Xiaoyu, et al.
Published: (2024) -
Silver medal Solution for Image Matching Challenge 2024
by: Wang, Yian
Published: (2024)