Agentic Discovery with Active Hypothesis Exploration for Visual Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Koo, Jaywon, Hernandez, Jefferson, He, Ruozhen, Chen, Hanjie, Wei, Chen, Ordonez, Vicente |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Referring Expressions: Scenario Comprehension Visual Grounding
by: He, Ruozhen, et al.
Published: (2026)
by: He, Ruozhen, et al.
Published: (2026)
GViT: Representing Images as Gaussians for Visual Recognition
by: Hernandez, Jefferson, et al.
Published: (2025)
by: Hernandez, Jefferson, et al.
Published: (2025)
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
by: Koo, Jaywon, et al.
Published: (2025)
by: Koo, Jaywon, et al.
Published: (2025)
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
by: Xiao, Zilin, et al.
Published: (2025)
by: Xiao, Zilin, et al.
Published: (2025)
PropTest: Automatic Property Testing for Improved Visual Programming
by: Koo, Jaywon, et al.
Published: (2024)
by: Koo, Jaywon, et al.
Published: (2024)
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
by: He, Ruozhen, et al.
Published: (2026)
by: He, Ruozhen, et al.
Published: (2026)
Generative Visual Instruction Tuning
by: Hernandez, Jefferson, et al.
Published: (2024)
by: Hernandez, Jefferson, et al.
Published: (2024)
Learning from Synthetic Data for Visual Grounding
by: He, Ruozhen, et al.
Published: (2024)
by: He, Ruozhen, et al.
Published: (2024)
NoiseShift: Resolution-Aware Noise Recalibration for Better Low-Resolution Image Generation
by: He, Ruozhen, et al.
Published: (2025)
by: He, Ruozhen, et al.
Published: (2025)
ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders
by: Hernandez, Jefferson, et al.
Published: (2023)
by: Hernandez, Jefferson, et al.
Published: (2023)
ActFER: Agentic Facial Expression Recognition via Active Tool-Augmented Visual Reasoning
by: Liu, Shifeng, et al.
Published: (2026)
by: Liu, Shifeng, et al.
Published: (2026)
Fairness and Bias Mitigation in Computer Vision: A Survey
by: Dehdashtian, Sepehr, et al.
Published: (2024)
by: Dehdashtian, Sepehr, et al.
Published: (2024)
DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning
by: Zou, Junbo, et al.
Published: (2025)
by: Zou, Junbo, et al.
Published: (2025)
Improving Large Vision and Language Models by Learning from a Panel of Peers
by: Hernandez, Jefferson, et al.
Published: (2025)
by: Hernandez, Jefferson, et al.
Published: (2025)
Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation
by: Chen, Peixin, et al.
Published: (2026)
by: Chen, Peixin, et al.
Published: (2026)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition
by: Chen, Shunpeng, et al.
Published: (2025)
by: Chen, Shunpeng, et al.
Published: (2025)
Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
by: Koo, Juil, et al.
Published: (2025)
by: Koo, Juil, et al.
Published: (2025)
Visual Prompt Discovery via Semantic Exploration
by: Kim, Jaechang, et al.
Published: (2026)
by: Kim, Jaechang, et al.
Published: (2026)
AdaGlimpse: Active Visual Exploration with Arbitrary Glimpse Position and Scale
by: Pardyl, Adam, et al.
Published: (2024)
by: Pardyl, Adam, et al.
Published: (2024)
Landmark Guided Visual Feature Extractor for Visual Speech Recognition with Limited Resource
by: Yang, Lei, et al.
Published: (2025)
by: Yang, Lei, et al.
Published: (2025)
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
by: Somayazulu, Arjun, et al.
Published: (2024)
by: Somayazulu, Arjun, et al.
Published: (2024)
T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition
by: Yeh, Chen, et al.
Published: (2024)
by: Yeh, Chen, et al.
Published: (2024)
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
by: Mi, Li, et al.
Published: (2025)
by: Mi, Li, et al.
Published: (2025)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
by: Du, Yingjun, et al.
Published: (2023)
by: Du, Yingjun, et al.
Published: (2023)
Active Generation Network of Human Skeleton for Action Recognition
by: Liu, Long, et al.
Published: (2024)
by: Liu, Long, et al.
Published: (2024)
CauSight: Learning to Supersense for Visual Causal Discovery
by: Zhang, Yize, et al.
Published: (2025)
by: Zhang, Yize, et al.
Published: (2025)
A Survey on Interpretability in Visual Recognition
by: Wan, Qiyang, et al.
Published: (2025)
by: Wan, Qiyang, et al.
Published: (2025)
Visual Language Hypothesis
by: Li, Xiu
Published: (2025)
by: Li, Xiu
Published: (2025)
Active Generalized Category Discovery
by: Ma, Shijie, et al.
Published: (2024)
by: Ma, Shijie, et al.
Published: (2024)
CompAgent: An Agentic Framework for Visual Compliance Verification
by: Ghosh, Rahul, et al.
Published: (2025)
by: Ghosh, Rahul, et al.
Published: (2025)
Re-Aligning Language to Visual Objects with an Agentic Workflow
by: Chen, Yuming, et al.
Published: (2025)
by: Chen, Yuming, et al.
Published: (2025)
NYC-Event-VPR: A Large-Scale High-Resolution Event-Based Visual Place Recognition Dataset in Dense Urban Environments
by: Pan, Taiyi, et al.
Published: (2024)
by: Pan, Taiyi, et al.
Published: (2024)
Flexible and Efficient Spatio-Temporal Transformer for Sequential Visual Place Recognition
by: Kiu, Yu, et al.
Published: (2025)
by: Kiu, Yu, et al.
Published: (2025)
FailureAtlas:Mapping the Failure Landscape of T2I Models via Active Exploration
by: Chen, Muxi, et al.
Published: (2025)
by: Chen, Muxi, et al.
Published: (2025)
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
by: Luo, Liqin, et al.
Published: (2025)
by: Luo, Liqin, et al.
Published: (2025)
Distributed Zero-Shot Learning for Visual Recognition
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
by: Chen, Jiahua, et al.
Published: (2026)
by: Chen, Jiahua, et al.
Published: (2026)
Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition
by: Cheng, Jintao, et al.
Published: (2026)
by: Cheng, Jintao, et al.
Published: (2026)
Similar Items
-
Beyond Referring Expressions: Scenario Comprehension Visual Grounding
by: He, Ruozhen, et al.
Published: (2026) -
GViT: Representing Images as Gaussians for Visual Recognition
by: Hernandez, Jefferson, et al.
Published: (2025) -
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
by: Koo, Jaywon, et al.
Published: (2025) -
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
by: Xiao, Zilin, et al.
Published: (2025) -
PropTest: Automatic Property Testing for Improved Visual Programming
by: Koo, Jaywon, et al.
Published: (2024)