Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Junhan, Zhou, Zilu, Tong, Yujun, Chang, Dongliang, Luo, Yitao, Ma, Zhanyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Privacy-Preserving Fine-Grained Visual Classification via Hierarchical Learning from Label Proportions
by: Chang, Jinyi, et al.
Published: (2025)
by: Chang, Jinyi, et al.
Published: (2025)
Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models
by: Tong, Yujun, et al.
Published: (2026)
by: Tong, Yujun, et al.
Published: (2026)
Recolour What Matters: Region-Aware Colour Editing via Token-Level Diffusion
by: Yang, Yuqi, et al.
Published: (2026)
by: Yang, Yuqi, et al.
Published: (2026)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
by: Zhang, Zhengxuan, et al.
Published: (2025)
by: Zhang, Zhengxuan, et al.
Published: (2025)
Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection
by: Qin, Haotian, et al.
Published: (2025)
by: Qin, Haotian, et al.
Published: (2025)
IncreFA: Breaking the Static Wall of Generative Model Attribution
by: Qin, Haotian, et al.
Published: (2026)
by: Qin, Haotian, et al.
Published: (2026)
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
by: Cao, Shuo, et al.
Published: (2025)
by: Cao, Shuo, et al.
Published: (2025)
ShiftedBronzes: Benchmarking and Analysis of Domain Fine-Grained Classification in Open-World Settings
by: Zhou, Rixin, et al.
Published: (2024)
by: Zhou, Rixin, et al.
Published: (2024)
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
by: Li, You, et al.
Published: (2026)
by: Li, You, et al.
Published: (2026)
S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
by: Li, Yingshu, et al.
Published: (2025)
by: Li, Yingshu, et al.
Published: (2025)
Controllable-Continuous Color Editing in Diffusion Model via Color Mapping
by: Yang, Yuqi, et al.
Published: (2025)
by: Yang, Yuqi, et al.
Published: (2025)
Expert Knowledge-Guided Decision Calibration for Accurate Fine-Grained Tree Species Classification
by: Long, Chen, et al.
Published: (2026)
by: Long, Chen, et al.
Published: (2026)
GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection
by: Li, Jiaming, et al.
Published: (2026)
by: Li, Jiaming, et al.
Published: (2026)
Streamlined Open-Vocabulary Human-Object Interaction Detection
by: Sun, Chang, et al.
Published: (2026)
by: Sun, Chang, et al.
Published: (2026)
Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion Guidance
by: Li, Wenjie, et al.
Published: (2025)
by: Li, Wenjie, et al.
Published: (2025)
Complementary Frequency-Varying Awareness Network for Open-Set Fine-Grained Image Recognition
by: Dong, Qiulei, et al.
Published: (2023)
by: Dong, Qiulei, et al.
Published: (2023)
SGIA: Enhancing Fine-Grained Visual Classification with Sequence Generative Image Augmentation
by: Liao, Qiyu, et al.
Published: (2024)
by: Liao, Qiyu, et al.
Published: (2024)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
by: Wang, Ziteng, et al.
Published: (2025)
by: Wang, Ziteng, et al.
Published: (2025)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
by: Lu, Yehao, et al.
Published: (2025)
by: Lu, Yehao, et al.
Published: (2025)
From See to Shield: ML-Assisted Fine-Grained Access Control for Visual Data
by: Akcay, Mete Harun, et al.
Published: (2025)
by: Akcay, Mete Harun, et al.
Published: (2025)
Toward Generalizable Forgery Detection and Reasoning
by: Gao, Yueying, et al.
Published: (2025)
by: Gao, Yueying, et al.
Published: (2025)
Beyond Frequency: Seeing Subtle Cues Through the Lens of Spatial Decomposition for Fine-Grained Visual Classification
by: Xu, Qin, et al.
Published: (2025)
by: Xu, Qin, et al.
Published: (2025)
Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained Experts
by: Xu, Yangyang, et al.
Published: (2025)
by: Xu, Yangyang, et al.
Published: (2025)
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
by: Xue, Junxiao, et al.
Published: (2026)
by: Xue, Junxiao, et al.
Published: (2026)
Detail Reinforcement Diffusion Model: Augmentation Fine-Grained Visual Categorization in Few-Shot Conditions
by: Wu, Tianxu, et al.
Published: (2023)
by: Wu, Tianxu, et al.
Published: (2023)
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classification
by: Luo, Qiuming, et al.
Published: (2026)
by: Luo, Qiuming, et al.
Published: (2026)
Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
Seeing and Knowing in the Wild: Open-domain Visual Entity Recognition with Large-scale Knowledge Graphs via Contrastive Learning
by: Zhou, Hongkuan, et al.
Published: (2025)
by: Zhou, Hongkuan, et al.
Published: (2025)
An LLM-LVLM Driven Agent for Iterative and Fine-Grained Image Editing
by: Liang, Zihan, et al.
Published: (2025)
by: Liang, Zihan, et al.
Published: (2025)
OpenObj: Open-Vocabulary Object-Level Neural Radiance Fields with Fine-Grained Understanding
by: Deng, Yinan, et al.
Published: (2024)
by: Deng, Yinan, et al.
Published: (2024)
Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
by: Ghosh, Dhruba, et al.
Published: (2026)
by: Ghosh, Dhruba, et al.
Published: (2026)
Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding
by: Guo, Leilei, et al.
Published: (2025)
by: Guo, Leilei, et al.
Published: (2025)
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding
by: Zheng, Lihao, et al.
Published: (2026)
by: Zheng, Lihao, et al.
Published: (2026)
MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
by: Li, Geng, et al.
Published: (2025)
by: Li, Geng, et al.
Published: (2025)
Fine-Grained Zero-Shot Object Detection
by: Ma, Hongxu, et al.
Published: (2025)
by: Ma, Hongxu, et al.
Published: (2025)
FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding
by: Feng, Kaidong, et al.
Published: (2026)
by: Feng, Kaidong, et al.
Published: (2026)
Similar Items
-
Towards Privacy-Preserving Fine-Grained Visual Classification via Hierarchical Learning from Label Proportions
by: Chang, Jinyi, et al.
Published: (2025) -
Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models
by: Tong, Yujun, et al.
Published: (2026) -
Recolour What Matters: Region-Aware Colour Editing via Token-Level Diffusion
by: Yang, Yuqi, et al.
Published: (2026) -
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
by: Zhang, Zhengxuan, et al.
Published: (2025) -
Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection
by: Qin, Haotian, et al.
Published: (2025)