VisKnow: Constructing Visual Knowledge Base for Object Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Ziwei, Wan, Qiyang, Wang, Ruiping, Chen, Xilin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey on Interpretability in Visual Recognition
by: Wan, Qiyang, et al.
Published: (2025)
by: Wan, Qiyang, et al.
Published: (2025)
Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models
by: Fu, Bin, et al.
Published: (2024)
by: Fu, Bin, et al.
Published: (2024)
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
by: Bai, Ziyi, et al.
Published: (2024)
by: Bai, Ziyi, et al.
Published: (2024)
GEAR: GEometry-motion Alternating Refinement for Articulated Object Modeling with Gaussian Splatting
by: Li, Jialin, et al.
Published: (2026)
by: Li, Jialin, et al.
Published: (2026)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026)
by: Xiao, Zhanqi, et al.
Published: (2026)
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
by: Zhang, Bin, et al.
Published: (2025)
by: Zhang, Bin, et al.
Published: (2025)
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
GS-LTS: 3D Gaussian Splatting-Based Adaptive Modeling for Long-Term Service Robots
by: Fu, Bin, et al.
Published: (2025)
by: Fu, Bin, et al.
Published: (2025)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
by: Xie, Senwei, et al.
Published: (2025)
by: Xie, Senwei, et al.
Published: (2025)
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
by: Chen, Jian, et al.
Published: (2025)
by: Chen, Jian, et al.
Published: (2025)
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024)
by: Awal, Rabiul, et al.
Published: (2024)
ObjectVisA-120: Object-based Visual Attention Prediction in Interactive Street-crossing Environments
by: Vozniak, Igor, et al.
Published: (2026)
by: Vozniak, Igor, et al.
Published: (2026)
M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models
by: Wang, Hongyu, et al.
Published: (2024)
by: Wang, Hongyu, et al.
Published: (2024)
Improving VisNet for Object Recognition
by: Serj, Mehdi Fatan, et al.
Published: (2025)
by: Serj, Mehdi Fatan, et al.
Published: (2025)
VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation
by: Deng, Jie, et al.
Published: (2026)
by: Deng, Jie, et al.
Published: (2026)
SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
by: Liu, Can, et al.
Published: (2025)
by: Liu, Can, et al.
Published: (2025)
UAV-VisLoc: A Large-scale Dataset for UAV Visual Localization
by: Xu, Wenjia, et al.
Published: (2024)
by: Xu, Wenjia, et al.
Published: (2024)
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
by: Wu, Wenhao, et al.
Published: (2023)
by: Wu, Wenhao, et al.
Published: (2023)
VisGuard: Securing Visualization Dissemination through Tamper-Resistant Data Retrieval
by: Ye, Huayuan, et al.
Published: (2025)
by: Ye, Huayuan, et al.
Published: (2025)
RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations
by: Chen, I-Hsiang, et al.
Published: (2026)
by: Chen, I-Hsiang, et al.
Published: (2026)
CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding
by: Deng, Xiaoyu, et al.
Published: (2024)
by: Deng, Xiaoyu, et al.
Published: (2024)
GLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition
by: Wang, Tianyue, et al.
Published: (2025)
by: Wang, Tianyue, et al.
Published: (2025)
DynamicVis: Dynamic Visual Perception for Efficient Remote Sensing Foundation Models
by: Chen, Keyan, et al.
Published: (2025)
by: Chen, Keyan, et al.
Published: (2025)
DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding
by: Tao, Mingzhe, et al.
Published: (2026)
by: Tao, Mingzhe, et al.
Published: (2026)
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
by: Huang, Zeyi, et al.
Published: (2025)
by: Huang, Zeyi, et al.
Published: (2025)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
by: Cha, Sungguk, et al.
Published: (2024)
by: Cha, Sungguk, et al.
Published: (2024)
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
by: Xiang, Wenzhao, et al.
Published: (2026)
by: Xiang, Wenzhao, et al.
Published: (2026)
Wavelet-Driven Masked Image Modeling: A Path to Efficient Visual Representation
by: Xiang, Wenzhao, et al.
Published: (2025)
by: Xiang, Wenzhao, et al.
Published: (2025)
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
by: Yao, Jianhang, et al.
Published: (2025)
by: Yao, Jianhang, et al.
Published: (2025)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
by: Jiang, Tianxiang, et al.
Published: (2025)
by: Jiang, Tianxiang, et al.
Published: (2025)
Seeking and Updating with Live Visual Knowledge
by: Fu, Mingyang, et al.
Published: (2025)
by: Fu, Mingyang, et al.
Published: (2025)
VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
by: Li, Zhaonan, et al.
Published: (2026)
by: Li, Zhaonan, et al.
Published: (2026)
EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy
by: Ge, Xuanyu, et al.
Published: (2026)
by: Ge, Xuanyu, et al.
Published: (2026)
Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit
by: Chen, Qizhou, et al.
Published: (2024)
by: Chen, Qizhou, et al.
Published: (2024)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
by: Ji, Haonian, et al.
Published: (2025)
by: Ji, Haonian, et al.
Published: (2025)
Object Retrieval for Visual Question Answering with Outside Knowledge
by: Kan, Shichao, et al.
Published: (2024)
by: Kan, Shichao, et al.
Published: (2024)
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models
by: Sima, Bingrui, et al.
Published: (2025)
by: Sima, Bingrui, et al.
Published: (2025)
Similar Items
-
A Survey on Interpretability in Visual Recognition
by: Wan, Qiyang, et al.
Published: (2025) -
Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models
by: Fu, Bin, et al.
Published: (2024) -
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
by: Bai, Ziyi, et al.
Published: (2024) -
GEAR: GEometry-motion Alternating Refinement for Articulated Object Modeling with Gaussian Splatting
by: Li, Jialin, et al.
Published: (2026) -
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026)