Context-Infused Visual Grounding for Art
Fuente:
arXiv
Saved in:
| Main Authors: | Khan, Selina, van Noord, Nanne |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stylistic Multi-Task Analysis of Ukiyo-e Woodblock Prints
by: Khan, Selina, et al.
Published: (2024)
by: Khan, Selina, et al.
Published: (2024)
The Iconicity of the Generated Image
by: van Noord, Nanne, et al.
Published: (2025)
by: van Noord, Nanne, et al.
Published: (2025)
EMPLACE: Self-Supervised Urban Scene Change Detection
by: Alpherts, Tim, et al.
Published: (2025)
by: Alpherts, Tim, et al.
Published: (2025)
Artifacts of Idiosyncracy in Global Street View Data
by: Alpherts, Tim, et al.
Published: (2025)
by: Alpherts, Tim, et al.
Published: (2025)
Find the Cliffhanger: Multi-Modal Trailerness in Soap Operas
by: Bretti, Carlo, et al.
Published: (2024)
by: Bretti, Carlo, et al.
Published: (2024)
GO4Align: Group Optimization for Multi-Task Alignment
by: Shen, Jiayi, et al.
Published: (2024)
by: Shen, Jiayi, et al.
Published: (2024)
TULIP: Token-length Upgraded CLIP
by: Najdenkoska, Ivona, et al.
Published: (2024)
by: Najdenkoska, Ivona, et al.
Published: (2024)
No Annotations for Object Detection in Art through Stable Diffusion
by: Ramos, Patrick, et al.
Published: (2024)
by: Ramos, Patrick, et al.
Published: (2024)
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
Infusing Environmental Captions for Long-Form Video Language Grounding
by: Lee, Hyogun, et al.
Published: (2024)
by: Lee, Hyogun, et al.
Published: (2024)
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
by: Shen, Ying, et al.
Published: (2026)
by: Shen, Ying, et al.
Published: (2026)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
by: Xu, Yunqiu, et al.
Published: (2024)
by: Xu, Yunqiu, et al.
Published: (2024)
The Art of Deception: Color Visual Illusions and Diffusion Models
by: Gomez-Villa, Alex, et al.
Published: (2024)
by: Gomez-Villa, Alex, et al.
Published: (2024)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
by: Demidov, Dmitry, et al.
Published: (2025)
by: Demidov, Dmitry, et al.
Published: (2025)
LidaRefer: Context-aware Outdoor 3D Visual Grounding for Autonomous Driving
by: Baek, Yeong-Seung, et al.
Published: (2024)
by: Baek, Yeong-Seung, et al.
Published: (2024)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
by: Munasinghe, Shehan, et al.
Published: (2024)
by: Munasinghe, Shehan, et al.
Published: (2024)
Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
by: Wang, Ruiyu, et al.
Published: (2025)
by: Wang, Ruiyu, et al.
Published: (2025)
Infused Suppression Of Magnification Artefacts For Micro-AU Detection
by: Khor, Huai-Qian, et al.
Published: (2025)
by: Khor, Huai-Qian, et al.
Published: (2025)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
The Context of Crash Occurrence: A Complexity-Infused Approach Integrating Semantic, Contextual, and Kinematic Features
by: Wang, Meng, et al.
Published: (2024)
by: Wang, Meng, et al.
Published: (2024)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
by: Jiang, Yubo, et al.
Published: (2026)
by: Jiang, Yubo, et al.
Published: (2026)
"Jutters"
by: Driessen, Meike, et al.
Published: (2025)
by: Driessen, Meike, et al.
Published: (2025)
An Attention Infused Deep Learning System with Grad-CAM Visualization for Early Screening of Glaucoma
by: Swaminathan, Ramanathan
Published: (2025)
by: Swaminathan, Ramanathan
Published: (2025)
On the Role of Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2024)
by: Reich, Daniel, et al.
Published: (2024)
Infusing fine-grained visual knowledge to Vision-Language Models
by: Ypsilantis, Nikolaos-Antonios, et al.
Published: (2025)
by: Ypsilantis, Nikolaos-Antonios, et al.
Published: (2025)
SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning
by: Thoker, Fida Mohammad, et al.
Published: (2025)
by: Thoker, Fida Mohammad, et al.
Published: (2025)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
Context-Guided Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
Direct Visual Grounding by Directing Attention of Visual Tokens
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
by: Cuttano, Claudia, et al.
Published: (2024)
by: Cuttano, Claudia, et al.
Published: (2024)
Learning Implicit Features with Flow Infused Attention for Realistic Virtual Try-On
by: Zhang, Delong, et al.
Published: (2024)
by: Zhang, Delong, et al.
Published: (2024)
InPK: Infusing Prior Knowledge into Prompt for Vision-Language Models
by: Zhou, Shuchang, et al.
Published: (2025)
by: Zhou, Shuchang, et al.
Published: (2025)
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
by: Li, Peizheng, et al.
Published: (2025)
by: Li, Peizheng, et al.
Published: (2025)
Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
Towards Visual Grounding: A Survey
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
Grounded Reinforcement Learning for Visual Reasoning
by: Sarch, Gabriel, et al.
Published: (2025)
by: Sarch, Gabriel, et al.
Published: (2025)
Visual Intention Grounding for Egocentric Assistants
by: Sun, Pengzhan, et al.
Published: (2025)
by: Sun, Pengzhan, et al.
Published: (2025)
Similar Items
-
Stylistic Multi-Task Analysis of Ukiyo-e Woodblock Prints
by: Khan, Selina, et al.
Published: (2024) -
The Iconicity of the Generated Image
by: van Noord, Nanne, et al.
Published: (2025) -
EMPLACE: Self-Supervised Urban Scene Change Detection
by: Alpherts, Tim, et al.
Published: (2025) -
Artifacts of Idiosyncracy in Global Street View Data
by: Alpherts, Tim, et al.
Published: (2025) -
Find the Cliffhanger: Multi-Modal Trailerness in Soap Operas
by: Bretti, Carlo, et al.
Published: (2024)