ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Fanelli, Nicola, Vessio, Gennaro, Castellano, Giovanna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting
by: Fanelli, Nicola, et al.
Published: (2024)
by: Fanelli, Nicola, et al.
Published: (2024)
Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation
by: Rinaldi, Ivan, et al.
Published: (2024)
by: Rinaldi, Ivan, et al.
Published: (2024)
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
by: Rinaldi, Ivan, et al.
Published: (2026)
by: Rinaldi, Ivan, et al.
Published: (2026)
Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA
by: De Marinis, Pasquale, et al.
Published: (2025)
by: De Marinis, Pasquale, et al.
Published: (2025)
RoWeeder: Unsupervised Weed Mapping through Crop-Row Detection
by: De Marinis, Pasquale, et al.
Published: (2024)
by: De Marinis, Pasquale, et al.
Published: (2024)
Label Anything: Multi-Class Few-Shot Semantic Segmentation with Visual Prompts
by: De Marinis, Pasquale, et al.
Published: (2024)
by: De Marinis, Pasquale, et al.
Published: (2024)
Matching-Based Few-Shot Semantic Segmentation Models Are Interpretable by Design
by: De Marinis, Pasquale, et al.
Published: (2025)
by: De Marinis, Pasquale, et al.
Published: (2025)
DistillFSS: Synthesizing Few-Shot Knowledge into a Lightweight Segmentation Model
by: De Marinis, Pasquale, et al.
Published: (2025)
by: De Marinis, Pasquale, et al.
Published: (2025)
Explainable offline automatic signature verifier to support forensic handwriting examiners
by: Diaz, Moises, et al.
Published: (2024)
by: Diaz, Moises, et al.
Published: (2024)
Dynamically enhanced static handwriting representation for Parkinson's disease detection
by: Diaz, Moises, et al.
Published: (2024)
by: Diaz, Moises, et al.
Published: (2024)
Smell and Emotion: Recognising emotions in smell-related artworks
by: Patoliya, Vishal, et al.
Published: (2024)
by: Patoliya, Vishal, et al.
Published: (2024)
HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
by: Peirone, Simone Alberto, et al.
Published: (2025)
by: Peirone, Simone Alberto, et al.
Published: (2025)
Disentangle and denoise: Tackling context misalignment for video moment retrieval
by: Ma, Kaijing, et al.
Published: (2024)
by: Ma, Kaijing, et al.
Published: (2024)
Speaking images. A novel framework for the automated self-description of artworks
by: Bernasconi, Valentine, et al.
Published: (2025)
by: Bernasconi, Valentine, et al.
Published: (2025)
PReP: Efficient context-based shape retrieval for missing parts
by: Fotis, Vlassis, et al.
Published: (2024)
by: Fotis, Vlassis, et al.
Published: (2024)
EHWGesture -- A dataset for multimodal understanding of clinical gestures
by: Amprimo, Gianluca, et al.
Published: (2025)
by: Amprimo, Gianluca, et al.
Published: (2025)
The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding
by: Bianchi, Lorenzo, et al.
Published: (2023)
by: Bianchi, Lorenzo, et al.
Published: (2023)
Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization
by: Sargent, Kyle, et al.
Published: (2025)
by: Sargent, Kyle, et al.
Published: (2025)
Towards a multimodal framework for remote sensing image change retrieval and captioning
by: Ferrod, Roger, et al.
Published: (2024)
by: Ferrod, Roger, et al.
Published: (2024)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
by: Wu, Yiqi, et al.
Published: (2024)
by: Wu, Yiqi, et al.
Published: (2024)
In-context learning enables multimodal large language models to classify cancer pathology images
by: Ferber, Dyke, et al.
Published: (2024)
by: Ferber, Dyke, et al.
Published: (2024)
Explaining multimodal LLMs via intra-modal token interactions
by: Liang, Jiawei, et al.
Published: (2025)
by: Liang, Jiawei, et al.
Published: (2025)
What do vision-language models see in the context? Investigating multimodal in-context learning
by: Santos, Gabriel O. dos, et al.
Published: (2025)
by: Santos, Gabriel O. dos, et al.
Published: (2025)
DeepSeek-OCR: Contexts Optical Compression
by: Wei, Haoran, et al.
Published: (2025)
by: Wei, Haoran, et al.
Published: (2025)
A multimodal gesture recognition dataset for desktop human-computer interaction
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
DeepSeek-OCR 2: Visual Causal Flow
by: Wei, Haoran, et al.
Published: (2026)
by: Wei, Haoran, et al.
Published: (2026)
A multimodal dataset for understanding the impact of mobile phones on remote online virtual education
by: Daza, Roberto, et al.
Published: (2024)
by: Daza, Roberto, et al.
Published: (2024)
PathReasoning: A multimodal reasoning agent for query-based ROI navigation on whole-slide images
by: Zhang, Kunpeng, et al.
Published: (2025)
by: Zhang, Kunpeng, et al.
Published: (2025)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
by: Xu, Lijian, et al.
Published: (2024)
by: Xu, Lijian, et al.
Published: (2024)
SalsaAgent: A multimodal embodied language model for interactive dance generation
by: Yazdian, Payam Jome, et al.
Published: (2026)
by: Yazdian, Payam Jome, et al.
Published: (2026)
Similarity over Factuality: Are we making progress on multimodal out-of-context misinformation detection?
by: Papadopoulos, Stefanos-Iordanis, et al.
Published: (2024)
by: Papadopoulos, Stefanos-Iordanis, et al.
Published: (2024)
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter
by: Yuan, Zhengqing, et al.
Published: (2023)
by: Yuan, Zhengqing, et al.
Published: (2023)
Seek-CAD: A Self-refined Generative Modeling for 3D Parametric CAD Using Local Inference via DeepSeek
by: Li, Xueyang, et al.
Published: (2025)
by: Li, Xueyang, et al.
Published: (2025)
CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation
by: Sharma, Sonali, et al.
Published: (2026)
by: Sharma, Sonali, et al.
Published: (2026)
DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities
by: Islam, Chashi Mahiul, et al.
Published: (2025)
by: Islam, Chashi Mahiul, et al.
Published: (2025)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
by: Xing, Yang, et al.
Published: (2026)
by: Xing, Yang, et al.
Published: (2026)
Attacks on multimodal models
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
Multi-SIGATnet: A multimodal schizophrenia MRI classification algorithm using sparse interaction mechanisms and graph attention networks
by: Jiao, Yuhong, et al.
Published: (2024)
by: Jiao, Yuhong, et al.
Published: (2024)
Can LLMs Assist Computer Education? an Empirical Case Study of DeepSeek
by: Xiao, Dongfu, et al.
Published: (2025)
by: Xiao, Dongfu, et al.
Published: (2025)
Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
by: Liang, Yunhao, et al.
Published: (2026)
by: Liang, Yunhao, et al.
Published: (2026)
Similar Items
-
I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting
by: Fanelli, Nicola, et al.
Published: (2024) -
Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation
by: Rinaldi, Ivan, et al.
Published: (2024) -
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
by: Rinaldi, Ivan, et al.
Published: (2026) -
Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA
by: De Marinis, Pasquale, et al.
Published: (2025) -
RoWeeder: Unsupervised Weed Mapping through Crop-Row Detection
by: De Marinis, Pasquale, et al.
Published: (2024)