If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Barbano, Carlo Alberto, Molinaro, Luca, Ciranni, Massimiliano, Aiello, Emanuele, Pastore, Vito Paolo, Grangetto, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Shape of Math To Come
by: Kontorovich, Alex
Published: (2025)
by: Kontorovich, Alex
Published: (2025)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
by: Brusnicki, Roberto, et al.
Published: (2026)
by: Brusnicki, Roberto, et al.
Published: (2026)
Unsupervised Learning of Unbiased Visual Representations
by: Barbano, Carlo Alberto, et al.
Published: (2022)
by: Barbano, Carlo Alberto, et al.
Published: (2022)
COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
by: Hassan, Umair
Published: (2025)
by: Hassan, Umair
Published: (2025)
Multimodal Structure-Aware Quantum Data Processing
by: Hawashin, Hala, et al.
Published: (2024)
by: Hawashin, Hala, et al.
Published: (2024)
Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
by: Shin, Suyeon, et al.
Published: (2024)
by: Shin, Suyeon, et al.
Published: (2024)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
by: Cui, Hejie, et al.
Published: (2024)
by: Cui, Hejie, et al.
Published: (2024)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
by: Yim, Wen-wai, et al.
Published: (2025)
by: Yim, Wen-wai, et al.
Published: (2025)
Model Surgery: Modulating LLM's Behavior Via Simple Parameter Editing
by: Wang, Huanqian, et al.
Published: (2024)
by: Wang, Huanqian, et al.
Published: (2024)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
by: Zhang, Junwen, et al.
Published: (2025)
by: Zhang, Junwen, et al.
Published: (2025)
Separate Before You Compress: The WWHO Tokenization Architecture
by: Darshana, Kusal
Published: (2026)
by: Darshana, Kusal
Published: (2026)
SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval
by: Wu, Ren-Di, et al.
Published: (2025)
by: Wu, Ren-Di, et al.
Published: (2025)
LLM-supported document separation for printed reviews from zbMATH Open
by: Pluzhnikov, Ivan, et al.
Published: (2026)
by: Pluzhnikov, Ivan, et al.
Published: (2026)
Does CLIP perceive art the same way we do?
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being
by: Ferdousi, Rahatara, et al.
Published: (2025)
by: Ferdousi, Rahatara, et al.
Published: (2025)
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
by: Zhang, Shengkai, et al.
Published: (2024)
by: Zhang, Shengkai, et al.
Published: (2024)
When can forward stable algorithms be composed stably?
by: Beltrán, Carlos, et al.
Published: (2021)
by: Beltrán, Carlos, et al.
Published: (2021)
A Practical Synthesis of Detecting AI-Generated Textual, Visual, and Audio Content
by: Cao, Lele
Published: (2025)
by: Cao, Lele
Published: (2025)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
by: Anam, Rizal Khoirul
Published: (2025)
by: Anam, Rizal Khoirul
Published: (2025)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
by: Yang, Baoyao, et al.
Published: (2025)
by: Yang, Baoyao, et al.
Published: (2025)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
The MSR-Video to Text Dataset with Clean Annotations
by: Chen, Haoran, et al.
Published: (2021)
by: Chen, Haoran, et al.
Published: (2021)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
by: Huo, Dongjie, et al.
Published: (2026)
by: Huo, Dongjie, et al.
Published: (2026)
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
by: Kornilov, Albert, et al.
Published: (2024)
by: Kornilov, Albert, et al.
Published: (2024)
Analysis of multivariate symbol statistics in primitive rational models
by: Goldwurm, Massimiliano, et al.
Published: (2026)
by: Goldwurm, Massimiliano, et al.
Published: (2026)
Depth Priors in Removal Neural Radiance Fields
by: Guo, Zhihao, et al.
Published: (2024)
by: Guo, Zhihao, et al.
Published: (2024)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
by: Andrylie, Lyzander Marciano, et al.
Published: (2025)
Supercharging Federated Intelligence Retrieval
by: Stripelis, Dimitris, et al.
Published: (2026)
by: Stripelis, Dimitris, et al.
Published: (2026)
RefineFormer3D: Efficient 3D Medical Image Segmentation via Adaptive Multi-Scale Transformer with Cross Attention Fusion
by: Tyagi, Kavyansh, et al.
Published: (2026)
by: Tyagi, Kavyansh, et al.
Published: (2026)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
by: Li, Mingda, et al.
Published: (2024)
by: Li, Mingda, et al.
Published: (2024)
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
by: Jin, Hang, et al.
Published: (2025)
by: Jin, Hang, et al.
Published: (2025)
Improving Recursive Transformers with Mixture of LoRAs
by: Nouriborji, Mohammadmahdi, et al.
Published: (2025)
by: Nouriborji, Mohammadmahdi, et al.
Published: (2025)
Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
by: Arabov, Mullosharaf K.
Published: (2026)
by: Arabov, Mullosharaf K.
Published: (2026)
Survey of Swarm Intelligence Approaches to Search Documents Based On Semantic Similarity
by: Muniyappa, Chandrashekar, et al.
Published: (2025)
by: Muniyappa, Chandrashekar, et al.
Published: (2025)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
by: Yao, Huaiyuan, et al.
Published: (2026)
by: Yao, Huaiyuan, et al.
Published: (2026)
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
by: Yu, Zhiqi, et al.
Published: (2026)
by: Yu, Zhiqi, et al.
Published: (2026)
AUTHENTICATION: Identifying Rare Failure Modes in Autonomous Vehicle Perception Systems using Adversarially Guided Diffusion Models
by: Zarei, Mohammad, et al.
Published: (2025)
by: Zarei, Mohammad, et al.
Published: (2025)
Similar Items
-
The Shape of Math To Come
by: Kontorovich, Alex
Published: (2025) -
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
by: Brusnicki, Roberto, et al.
Published: (2026) -
Unsupervised Learning of Unbiased Visual Representations
by: Barbano, Carlo Alberto, et al.
Published: (2022) -
COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
by: Hassan, Umair
Published: (2025) -
Multimodal Structure-Aware Quantum Data Processing
by: Hawashin, Hala, et al.
Published: (2024)