Visual symbolic mechanisms: Emergent symbol processing in vision language models
Fuente:
arXiv
Saved in:
| Main Authors: | Assouel, Rim, Campbell, Declan, Bengio, Yoshua, Webb, Taylor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Binding Visual Features Point by Point
by: Haputhanthri, Udith, et al.
Published: (2026)
by: Haputhanthri, Udith, et al.
Published: (2026)
Amortizing intractable inference in diffusion models for vision, language, and control
by: Venkatraman, Siddarth, et al.
Published: (2024)
by: Venkatraman, Siddarth, et al.
Published: (2024)
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
by: Assouel, Rim, et al.
Published: (2026)
by: Assouel, Rim, et al.
Published: (2026)
Object-centric Binding in Contrastive Language-Image Pretraining
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
Chameleon: Fast-slow Neuro-symbolic Lane Topology Extraction
by: Zhang, Zongzheng, et al.
Published: (2025)
by: Zhang, Zongzheng, et al.
Published: (2025)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
Searching for internal symbols underlying deep learning
by: Lee, Jung H., et al.
Published: (2024)
by: Lee, Jung H., et al.
Published: (2024)
Bringing SAM to new heights: Leveraging elevation data for tree crown segmentation from drone imagery
by: Teng, Mélisande, et al.
Published: (2025)
by: Teng, Mélisande, et al.
Published: (2025)
A Novel Neural-symbolic System under Statistical Relational Learning
by: Yu, Dongran, et al.
Published: (2023)
by: Yu, Dongran, et al.
Published: (2023)
An analysis of vision-language models for fabric retrieval
by: Giuliari, Francesco, et al.
Published: (2025)
by: Giuliari, Francesco, et al.
Published: (2025)
Memory Efficient Neural Processes via Constant Memory Attention Block
by: Feng, Leo, et al.
Published: (2023)
by: Feng, Leo, et al.
Published: (2023)
Efficient Diversity-Preserving Diffusion Alignment via Gradient-Informed GFlowNets
by: Liu, Zhen, et al.
Published: (2024)
by: Liu, Zhen, et al.
Published: (2024)
Visual hallucination detection in large vision-language models via evidential conflict
by: Huang, Tao, et al.
Published: (2025)
by: Huang, Tao, et al.
Published: (2025)
Hyperion -- A fast, versatile symbolic Gaussian Belief Propagation framework for Continuous-Time SLAM
by: Hug, David, et al.
Published: (2024)
by: Hug, David, et al.
Published: (2024)
Learning What Matters: Steering Diffusion via Spectrally Anisotropic Forward Noise
by: Scimeca, Luca, et al.
Published: (2025)
by: Scimeca, Luca, et al.
Published: (2025)
Zero-shot Sequential Neuro-symbolic Reasoning for Automatically Generating Architecture Schematic Designs
by: Kodnongbua, Milin, et al.
Published: (2024)
by: Kodnongbua, Milin, et al.
Published: (2024)
Do large language vision models understand 3D shapes?
by: Eppel, Sagi
Published: (2024)
by: Eppel, Sagi
Published: (2024)
Are vision-language models ready to zero-shot replace supervised classification models in agriculture?
by: Ranario, Earl, et al.
Published: (2025)
by: Ranario, Earl, et al.
Published: (2025)
Few-Shot Learning of Visual Compositional Concepts through Probabilistic Schema Induction
by: Lee, Andrew Jun, et al.
Published: (2025)
by: Lee, Andrew Jun, et al.
Published: (2025)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
bi-modal textual prompt learning for vision-language models in remote sensing
by: Kashyap, Pankhi, et al.
Published: (2026)
by: Kashyap, Pankhi, et al.
Published: (2026)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
by: Scimeca, Luca, et al.
Published: (2023)
by: Scimeca, Luca, et al.
Published: (2023)
Zero-Shot Object-Centric Representation Learning
by: Didolkar, Aniket, et al.
Published: (2024)
by: Didolkar, Aniket, et al.
Published: (2024)
Enhancing medical vision-language contrastive learning via inter-matching relation modelling
by: Li, Mingjian, et al.
Published: (2024)
by: Li, Mingjian, et al.
Published: (2024)
A multi-modal vision-language model for generalizable annotation-free pathology localization
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
by: Huang, Weijian, et al.
Published: (2024)
by: Huang, Weijian, et al.
Published: (2024)
Initialization matters in few-shot adaptation of vision-language models for histopathological image classification
by: Meseguer, Pablo, et al.
Published: (2026)
by: Meseguer, Pablo, et al.
Published: (2026)
Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation
by: Liu, Xiaohong, et al.
Published: (2024)
by: Liu, Xiaohong, et al.
Published: (2024)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Slot Abstractors: Toward Scalable Abstract Visual Reasoning
by: Mondal, Shanka Subhra, et al.
Published: (2024)
by: Mondal, Shanka Subhra, et al.
Published: (2024)
Assessing SAM for Tree Crown Instance Segmentation from Drone Imagery
by: Teng, Mélisande, et al.
Published: (2025)
by: Teng, Mélisande, et al.
Published: (2025)
Interpreting the linear structure of vision-language model embedding spaces
by: Papadimitriou, Isabel, et al.
Published: (2025)
by: Papadimitriou, Isabel, et al.
Published: (2025)
Zero-shot segmentation of skin tumors in whole-slide images with vision-language foundation models
by: Moreno, Santiago, et al.
Published: (2025)
by: Moreno, Santiago, et al.
Published: (2025)
EarthView: A Large Scale Remote Sensing Dataset for Self-Supervision
by: Velazquez, Diego, et al.
Published: (2025)
by: Velazquez, Diego, et al.
Published: (2025)
Frame Sampling Strategies Matter: A Benchmark for small vision language models
by: Brkic, Marija, et al.
Published: (2025)
by: Brkic, Marija, et al.
Published: (2025)
A vision-language model and platform for temporally mapping surgery from video
by: Kiyasseh, Dani
Published: (2026)
by: Kiyasseh, Dani
Published: (2026)
Thinker: A vision-language foundation model for embodied intelligence
by: Pan, Baiyu, et al.
Published: (2026)
by: Pan, Baiyu, et al.
Published: (2026)
(LiFT) Lightweight Fitness Transformer: A language-vision model for Remote Monitoring of Physical Training
by: Postlmayr, A., et al.
Published: (2025)
by: Postlmayr, A., et al.
Published: (2025)
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
by: Nassar, Ahmed, et al.
Published: (2025)
by: Nassar, Ahmed, et al.
Published: (2025)
Similar Items
-
Binding Visual Features Point by Point
by: Haputhanthri, Udith, et al.
Published: (2026) -
Amortizing intractable inference in diffusion models for vision, language, and control
by: Venkatraman, Siddarth, et al.
Published: (2024) -
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
by: Assouel, Rim, et al.
Published: (2026) -
Object-centric Binding in Contrastive Language-Image Pretraining
by: Assouel, Rim, et al.
Published: (2025) -
Chameleon: Fast-slow Neuro-symbolic Lane Topology Extraction
by: Zhang, Zongzheng, et al.
Published: (2025)