Saved in:
| Main Authors: | Gizdov, Andrey, Procopio, Andrea, Li, Yichen, Harari, Daniel, Ullman, Tomer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.00365 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Human-Like Coarse Object Representations in Vision Models
by: Gizdov, Andrey, et al.
Published: (2026)
by: Gizdov, Andrey, et al.
Published: (2026)
Time-Series at the Edge: Tiny Separable CNNs for Wearable Gait Detection and Optimal Sensor Placement
by: Procopio, Andrea, et al.
Published: (2025)
by: Procopio, Andrea, et al.
Published: (2025)
Chain of Time: In-Context Physical Simulation with Image Generation Models
by: Wang, YingQiao, et al.
Published: (2025)
by: Wang, YingQiao, et al.
Published: (2025)
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
by: Englebert, Alexandre, et al.
Published: (2024)
by: Englebert, Alexandre, et al.
Published: (2024)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Hallucination-aware intermediate representation edit in large vision-language models
by: Suo, Wei, et al.
Published: (2026)
by: Suo, Wei, et al.
Published: (2026)
Improving vision-language alignment with graph spiking hybrid Networks
by: Zhang, Siyu, et al.
Published: (2025)
by: Zhang, Siyu, et al.
Published: (2025)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
by: Ullman, Tomer
Published: (2024)
by: Ullman, Tomer
Published: (2024)
AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Human alignment of neural network representations
by: Muttenthaler, Lukas, et al.
Published: (2022)
by: Muttenthaler, Lukas, et al.
Published: (2022)
Untrained neural networks can demonstrate memorization-independent abstract reasoning
by: Barak, Tomer, et al.
Published: (2024)
by: Barak, Tomer, et al.
Published: (2024)
Towards a vision foundation model for comprehensive assessment of Cardiac MRI
by: Jacob, Athira J, et al.
Published: (2024)
by: Jacob, Athira J, et al.
Published: (2024)
AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding
by: Jabin, Md Asaduzzaman, et al.
Published: (2025)
by: Jabin, Md Asaduzzaman, et al.
Published: (2025)
Generalizing vision-language models to novel domains: A comprehensive survey
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
Self-supervised video pretraining yields robust and more human-aligned visual representations
by: Parthasarathy, Nikhil, et al.
Published: (2022)
by: Parthasarathy, Nikhil, et al.
Published: (2022)
Evaluating alignment between humans and neural network representations in image-based learning tasks
by: Demircan, Can, et al.
Published: (2023)
by: Demircan, Can, et al.
Published: (2023)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
PEAR: Pixel-aligned Expressive humAn mesh Recovery
by: Wu, Jiahao, et al.
Published: (2026)
by: Wu, Jiahao, et al.
Published: (2026)
Interpreting vision transformers via residual replacement model
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection
by: Harari, Daniel, et al.
Published: (2025)
by: Harari, Daniel, et al.
Published: (2025)
Dimensions underlying the representational alignment of deep neural networks with humans
by: Mahner, Florian P., et al.
Published: (2024)
by: Mahner, Florian P., et al.
Published: (2024)
DOGR: Towards Versatile Visual Document Grounding and Referring
by: Zhou, Yinan, et al.
Published: (2024)
by: Zhou, Yinan, et al.
Published: (2024)
Auto-regressive transformation for image alignment
by: Lee, Kanggeon, et al.
Published: (2025)
by: Lee, Kanggeon, et al.
Published: (2025)
Thinker: A vision-language foundation model for embodied intelligence
by: Pan, Baiyu, et al.
Published: (2026)
by: Pan, Baiyu, et al.
Published: (2026)
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026)
by: Yu, Kang, et al.
Published: (2026)
RadEdit: stress-testing biomedical vision models via diffusion image editing
by: Pérez-García, Fernando, et al.
Published: (2023)
by: Pérez-García, Fernando, et al.
Published: (2023)
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein
by: Guo, Xiaotong, et al.
Published: (2025)
by: Guo, Xiaotong, et al.
Published: (2025)
Phantom: Subject-consistent video generation via cross-modal alignment
by: Liu, Lijie, et al.
Published: (2025)
by: Liu, Lijie, et al.
Published: (2025)
Building and better understanding vision-language models: insights and future directions
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
MMToM-QA: Multimodal Theory of Mind Question Answering
by: Jin, Chuanyang, et al.
Published: (2024)
by: Jin, Chuanyang, et al.
Published: (2024)
Towards Self-Improvement of Diffusion Models via Group Preference Optimization
by: Chen, Renjie, et al.
Published: (2025)
by: Chen, Renjie, et al.
Published: (2025)
Dilated Convolution with Learnable Spacings makes visual models more aligned with humans: a Grad-CAM study
by: Chamas, Rabih, et al.
Published: (2024)
by: Chamas, Rabih, et al.
Published: (2024)
Spread them Apart: Towards Robust Watermarking of Generated Content
by: Pautov, Mikhail, et al.
Published: (2025)
by: Pautov, Mikhail, et al.
Published: (2025)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
Addressing a fundamental limitation in deep vision models: lack of spatial attention
by: Borji, Ali
Published: (2024)
by: Borji, Ali
Published: (2024)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024)
by: Pariza, Valentinos, et al.
Published: (2024)
Representation geometry shapes task performance in vision-language modeling for CT enterography
by: Minoccheri, Cristian, et al.
Published: (2026)
by: Minoccheri, Cristian, et al.
Published: (2026)
Beyond the Hype: A dispassionate look at vision-language models in medical scenario
by: Nan, Yang, et al.
Published: (2024)
by: Nan, Yang, et al.
Published: (2024)
Similar Items
-
Human-Like Coarse Object Representations in Vision Models
by: Gizdov, Andrey, et al.
Published: (2026) -
Time-Series at the Edge: Tiny Separable CNNs for Wearable Gait Detection and Optimal Sensor Placement
by: Procopio, Andrea, et al.
Published: (2025) -
Chain of Time: In-Context Physical Simulation with Image Generation Models
by: Wang, YingQiao, et al.
Published: (2025) -
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
by: Englebert, Alexandre, et al.
Published: (2024) -
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)