Perception of Visual Content: Differences Between Humans and Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Pratama, Nardiena A., Fan, Shaoyang, Demartini, Gianluca |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-Time Canonicalization by Foundation Models for Robust Perception
by: Singhal, Utkarsh, et al.
Published: (2025)
by: Singhal, Utkarsh, et al.
Published: (2025)
Ideology-Based LLMs for Content Moderation
by: Civelli, Stefano, et al.
Published: (2025)
by: Civelli, Stefano, et al.
Published: (2025)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception
by: Garcia, Kathy, et al.
Published: (2025)
by: Garcia, Kathy, et al.
Published: (2025)
Cross-Domain Few-Shot Learning for Hyperspectral Image Classification Based on Mixup Foundation Model
by: Paeedeh, Naeem, et al.
Published: (2026)
by: Paeedeh, Naeem, et al.
Published: (2026)
Semi-Supervised Fine-Tuning of Vision Foundation Models with Content-Style Decomposition
by: Drozdova, Mariia, et al.
Published: (2024)
by: Drozdova, Mariia, et al.
Published: (2024)
Introducing Visual Perception Token into Multimodal Large Language Model
by: Yu, Runpeng, et al.
Published: (2025)
by: Yu, Runpeng, et al.
Published: (2025)
PathoTune: Adapting Visual Foundation Model to Pathological Specialists
by: Lu, Jiaxuan, et al.
Published: (2024)
by: Lu, Jiaxuan, et al.
Published: (2024)
Robust Adaptation of Foundation Models with Black-Box Visual Prompting
by: Oh, Changdae, et al.
Published: (2024)
by: Oh, Changdae, et al.
Published: (2024)
MoFM: A Large-Scale Human Motion Foundation Model
by: Baharani, Mohammadreza, et al.
Published: (2025)
by: Baharani, Mohammadreza, et al.
Published: (2025)
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
by: Wu, Yecheng, et al.
Published: (2024)
by: Wu, Yecheng, et al.
Published: (2024)
Semantic Mosaicing of Histo-Pathology Image Fragments using Visual Foundation Models
by: Brandstätter, Stefan, et al.
Published: (2025)
by: Brandstätter, Stefan, et al.
Published: (2025)
Exploiting Layer Normalization Fine-tuning in Visual Transformer Foundation Models for Classification
by: Tan, Zhaorui, et al.
Published: (2025)
by: Tan, Zhaorui, et al.
Published: (2025)
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
by: Kargin, Turhan Can, et al.
Published: (2026)
by: Kargin, Turhan Can, et al.
Published: (2026)
An Investigation of Visual Foundation Models Robustness
by: Gupta, Sandeep, et al.
Published: (2025)
by: Gupta, Sandeep, et al.
Published: (2025)
Are Large Language Models Good Data Preprocessors?
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
Enhancing Social Media Post Popularity Prediction with Visual Content
by: Jeong, Dahyun, et al.
Published: (2024)
by: Jeong, Dahyun, et al.
Published: (2024)
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
by: Mamaghan, Amir Mohammad Karimi, et al.
Published: (2024)
by: Mamaghan, Amir Mohammad Karimi, et al.
Published: (2024)
Toward Inherently Robust VLMs Against Visual Perception Attacks
by: MohajerAnsari, Pedram, et al.
Published: (2025)
by: MohajerAnsari, Pedram, et al.
Published: (2025)
Human-Aligned Image Models Improve Visual Decoding from the Brain
by: Rajabi, Nona, et al.
Published: (2025)
by: Rajabi, Nona, et al.
Published: (2025)
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs
by: Janjua, Muhammad Kamran, et al.
Published: (2026)
by: Janjua, Muhammad Kamran, et al.
Published: (2026)
Temporal Visual Semantics-Induced Human Motion Understanding with Large Language Models
by: Xing, Zheng, et al.
Published: (2025)
by: Xing, Zheng, et al.
Published: (2025)
FodFoM: Fake Outlier Data by Foundation Models Creates Stronger Visual Out-of-Distribution Detector
by: Chen, Jiankang, et al.
Published: (2024)
by: Chen, Jiankang, et al.
Published: (2024)
NeurAll: Towards a Unified Visual Perception Model for Automated Driving
by: Sistu, Ganesh, et al.
Published: (2019)
by: Sistu, Ganesh, et al.
Published: (2019)
Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering
by: Hu, Xinyue, et al.
Published: (2023)
by: Hu, Xinyue, et al.
Published: (2023)
Reversible Unfolding Network for Concealed Visual Perception with Generative Refinement
by: He, Chunming, et al.
Published: (2025)
by: He, Chunming, et al.
Published: (2025)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
Adaptive Routing of Text-to-Image Generation Requests Between Large Cloud Model and Light-Weight Edge Model
by: Xin, Zewei, et al.
Published: (2024)
by: Xin, Zewei, et al.
Published: (2024)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
by: Wang, Xiyao, et al.
Published: (2025)
by: Wang, Xiyao, et al.
Published: (2025)
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
by: Li, Kaican, et al.
Published: (2025)
by: Li, Kaican, et al.
Published: (2025)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
Bridging the Gap Between Multimodal Foundation Models and World Models
by: He, Xuehai
Published: (2025)
by: He, Xuehai
Published: (2025)
Context Sensitivity Improves Human-Machine Visual Alignment
by: Born, Frieda, et al.
Published: (2026)
by: Born, Frieda, et al.
Published: (2026)
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
by: Zhou, Wenhao, et al.
Published: (2025)
by: Zhou, Wenhao, et al.
Published: (2025)
Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception
by: Loukovitis, Spyridon, et al.
Published: (2025)
by: Loukovitis, Spyridon, et al.
Published: (2025)
Vision Foundation Models in Remote Sensing: A Survey
by: Lu, Siqi, et al.
Published: (2024)
by: Lu, Siqi, et al.
Published: (2024)
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
by: Cho, Jang Hyun, et al.
Published: (2025)
by: Cho, Jang Hyun, et al.
Published: (2025)
Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction
by: Gustafsson, Fredrik K., et al.
Published: (2026)
by: Gustafsson, Fredrik K., et al.
Published: (2026)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
by: Wang, Yubo, et al.
Published: (2024)
by: Wang, Yubo, et al.
Published: (2024)
Similar Items
-
Test-Time Canonicalization by Foundation Models for Robust Perception
by: Singhal, Utkarsh, et al.
Published: (2025) -
Ideology-Based LLMs for Content Moderation
by: Civelli, Stefano, et al.
Published: (2025) -
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025) -
Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception
by: Garcia, Kathy, et al.
Published: (2025) -
Cross-Domain Few-Shot Learning for Hyperspectral Image Classification Based on Mixup Foundation Model
by: Paeedeh, Naeem, et al.
Published: (2026)