Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
Fuente:
arXiv
Saved in:
| Main Authors: | Abdullahu, Fitim, Grabner, Helmut |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Commonly Interesting Images
by: Abdullahu, Fitim, et al.
Published: (2024)
by: Abdullahu, Fitim, et al.
Published: (2024)
Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers
by: Immertreu, Mathis, et al.
Published: (2026)
by: Immertreu, Mathis, et al.
Published: (2026)
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
by: Zhang, Shilong, et al.
Published: (2023)
by: Zhang, Shilong, et al.
Published: (2023)
Mirror-Aware Neural Humans
by: Ajisafe, Daniel, et al.
Published: (2023)
by: Ajisafe, Daniel, et al.
Published: (2023)
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
by: Wu, Wenhao, et al.
Published: (2023)
by: Wu, Wenhao, et al.
Published: (2023)
MirrorCalib: Utilizing Human Pose Information for Mirror-based Virtual Camera Calibration
by: Liao, Longyun, et al.
Published: (2023)
by: Liao, Longyun, et al.
Published: (2023)
Visual Instruction Tuning with Chain of Region-of-Interest
by: Chen, Yixin, et al.
Published: (2025)
by: Chen, Yixin, et al.
Published: (2025)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
by: Wu, Yiqi, et al.
Published: (2024)
by: Wu, Yiqi, et al.
Published: (2024)
Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
by: Masri, Sari, et al.
Published: (2025)
by: Masri, Sari, et al.
Published: (2025)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
Illusions in Humans and AI: How Visual Perception Aligns and Diverges
by: Yang, Jianyi, et al.
Published: (2025)
by: Yang, Jianyi, et al.
Published: (2025)
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
by: Lu, Hao, et al.
Published: (2024)
by: Lu, Hao, et al.
Published: (2024)
An Empirical Study of GPT-4o Image Generation Capabilities
by: Chen, Sixiang, et al.
Published: (2025)
by: Chen, Sixiang, et al.
Published: (2025)
Decoding Visual Sentiment of Political Imagery
by: Gasparyan, Olga, et al.
Published: (2024)
by: Gasparyan, Olga, et al.
Published: (2024)
MirrorSAM2: Segment Mirror in Videos with Depth Perception
by: Xu, Mingchen, et al.
Published: (2025)
by: Xu, Mingchen, et al.
Published: (2025)
Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
by: Guo, Haonan, et al.
Published: (2024)
by: Guo, Haonan, et al.
Published: (2024)
Exploring Visual Culture Awareness in GPT-4V: A Comprehensive Probing
by: Cao, Yong, et al.
Published: (2024)
by: Cao, Yong, et al.
Published: (2024)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Leveraging YOLO-World and GPT-4V LMMs for Zero-Shot Person Detection and Action Recognition in Drone Imagery
by: Limberg, Christian, et al.
Published: (2024)
by: Limberg, Christian, et al.
Published: (2024)
Rejuvenating image-GPT as Strong Visual Representation Learners
by: Ren, Sucheng, et al.
Published: (2023)
by: Ren, Sucheng, et al.
Published: (2023)
MirrorGaussian: Reflecting 3D Gaussians for Reconstructing Mirror Reflections
by: Liu, Jiayue, et al.
Published: (2024)
by: Liu, Jiayue, et al.
Published: (2024)
Human-Aligned Image Models Improve Visual Decoding from the Brain
by: Rajabi, Nona, et al.
Published: (2025)
by: Rajabi, Nona, et al.
Published: (2025)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
by: Ataallah, Kirolos, et al.
Published: (2024)
by: Ataallah, Kirolos, et al.
Published: (2024)
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
by: Chen, Zhe, et al.
Published: (2024)
by: Chen, Zhe, et al.
Published: (2024)
Mirror-Yolo: A Novel Attention Focus, Instance Segmentation and Mirror Detection Model
by: Li, Fengze, et al.
Published: (2022)
by: Li, Fengze, et al.
Published: (2022)
Exploring The Visual Feature Space for Multimodal Neural Decoding
by: Xia, Weihao, et al.
Published: (2025)
by: Xia, Weihao, et al.
Published: (2025)
Decoding Visual Neural Representations by Multimodal with Dynamic Balancing
by: sun, Kaili, et al.
Published: (2025)
by: sun, Kaili, et al.
Published: (2025)
Decoding Functional Networks for Visual Categories via GNNs
by: Karmi, Shira, et al.
Published: (2026)
by: Karmi, Shira, et al.
Published: (2026)
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
by: He, Zhendong, et al.
Published: (2026)
by: He, Zhendong, et al.
Published: (2026)
EDTformer: An Efficient Decoder Transformer for Visual Place Recognition
by: Jin, Tong, et al.
Published: (2024)
by: Jin, Tong, et al.
Published: (2024)
Contextual Encoder-Decoder Network for Visual Saliency Prediction
by: Kroner, Alexander, et al.
Published: (2019)
by: Kroner, Alexander, et al.
Published: (2019)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
by: Shen, Xiaoqian, et al.
Published: (2023)
by: Shen, Xiaoqian, et al.
Published: (2023)
Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient
by: Chen, Zigeng, et al.
Published: (2024)
by: Chen, Zigeng, et al.
Published: (2024)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
by: Zhao, Jiahe, et al.
Published: (2025)
by: Zhao, Jiahe, et al.
Published: (2025)
SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance
by: Xia, Qi, et al.
Published: (2026)
by: Xia, Qi, et al.
Published: (2026)
Can GPT-4 Models Detect Misleading Visualizations?
by: Alexander, Jason, et al.
Published: (2024)
by: Alexander, Jason, et al.
Published: (2024)
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)
by: Zhang, Jiangning, et al.
Published: (2023)
Evaluation of GPT-4o and GPT-4o-mini's Vision Capabilities for Compositional Analysis from Dried Solution Drops
by: Dangi, Deven B., et al.
Published: (2024)
by: Dangi, Deven B., et al.
Published: (2024)
SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension
by: Jiang, Yue, et al.
Published: (2025)
by: Jiang, Yue, et al.
Published: (2025)
REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding
by: Tai, Yan, et al.
Published: (2025)
by: Tai, Yan, et al.
Published: (2025)
Similar Items
-
Commonly Interesting Images
by: Abdullahu, Fitim, et al.
Published: (2024) -
Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers
by: Immertreu, Mathis, et al.
Published: (2026) -
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
by: Zhang, Shilong, et al.
Published: (2023) -
Mirror-Aware Neural Humans
by: Ajisafe, Daniel, et al.
Published: (2023) -
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
by: Wu, Wenhao, et al.
Published: (2023)