Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Immertreu, Mathis, Abdullahu, Fitim, Kinfe, Thomas, Grabner, Helmut, Krauss, Patrick, Schilling, Achim |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
by: Abdullahu, Fitim, et al.
Published: (2025)
by: Abdullahu, Fitim, et al.
Published: (2025)
Commonly Interesting Images
by: Abdullahu, Fitim, et al.
Published: (2024)
by: Abdullahu, Fitim, et al.
Published: (2024)
Word Class Representations Spontaneously Emerge from Successor Representations Trained on Natural Language
by: Immertreu, Mathis, et al.
Published: (2026)
by: Immertreu, Mathis, et al.
Published: (2026)
Probing for Consciousness in Machines
by: Immertreu, Mathis, et al.
Published: (2024)
by: Immertreu, Mathis, et al.
Published: (2024)
Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
by: Hütten, Nils, et al.
Published: (2025)
by: Hütten, Nils, et al.
Published: (2025)
Convergent Representations of Linguistic Constructions in Human and Artificial Neural Systems
by: Ramezani, Pegah, et al.
Published: (2026)
by: Ramezani, Pegah, et al.
Published: (2026)
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
by: Huang, Zilin, et al.
Published: (2026)
by: Huang, Zilin, et al.
Published: (2026)
Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI
by: Georgenthum, Hugo, et al.
Published: (2025)
by: Georgenthum, Hugo, et al.
Published: (2025)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
by: Sinha, Rohit, et al.
Published: (2026)
by: Sinha, Rohit, et al.
Published: (2026)
NeuralDiffuser: Neuroscience-inspired Diffusion Guidance for fMRI Visual Reconstruction
by: Li, Haoyu, et al.
Published: (2024)
by: Li, Haoyu, et al.
Published: (2024)
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
by: Chen, Xinwang, et al.
Published: (2024)
by: Chen, Xinwang, et al.
Published: (2024)
Brain-Inspired Capture: Evidence-Driven Neuromimetic Perceptual Simulation for Visual Decoding
by: Shao, Feixue, et al.
Published: (2026)
by: Shao, Feixue, et al.
Published: (2026)
A Cognitive Process-Inspired Architecture for Subject-Agnostic Brain Visual Decoding
by: Lu, Jingyu, et al.
Published: (2025)
by: Lu, Jingyu, et al.
Published: (2025)
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
by: Liao, Haicheng, et al.
Published: (2025)
by: Liao, Haicheng, et al.
Published: (2025)
Analysis and Visualization of Linguistic Structures in Large Language Models: Neural Representations of Verb-Particle Constructions in BERT
by: Kissane, Hassane, et al.
Published: (2024)
by: Kissane, Hassane, et al.
Published: (2024)
TSOM: Small Object Motion Detection Neural Network Inspired by Avian Visual Circuit
by: Hu, Pignge, et al.
Published: (2024)
by: Hu, Pignge, et al.
Published: (2024)
Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective
by: Manh, Bui Duc, et al.
Published: (2025)
by: Manh, Bui Duc, et al.
Published: (2025)
Neural Brain: A Neuroscience-inspired Framework for Embodied Agents
by: Liu, Jian, et al.
Published: (2025)
by: Liu, Jian, et al.
Published: (2025)
Rethinking Visual Information Processing in Multimodal LLMs
by: Kim, Dongwan, et al.
Published: (2025)
by: Kim, Dongwan, et al.
Published: (2025)
Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models
by: Babaiee, Zahra, et al.
Published: (2025)
by: Babaiee, Zahra, et al.
Published: (2025)
No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Modality-Driven Design for Multi-Step Dexterous Manipulation: Insights from Neuroscience
by: Wake, Naoki, et al.
Published: (2024)
by: Wake, Naoki, et al.
Published: (2024)
AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents
by: Liang, Jiafeng, et al.
Published: (2025)
by: Liang, Jiafeng, et al.
Published: (2025)
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
by: Chi, Donghwan, et al.
Published: (2025)
by: Chi, Donghwan, et al.
Published: (2025)
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
True Multimodal In-Context Learning Needs Attention to the Visual Context
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
by: Liu, Yufang, et al.
Published: (2025)
by: Liu, Yufang, et al.
Published: (2025)
PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
by: Blume, Ansel, et al.
Published: (2025)
by: Blume, Ansel, et al.
Published: (2025)
Preserving Cross-Modal Stability for Visual Unlearning in Multimodal Scenarios
by: Li, Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu
Published: (2025)
by: Li, Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu
Published: (2025)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
by: Cornet, Clément, et al.
Published: (2025)
by: Cornet, Clément, et al.
Published: (2025)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
by: Che, Chang, et al.
Published: (2024)
by: Che, Chang, et al.
Published: (2024)
Training Transitive and Commutative Multimodal Transformers with LoReTTa
by: Tran, Manuel, et al.
Published: (2023)
by: Tran, Manuel, et al.
Published: (2023)
Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
by: Vlachogiannis, Dimitrios N., et al.
Published: (2025)
by: Vlachogiannis, Dimitrios N., et al.
Published: (2025)
Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
by: Deng, Boyang, et al.
Published: (2025)
by: Deng, Boyang, et al.
Published: (2025)
Cognitively-Inspired Tokens Overcome Egocentric Bias in Multimodal Models
by: Leonard, Bridget, et al.
Published: (2026)
by: Leonard, Bridget, et al.
Published: (2026)
Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model
by: She, Yifei, et al.
Published: (2025)
by: She, Yifei, et al.
Published: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025)
by: Choi, Changin, et al.
Published: (2025)
Similar Items
-
Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
by: Abdullahu, Fitim, et al.
Published: (2025) -
Commonly Interesting Images
by: Abdullahu, Fitim, et al.
Published: (2024) -
Word Class Representations Spontaneously Emerge from Successor Representations Trained on Natural Language
by: Immertreu, Mathis, et al.
Published: (2026) -
Probing for Consciousness in Machines
by: Immertreu, Mathis, et al.
Published: (2024) -
Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
by: Hütten, Nils, et al.
Published: (2025)