Learning to See the Elephant in the Room: Self-Supervised Context Reasoning in Humans and AI
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiao, Sarker, Soumick, Sikarwar, Ankur, Kiely, Bryan Atista, Kreiman, Gabriel, Shi, Zenglin, Zhang, Mengmi |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tuned Compositional Feature Replays for Efficient Stream Learning
by: Talbot, Morgan B., et al.
Published: (2021)
by: Talbot, Morgan B., et al.
Published: (2021)
Unveiling the Tapestry: the Interplay of Generalization and Forgetting in Continual Learning
by: Shi, Zenglin, et al.
Published: (2022)
by: Shi, Zenglin, et al.
Published: (2022)
Improving generalization by mimicking the human visual diet
by: Madan, Spandan, et al.
Published: (2022)
by: Madan, Spandan, et al.
Published: (2022)
Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
by: Cai, Yusen, et al.
Published: (2025)
by: Cai, Yusen, et al.
Published: (2025)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)
by: Yang, Qian, et al.
Published: (2026)
Learning to Perceive "Where": Spatial Pretext Tasks for Robust Self-Supervised Learning
by: Shen, Yang, et al.
Published: (2026)
by: Shen, Yang, et al.
Published: (2026)
Transcending the Annotation Bottleneck: AI-Powered Discovery in Biology and Medicine
by: Chatterjee, Soumick
Published: (2026)
by: Chatterjee, Soumick
Published: (2026)
Communicating about Space: Language-Mediated Spatial Integration Across Partial Views
by: Sikarwar, Ankur, et al.
Published: (2026)
by: Sikarwar, Ankur, et al.
Published: (2026)
MVEB: Self-Supervised Learning with Multi-View Entropy Bottleneck
by: Wen, Liangjian, et al.
Published: (2024)
by: Wen, Liangjian, et al.
Published: (2024)
Can Machines Imitate Humans? Integrative Turing-like tests for Language and Vision Demonstrate a Narrowing Gap
by: Zhang, Mengmi, et al.
Published: (2022)
by: Zhang, Mengmi, et al.
Published: (2022)
PRISM: Progressive Reasoning through Iterative Slot Memory for Vision
by: Wang, Ziyu, et al.
Published: (2026)
by: Wang, Ziyu, et al.
Published: (2026)
VANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre-Training
by: Nazeri, Mohammad, et al.
Published: (2024)
by: Nazeri, Mohammad, et al.
Published: (2024)
Facing the Elephant in the Room: Visual Prompt Tuning or Full Finetuning?
by: Han, Cheng, et al.
Published: (2024)
by: Han, Cheng, et al.
Published: (2024)
Adaptive Visual Scene Understanding: Incremental Scene Graph Generation
by: Khandelwal, Naitik, et al.
Published: (2023)
by: Khandelwal, Naitik, et al.
Published: (2023)
Seeing the Whole in the Parts in Self-Supervised Representation Learning
by: Aubret, Arthur, et al.
Published: (2025)
by: Aubret, Arthur, et al.
Published: (2025)
Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization
by: Jia, Yanhao, et al.
Published: (2025)
by: Jia, Yanhao, et al.
Published: (2025)
Pose Prior Learner: Unsupervised Categorical Prior Learning for Pose Estimation
by: Wang, Ziyu, et al.
Published: (2024)
by: Wang, Ziyu, et al.
Published: (2024)
Explicit Context Reasoning with Supervision for Visual Tracking
by: Zeng, Fansheng, et al.
Published: (2025)
by: Zeng, Fansheng, et al.
Published: (2025)
L-WISE: Boosting Human Visual Category Learning Through Model-Based Image Selection and Enhancement
by: Talbot, Morgan B., et al.
Published: (2024)
by: Talbot, Morgan B., et al.
Published: (2024)
Unraveling the geometry of visual relational reasoning
by: Shang, Jiaqi, et al.
Published: (2025)
by: Shang, Jiaqi, et al.
Published: (2025)
HumorDB: Can AI understand graphical humor?
by: Jain, Vedaant, et al.
Published: (2024)
by: Jain, Vedaant, et al.
Published: (2024)
Human-like Object Grouping in Self-supervised Vision Transformers
by: Adeli, Hossein, et al.
Published: (2026)
by: Adeli, Hossein, et al.
Published: (2026)
See, Think, Learn: A Self-Taught Multimodal Reasoner
by: Sharma, Sourabh, et al.
Published: (2025)
by: Sharma, Sourabh, et al.
Published: (2025)
Seeing Through Uncertainty: A Free-Energy Approach for Real-Time Perceptual Adaptation in Robust Visual Navigation
by: Piriyajitakonkij, Maytus, et al.
Published: (2024)
by: Piriyajitakonkij, Maytus, et al.
Published: (2024)
Addressing the Elephant in the Room: Robust Animal Re-Identification with Unsupervised Part-Based Feature Alignment
by: Yu, Yingxue, et al.
Published: (2024)
by: Yu, Yingxue, et al.
Published: (2024)
Self-supervised Learning via Cluster Distance Prediction for Operating Room Context Awareness
by: Hamoud, Idris, et al.
Published: (2024)
by: Hamoud, Idris, et al.
Published: (2024)
Robust Representation Learning with Self-Distillation for Domain Generalization
by: Singh, Ankur, et al.
Published: (2023)
by: Singh, Ankur, et al.
Published: (2023)
Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room
by: Chen, Keqi, et al.
Published: (2026)
by: Chen, Keqi, et al.
Published: (2026)
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
by: Fuller, Anthony, et al.
Published: (2025)
by: Fuller, Anthony, et al.
Published: (2025)
Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion Perception
by: Han, Shuangpeng, et al.
Published: (2024)
by: Han, Shuangpeng, et al.
Published: (2024)
Transfer Learning and Explainable AI for Brain Tumor Classification: A Study Using MRI Data from Bangladesh
by: Sarker, Shuvashis
Published: (2025)
by: Sarker, Shuvashis
Published: (2025)
Self-Supervised Learning for Detecting AI-Generated Faces as Anomalies
by: Zou, Mian, et al.
Published: (2025)
by: Zou, Mian, et al.
Published: (2025)
Multi-View Crowd Counting With Self-Supervised Learning
by: Mo, Hong, et al.
Published: (2025)
by: Mo, Hong, et al.
Published: (2025)
Peering into the Unknown: Active View Selection with Neural Uncertainty Maps for 3D Reconstruction
by: Zhang, Zhengquan, et al.
Published: (2025)
by: Zhang, Zhengquan, et al.
Published: (2025)
Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual Foraging
by: Wang, Bo, et al.
Published: (2024)
by: Wang, Bo, et al.
Published: (2024)
Explorations in Self-Supervised Learning: Dataset Composition Testing for Object Classification
by: Chavez, Raynor Kirkson E., et al.
Published: (2024)
by: Chavez, Raynor Kirkson E., et al.
Published: (2024)
In-Context Symmetries: Self-Supervised Learning through Contextual World Models
by: Gupta, Sharut, et al.
Published: (2024)
by: Gupta, Sharut, et al.
Published: (2024)
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
Pyramid Pixel Context Adaption Network for Medical Image Classification with Supervised Contrastive Learning
by: Zhang, Xiaoqing, et al.
Published: (2023)
by: Zhang, Xiaoqing, et al.
Published: (2023)
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
by: Caffagni, Davide, et al.
Published: (2025)
by: Caffagni, Davide, et al.
Published: (2025)
Similar Items
-
Tuned Compositional Feature Replays for Efficient Stream Learning
by: Talbot, Morgan B., et al.
Published: (2021) -
Unveiling the Tapestry: the Interplay of Generalization and Forgetting in Continual Learning
by: Shi, Zenglin, et al.
Published: (2022) -
Improving generalization by mimicking the human visual diet
by: Madan, Spandan, et al.
Published: (2022) -
Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
by: Cai, Yusen, et al.
Published: (2025) -
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)