A solution to generalized learning from small training sets found in infant repeated visual experiences of individual objects
Fuente:
arXiv
Saved in:
| Main Authors: | Ramirez, Frangil, Clerkin, Elizabeth, Crandall, David J., Smith, Linda B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
EgoVIS@CVPR: What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
by: Kung, Chi-Hsi, et al.
Published: (2025)
by: Kung, Chi-Hsi, et al.
Published: (2025)
What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
by: Kung, Chi-Hsi, et al.
Published: (2025)
by: Kung, Chi-Hsi, et al.
Published: (2025)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
EgoVIS@CVPR: PAIR-Net: Enhancing Egocentric Speaker Detection via Pretrained Audio-Visual Fusion and Alignment Loss
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
An analysis of HOI: using a training-free method with multimodal visual foundation models when only the test set is available, without the training set
by: Ai, Chaoyi
Published: (2024)
by: Ai, Chaoyi
Published: (2024)
LoCoNet: Long-Short Context Network for Active Speaker Detection
by: Wang, Xizi, et al.
Published: (2023)
by: Wang, Xizi, et al.
Published: (2023)
Geometric Constraints in Deep Learning Frameworks: A Survey
by: Vats, Vibhas K, et al.
Published: (2024)
by: Vats, Vibhas K, et al.
Published: (2024)
Deep learning-based neurodevelopmental assessment in preterm infants
by: Ren, Lexin, et al.
Published: (2026)
by: Ren, Lexin, et al.
Published: (2026)
Do ImageNet-trained models learn shortcuts? The impact of frequency shortcuts on generalization
by: Wang, Shunxin, et al.
Published: (2025)
by: Wang, Shunxin, et al.
Published: (2025)
Affine transformation estimation improves visual self-supervised learning
by: Torpey, David, et al.
Published: (2024)
by: Torpey, David, et al.
Published: (2024)
Less is more: concatenating videos for Sign Language Translation from a small set of signs
by: da Silva, David Vinicius, et al.
Published: (2024)
by: da Silva, David Vinicius, et al.
Published: (2024)
Case-Enhanced Vision Transformer: Improving Explanations of Image Similarity with a ViT-based Similarity Metric
by: Zhao, Ziwei, et al.
Published: (2024)
by: Zhao, Ziwei, et al.
Published: (2024)
Sequence-Based Identification of First-Person Camera Wearers in Third-Person Views
by: Zhao, Ziwei, et al.
Published: (2025)
by: Zhao, Ziwei, et al.
Published: (2025)
Constrained Transformer-Based Porous Media Generation to Spatial Distribution of Rock Properties
by: Ren, Zihan, et al.
Published: (2024)
by: Ren, Zihan, et al.
Published: (2024)
Deep learning based infrared small object segmentation: Challenges and future directions
by: Yang, Zhengeng, et al.
Published: (2025)
by: Yang, Zhengeng, et al.
Published: (2025)
Self-supervised visual learning from interactions with objects
by: Aubret, Arthur, et al.
Published: (2024)
by: Aubret, Arthur, et al.
Published: (2024)
Degradation-based augmented training for robust individual animal re-identification
by: Polychronou, Thanos, et al.
Published: (2026)
by: Polychronou, Thanos, et al.
Published: (2026)
The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences
by: Long, Bria, et al.
Published: (2024)
by: Long, Bria, et al.
Published: (2024)
Weakly supervised training of universal visual concepts for multi-domain semantic segmentation
by: Bevandić, Petra, et al.
Published: (2022)
by: Bevandić, Petra, et al.
Published: (2022)
DeepSet SimCLR: Self-supervised deep sets for improved pathology representation learning
by: Torpey, David, et al.
Published: (2024)
by: Torpey, David, et al.
Published: (2024)
GAQAT: gradient-adaptive quantization-aware training for domain generalization
by: Jiang, Jiacheng, et al.
Published: (2024)
by: Jiang, Jiacheng, et al.
Published: (2024)
Do text-free diffusion models learn discriminative visual representations?
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
Bridging the gap to real-world language-grounded visual concept learning
by: Jung, Whie, et al.
Published: (2025)
by: Jung, Whie, et al.
Published: (2025)
Dense outlier detection and open-set recognition based on training with noisy negative images
by: Bevandić, Petra, et al.
Published: (2021)
by: Bevandić, Petra, et al.
Published: (2021)
HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
A Markovian View of Iterative-Feedback Loops in Image Generative Models: Neural Resonance and Model Collapse
by: Vats, Vibhas Kumar, et al.
Published: (2026)
by: Vats, Vibhas Kumar, et al.
Published: (2026)
Self-supervised visual learning in the low-data regime: a comparative evaluation
by: Konstantakos, Sotirios, et al.
Published: (2024)
by: Konstantakos, Sotirios, et al.
Published: (2024)
Human-like compositional learning of visually-grounded concepts using synthetic environments
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
An extremely coarse feedback signal is sufficient for learning human-aligned visual representations
by: Mehta, Yash, et al.
Published: (2026)
by: Mehta, Yash, et al.
Published: (2026)
Shadow loss: Memory-linear deep metric learning for efficient training
by: Khan, Alif Elham, et al.
Published: (2023)
by: Khan, Alif Elham, et al.
Published: (2023)
Characterizing the visual representation of objects from the child's view
by: Yang, Jane, et al.
Published: (2026)
by: Yang, Jane, et al.
Published: (2026)
Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning
by: Kolner, Oleh, et al.
Published: (2024)
by: Kolner, Oleh, et al.
Published: (2024)
Attention is All They Need: Exploring the Media Archaeology of the Computer Vision Research Paper
by: Goree, Samuel, et al.
Published: (2022)
by: Goree, Samuel, et al.
Published: (2022)
Audio-visual training for improved grounding in video-text LLMs
by: Sagare, Shivprasad, et al.
Published: (2024)
by: Sagare, Shivprasad, et al.
Published: (2024)
Markerless retro-identification complements re-identification of individual insect subjects in archived image data of biological experiments
by: Zaman, Asaduz, et al.
Published: (2024)
by: Zaman, Asaduz, et al.
Published: (2024)
Performance evaluation of deep learning models for image analysis: considerations for visual control and statistical metrics
by: Bertram, Christof A., et al.
Published: (2026)
by: Bertram, Christof A., et al.
Published: (2026)
Do computer vision foundation models learn the low-level characteristics of the human visual system?
by: Cai, Yancheng, et al.
Published: (2025)
by: Cai, Yancheng, et al.
Published: (2025)
Improving generalization by mimicking the human visual diet
by: Madan, Spandan, et al.
Published: (2022)
by: Madan, Spandan, et al.
Published: (2022)
Multi-resolution Guided 3D GANs for Medical Image Translation
by: Ha, Juhyung, et al.
Published: (2024)
by: Ha, Juhyung, et al.
Published: (2024)
Similar Items
-
GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection
by: Wang, Yu, et al.
Published: (2025) -
EgoVIS@CVPR: What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
by: Kung, Chi-Hsi, et al.
Published: (2025) -
What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning
by: Kung, Chi-Hsi, et al.
Published: (2025) -
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025) -
EgoVIS@CVPR: PAIR-Net: Enhancing Egocentric Speaker Detection via Pretrained Audio-Visual Fusion and Alignment Loss
by: Wang, Yu, et al.
Published: (2025)