BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shengao, Wang, Wenqi, Wang, Zecheng, Whitton, Max, Wakeham, Michael, Chandra, Arjun, Huang, Joey, Zhu, Pengyue, Chen, Helen, Li, David, Li, Jeffrey, Li, Shawn, Zagula, Andrew, Zhao, Amy, Zhu, Andrew, Nakamura, Sayaka, Yamamoto, Yuki, Yokono, Jerry Jun, Mueller, Aaron, Plummer, Bryan A., Saenko, Kate, Saligrama, Venkatesh, Gong, Boqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
by: Wang, Shengao, et al.
Published: (2025)
by: Wang, Shengao, et al.
Published: (2025)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
by: Liu, Aoming, et al.
Published: (2025)
by: Liu, Aoming, et al.
Published: (2025)
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
by: Mishra, Samarth, et al.
Published: (2025)
by: Mishra, Samarth, et al.
Published: (2025)
Deep Companion Learning: Enhancing Generalization Through Historical Consistency
by: Zhu, Ruizhao, et al.
Published: (2024)
by: Zhu, Ruizhao, et al.
Published: (2024)
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
by: Mishra, Samarth, et al.
Published: (2023)
by: Mishra, Samarth, et al.
Published: (2023)
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
by: Miller, Kevin, et al.
Published: (2025)
by: Miller, Kevin, et al.
Published: (2025)
TRACE: AI-Assisted Assessment of Collaborative Projects in Computer Science Education
by: Yu, Songmei, et al.
Published: (2025)
by: Yu, Songmei, et al.
Published: (2025)
Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation
by: Chandra, Arjun, et al.
Published: (2026)
by: Chandra, Arjun, et al.
Published: (2026)
Constrained Linear Thompson Sampling
by: Gangrade, Aditya, et al.
Published: (2025)
by: Gangrade, Aditya, et al.
Published: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025)
by: Reddy, Aashray, et al.
Published: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025)
by: Reddy, Aashray, et al.
Published: (2025)
SITE: towards Spatial Intelligence Thorough Evaluation
by: Wang, Wenqi, et al.
Published: (2025)
by: Wang, Wenqi, et al.
Published: (2025)
Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification
by: Lutz, Patrick, et al.
Published: (2026)
by: Lutz, Patrick, et al.
Published: (2026)
Safe Linear Bandits over Unknown Polytopes
by: Gangrade, Aditya, et al.
Published: (2022)
by: Gangrade, Aditya, et al.
Published: (2022)
Breaking the Assistant Mold: Modeling Behavioral Variation in LLM Based Procedural Character Generation
by: Qraitem, Maan, et al.
Published: (2026)
by: Qraitem, Maan, et al.
Published: (2026)
Tell Me What's Next: Textual Foresight for Generic UI Representations
by: Burns, Andrea, et al.
Published: (2024)
by: Burns, Andrea, et al.
Published: (2024)
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
by: Qraitem, Maan, et al.
Published: (2023)
by: Qraitem, Maan, et al.
Published: (2023)
Navigation with VLM framework: Towards Going to Any Language
by: Yin, Zecheng, et al.
Published: (2024)
by: Yin, Zecheng, et al.
Published: (2024)
Data Deletion Can Help in Adaptive RL
by: Budhraja, Param, et al.
Published: (2026)
by: Budhraja, Param, et al.
Published: (2026)
Testing the Feasibility of Linear Programs with Bandit Feedback
by: Gangrade, Aditya, et al.
Published: (2024)
by: Gangrade, Aditya, et al.
Published: (2024)
Linear Transformers Implicitly Discover Unified Numerical Algorithms
by: Lutz, Patrick, et al.
Published: (2025)
by: Lutz, Patrick, et al.
Published: (2025)
SLANT: Spurious Logo ANalysis Toolkit
by: Qraitem, Maan, et al.
Published: (2024)
by: Qraitem, Maan, et al.
Published: (2024)
Web Artifact Attacks Disrupt Vision Language Models
by: Qraitem, Maan, et al.
Published: (2025)
by: Qraitem, Maan, et al.
Published: (2025)
Time-Dependent Hamiltonian Simulation via Time-Independent Dynamics in a Larger Space
by: Li, Zecheng, et al.
Published: (2025)
by: Li, Zecheng, et al.
Published: (2025)
Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs
by: Liu, Aoming, et al.
Published: (2025)
by: Liu, Aoming, et al.
Published: (2025)
AVA: Attentive VLM Agent for Mastering StarCraft II
by: Ma, Weiyu, et al.
Published: (2025)
by: Ma, Weiyu, et al.
Published: (2025)
OP-LoRA: The Blessing of Dimensionality
by: Teterwak, Piotr, et al.
Published: (2024)
by: Teterwak, Piotr, et al.
Published: (2024)
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
by: Dhole, Kaustubh D.
Published: (2026)
by: Dhole, Kaustubh D.
Published: (2026)
Planning for Cooler Cities: A Multimodal AI Framework for Predicting and Mitigating Urban Heat Stress through Urban Landscape Transformation
by: Yi, Shengao, et al.
Published: (2025)
by: Yi, Shengao, et al.
Published: (2025)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
by: Teterwak, Piotr, et al.
Published: (2024)
by: Teterwak, Piotr, et al.
Published: (2024)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
by: Qraitem, Maan, et al.
Published: (2024)
by: Qraitem, Maan, et al.
Published: (2024)
ERM++: An Improved Baseline for Domain Generalization
by: Teterwak, Piotr, et al.
Published: (2023)
by: Teterwak, Piotr, et al.
Published: (2023)
Baby Universe in a Coupled SYK Model
by: Sontag, Andrew, et al.
Published: (2026)
by: Sontag, Andrew, et al.
Published: (2026)
Polyoxometalate‐Based MOF as a Highly Sensitive and Stable X‐ray Detector for Imaging
by: Yanli Yang, et al.
Published: (2025)
by: Yanli Yang, et al.
Published: (2025)
CLAMP: Contrastive LAnguage Model Prompt-tuning
by: Teterwak, Piotr, et al.
Published: (2023)
by: Teterwak, Piotr, et al.
Published: (2023)
Recover Cell Tensor: Diffusion-Equivalent Tensor Completion for Fluorescence Microscopy Imaging
by: Wang, Chenwei, et al.
Published: (2026)
by: Wang, Chenwei, et al.
Published: (2026)
Plastic inorganic Sn2BiS2I3 semiconductor enabled deformable and flexible electronic tongue for heavy metal detection
by: Wang, Qiao, et al.
Published: (2025)
by: Wang, Qiao, et al.
Published: (2025)
An Information-Theoretic Method for Dynamic System Identification With Output-Only Damping Estimation
by: Impraimakis, Marios, et al.
Published: (2026)
by: Impraimakis, Marios, et al.
Published: (2026)
Ac34deoGlcNAz: A Selective Probe for Identifying O‐GlcNAc‐Modified Proteins in Prostate Cancer
by: Guoliang Xu, et al.
Published: (2024)
by: Guoliang Xu, et al.
Published: (2024)
Control Synthesis of Multi‐Equilibrium Asynchronously Switched Systems Under Control Constraints
by: Yuejiang Han, et al.
Published: (2025)
by: Yuejiang Han, et al.
Published: (2025)
Similar Items
-
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
by: Wang, Shengao, et al.
Published: (2025) -
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
by: Liu, Aoming, et al.
Published: (2025) -
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
by: Mishra, Samarth, et al.
Published: (2025) -
Deep Companion Learning: Enhancing Generalization Through Historical Consistency
by: Zhu, Ruizhao, et al.
Published: (2024) -
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
by: Mishra, Samarth, et al.
Published: (2023)