BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Shengao, Chandra, Arjun, Liu, Aoming, Saligrama, Venkatesh, Gong, Boqing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
por: Wang, Shengao, et al.
Publicado: (2025)
por: Wang, Shengao, et al.
Publicado: (2025)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
por: Liu, Aoming, et al.
Publicado: (2025)
por: Liu, Aoming, et al.
Publicado: (2025)
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
por: Mishra, Samarth, et al.
Publicado: (2025)
por: Mishra, Samarth, et al.
Publicado: (2025)
EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
por: Lin, Dongyan, et al.
Publicado: (2026)
por: Lin, Dongyan, et al.
Publicado: (2026)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
por: Tan, Yuwen, et al.
Publicado: (2025)
por: Tan, Yuwen, et al.
Publicado: (2025)
Deep Companion Learning: Enhancing Generalization Through Historical Consistency
por: Zhu, Ruizhao, et al.
Publicado: (2024)
por: Zhu, Ruizhao, et al.
Publicado: (2024)
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
por: Mishra, Samarth, et al.
Publicado: (2023)
por: Mishra, Samarth, et al.
Publicado: (2023)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
por: Lokesh, K, et al.
Publicado: (2026)
por: Lokesh, K, et al.
Publicado: (2026)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
por: Liu, Peng, et al.
Publicado: (2025)
por: Liu, Peng, et al.
Publicado: (2025)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
por: Lu, Meng, et al.
Publicado: (2025)
por: Lu, Meng, et al.
Publicado: (2025)
On Discrete Prompt Optimization for Diffusion Models
por: Wang, Ruochen, et al.
Publicado: (2024)
por: Wang, Ruochen, et al.
Publicado: (2024)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
por: Zhou, Guanyu, et al.
Publicado: (2026)
por: Zhou, Guanyu, et al.
Publicado: (2026)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
por: Yu, Shoubin, et al.
Publicado: (2026)
por: Yu, Shoubin, et al.
Publicado: (2026)
Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs
por: Hong, Weihao, et al.
Publicado: (2026)
por: Hong, Weihao, et al.
Publicado: (2026)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
por: Li, Shuo, et al.
Publicado: (2024)
por: Li, Shuo, et al.
Publicado: (2024)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
por: Mayer, Julius, et al.
Publicado: (2025)
por: Mayer, Julius, et al.
Publicado: (2025)
Can VLMs Recall Factual Associations From Visual References?
por: Ashok, Dhananjay, et al.
Publicado: (2025)
por: Ashok, Dhananjay, et al.
Publicado: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
por: Wang, Dianyi, et al.
Publicado: (2025)
por: Wang, Dianyi, et al.
Publicado: (2025)
BabyVision: Visual Reasoning Beyond Language
por: Chen, Liang, et al.
Publicado: (2026)
por: Chen, Liang, et al.
Publicado: (2026)
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
por: Zhan, Qiusi, et al.
Publicado: (2025)
por: Zhan, Qiusi, et al.
Publicado: (2025)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
por: Avogaro, Niccolo, et al.
Publicado: (2026)
por: Avogaro, Niccolo, et al.
Publicado: (2026)
Navigation with VLM framework: Towards Going to Any Language
por: Yin, Zecheng, et al.
Publicado: (2024)
por: Yin, Zecheng, et al.
Publicado: (2024)
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
por: Penamakuri, Abhirama Subramanyam, et al.
Publicado: (2025)
por: Penamakuri, Abhirama Subramanyam, et al.
Publicado: (2025)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage
por: Khan, Ufaq, et al.
Publicado: (2026)
por: Khan, Ufaq, et al.
Publicado: (2026)
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
por: Huang, Yihong, et al.
Publicado: (2026)
por: Huang, Yihong, et al.
Publicado: (2026)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
por: Xu, Ruyi, et al.
Publicado: (2025)
por: Xu, Ruyi, et al.
Publicado: (2025)
FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
por: Xian, Ruiqi, et al.
Publicado: (2024)
por: Xian, Ruiqi, et al.
Publicado: (2024)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
por: Zhang, Yue, et al.
Publicado: (2026)
por: Zhang, Yue, et al.
Publicado: (2026)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
por: Oh, Youngtaek, et al.
Publicado: (2024)
por: Oh, Youngtaek, et al.
Publicado: (2024)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
por: Liu, Tianhui, et al.
Publicado: (2026)
por: Liu, Tianhui, et al.
Publicado: (2026)
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
por: Pan, Jiazhen, et al.
Publicado: (2025)
por: Pan, Jiazhen, et al.
Publicado: (2025)
CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries
por: Liu, Shudong, et al.
Publicado: (2025)
por: Liu, Shudong, et al.
Publicado: (2025)
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
por: Singh, Anshul, et al.
Publicado: (2025)
por: Singh, Anshul, et al.
Publicado: (2025)
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
por: Pang, Zirui, et al.
Publicado: (2025)
por: Pang, Zirui, et al.
Publicado: (2025)
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
por: Ebouky, Brown, et al.
Publicado: (2026)
por: Ebouky, Brown, et al.
Publicado: (2026)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
por: Yu, Seonghoon, et al.
Publicado: (2026)
por: Yu, Seonghoon, et al.
Publicado: (2026)
Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
por: Omasa, Takamitsu, et al.
Publicado: (2025)
por: Omasa, Takamitsu, et al.
Publicado: (2025)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
por: Singh, Ayush, et al.
Publicado: (2024)
por: Singh, Ayush, et al.
Publicado: (2024)
VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT
por: Kang, Hyeonsu, et al.
Publicado: (2025)
por: Kang, Hyeonsu, et al.
Publicado: (2025)
Ejemplares similares
-
BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
por: Wang, Shengao, et al.
Publicado: (2025) -
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
por: Liu, Aoming, et al.
Publicado: (2025) -
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
por: Mishra, Samarth, et al.
Publicado: (2025) -
EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
por: Lin, Dongyan, et al.
Publicado: (2026) -
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
por: Tan, Yuwen, et al.
Publicado: (2025)