Dreaming Out Loud: A Self-Synthesis Approach For Training Vision-Language Models With Developmentally Plausible Data
Fuente:
arXiv
Saved in:
| Main Authors: | AlKhamissi, Badr, Tang, Yingtian, Gökce, Abdülkadir, Mehrer, Johannes, Schrimpf, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding
by: Gokce, Abdulkadir, et al.
Published: (2026)
by: Gokce, Abdulkadir, et al.
Published: (2026)
Inducing Dyslexia in Vision Language Models
by: Honarmand, Melika, et al.
Published: (2025)
by: Honarmand, Melika, et al.
Published: (2025)
From Language to Cognition: How LLMs Outgrow the Human Language Network
by: AlKhamissi, Badr, et al.
Published: (2025)
by: AlKhamissi, Badr, et al.
Published: (2025)
The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
TopoLM: brain-like spatio-functional organization in a topographic language model
by: Rathi, Neil, et al.
Published: (2024)
by: Rathi, Neil, et al.
Published: (2024)
Brain-Like Language Processing via a Shallow Untrained Multihead Attention Network
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models
by: Jamaa, Yassine, et al.
Published: (2025)
by: Jamaa, Yassine, et al.
Published: (2025)
Model-Guided Microstimulation Steers Primate Visual Behavior
by: Mehrer, Johannes, et al.
Published: (2025)
by: Mehrer, Johannes, et al.
Published: (2025)
Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream
by: Gokce, Abdulkadir, et al.
Published: (2024)
by: Gokce, Abdulkadir, et al.
Published: (2024)
Investigating Cultural Alignment of Large Language Models
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
A Context-Contrastive Inference Approach To Partial Diacritization
by: ElNokrashy, Muhammad, et al.
Published: (2024)
by: ElNokrashy, Muhammad, et al.
Published: (2024)
Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens
by: AlKhamissi, Mai, et al.
Published: (2025)
by: AlKhamissi, Mai, et al.
Published: (2025)
Instruction-tuning Aligns LLMs to the Human Brain
by: Aw, Khai Loong, et al.
Published: (2023)
by: Aw, Khai Loong, et al.
Published: (2023)
Depth-Wise Attention (DWAtt): A Layer Fusion Method for Data-Efficient Classification
by: ElNokrashy, Muhammad, et al.
Published: (2022)
by: ElNokrashy, Muhammad, et al.
Published: (2022)
Contour Integration Underlies Human-Like Vision
by: Lonnqvist, Ben, et al.
Published: (2025)
by: Lonnqvist, Ben, et al.
Published: (2025)
Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization
by: AlKhamissi, Badr, et al.
Published: (2025)
by: AlKhamissi, Badr, et al.
Published: (2025)
"Flex Tape Can't Fix That": Bias and Misinformation in Edited Language Models
by: Halevy, Karina, et al.
Published: (2024)
by: Halevy, Karina, et al.
Published: (2024)
Rosetta Stone at KSAA-RD Shared Task: A Hop From Language Modeling To Word--Definition Alignment
by: ElBakry, Ahmed, et al.
Published: (2023)
by: ElBakry, Ahmed, et al.
Published: (2023)
Large Language Models Align with the Human Brain during Creative Thinking
by: Ismayilzada, Mete, et al.
Published: (2026)
by: Ismayilzada, Mete, et al.
Published: (2026)
Rational Metareasoning for Large Language Models
by: De Sabbata, C. Nicolò, et al.
Published: (2024)
by: De Sabbata, C. Nicolò, et al.
Published: (2024)
Khattat: Enhancing Readability and Concept Representation of Semantic Typography
by: Hussein, Ahmed, et al.
Published: (2024)
by: Hussein, Ahmed, et al.
Published: (2024)
When are Lemons Purple? The Concept Association Bias of Vision-Language Models
by: Yamada, Yutaro, et al.
Published: (2022)
by: Yamada, Yutaro, et al.
Published: (2022)
Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models
by: Takmaz, Ece, et al.
Published: (2025)
by: Takmaz, Ece, et al.
Published: (2025)
MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
by: Wang, Shuo, et al.
Published: (2025)
by: Wang, Shuo, et al.
Published: (2025)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
by: Ye, Jiacheng, et al.
Published: (2025)
by: Ye, Jiacheng, et al.
Published: (2025)
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
by: Inal, Gokce, et al.
Published: (2026)
by: Inal, Gokce, et al.
Published: (2026)
Backdooring Vision-Language Models with Out-Of-Distribution Data
by: Lyu, Weimin, et al.
Published: (2024)
by: Lyu, Weimin, et al.
Published: (2024)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
Text-Only Data Synthesis for Vision Language Model Training
by: Yu, Xiaomin, et al.
Published: (2025)
by: Yu, Xiaomin, et al.
Published: (2025)
From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models
by: Gong, Weile, et al.
Published: (2026)
by: Gong, Weile, et al.
Published: (2026)
An Efficient Self-supervised Seismic Data Reconstruction Method Based on Self-Consistency Learning
by: Wang, Mingwei, et al.
Published: (2024)
by: Wang, Mingwei, et al.
Published: (2024)
Multisensor Data Fusion for Automatized Insect Monitoring (KInsecta)
by: Tschaikner, Martin, et al.
Published: (2024)
by: Tschaikner, Martin, et al.
Published: (2024)
A Self-Supervised Approach on Motion Calibration for Enhancing Physical Plausibility in Text-to-Motion
by: Shim, Gahyeon, et al.
Published: (2026)
by: Shim, Gahyeon, et al.
Published: (2026)
VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
by: Yang, Xindi, et al.
Published: (2025)
by: Yang, Xindi, et al.
Published: (2025)
Recite Your Ask Out Loud
Published: (2025)
Published: (2025)
Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection
by: Yu, Geng, et al.
Published: (2024)
by: Yu, Geng, et al.
Published: (2024)
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
by: Jumelet, Jaap, et al.
Published: (2025)
by: Jumelet, Jaap, et al.
Published: (2025)
Dream-Box: Object-wise Outlier Generation for Out-of-Distribution Detection
by: Isaac-Medina, Brian K. S., et al.
Published: (2025)
by: Isaac-Medina, Brian K. S., et al.
Published: (2025)
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2026)
by: Ye, Junjie, et al.
Published: (2026)
Low Cost Machine Vision for Insect Classification
by: Brandt, Danja, et al.
Published: (2024)
by: Brandt, Danja, et al.
Published: (2024)
Similar Items
-
MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding
by: Gokce, Abdulkadir, et al.
Published: (2026) -
Inducing Dyslexia in Vision Language Models
by: Honarmand, Melika, et al.
Published: (2025) -
From Language to Cognition: How LLMs Outgrow the Human Language Network
by: AlKhamissi, Badr, et al.
Published: (2025) -
The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
by: AlKhamissi, Badr, et al.
Published: (2024) -
TopoLM: brain-like spatio-functional organization in a topographic language model
by: Rathi, Neil, et al.
Published: (2024)