Assessing the alignment between infants' visual and linguistic experience using multimodal language models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Alvin Wei Ming, Yang, Jane, Sepuri, Tarun, Aw, Khai Loong, Sparks, Robert Z., Yin, Zi, Marchman, Virginia A., Frank, Michael C., Long, Bria |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Characterizing the visual representation of objects from the child's view
von: Yang, Jane, et al.
Veröffentlicht: (2026)
von: Yang, Jane, et al.
Veröffentlicht: (2026)
The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences
von: Long, Bria, et al.
Veröffentlicht: (2024)
von: Long, Bria, et al.
Veröffentlicht: (2024)
DevBench: A multimodal developmental benchmark for language learning
von: Tan, Alvin Wei Ming, et al.
Veröffentlicht: (2024)
von: Tan, Alvin Wei Ming, et al.
Veröffentlicht: (2024)
Instruction-tuning Aligns LLMs to the Human Brain
von: Aw, Khai Loong, et al.
Veröffentlicht: (2023)
von: Aw, Khai Loong, et al.
Veröffentlicht: (2023)
Unified 3D Scene Understanding Through Physical World Modeling
von: Lee, Wanhee, et al.
Veröffentlicht: (2026)
von: Lee, Wanhee, et al.
Veröffentlicht: (2026)
3D Scene Understanding Through Local Random Access Sequence Modeling
von: Lee, Wanhee, et al.
Veröffentlicht: (2025)
von: Lee, Wanhee, et al.
Veröffentlicht: (2025)
Protecting multimodal large language models against misleading visualizations
von: Tonglet, Jonathan, et al.
Veröffentlicht: (2025)
von: Tonglet, Jonathan, et al.
Veröffentlicht: (2025)
Context informs pragmatic interpretation in vision-language models
von: Tan, Alvin Wei Ming, et al.
Veröffentlicht: (2025)
von: Tan, Alvin Wei Ming, et al.
Veröffentlicht: (2025)
What to align in multimodal contrastive learning?
von: Dufumier, Benoit, et al.
Veröffentlicht: (2024)
von: Dufumier, Benoit, et al.
Veröffentlicht: (2024)
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
von: Chen, Yanyuan, et al.
Veröffentlicht: (2025)
von: Chen, Yanyuan, et al.
Veröffentlicht: (2025)
The role of translation equivalents in bilingual word learning
von: Alvin W. M. Tan, et al.
Veröffentlicht: (2024)
von: Alvin W. M. Tan, et al.
Veröffentlicht: (2024)
On the robustness of multimodal language model towards distractions
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment
von: Wu, Hanlin, et al.
Veröffentlicht: (2025)
von: Wu, Hanlin, et al.
Veröffentlicht: (2025)
Is your multimodal large language model a good science tutor?
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
Large language models and linguistic intentionality
von: Grindrod, Jumbly
Veröffentlicht: (2024)
von: Grindrod, Jumbly
Veröffentlicht: (2024)
How desirable is alignment between LLMs and linguistically diverse human users?
von: Knoeferle, Pia, et al.
Veröffentlicht: (2025)
von: Knoeferle, Pia, et al.
Veröffentlicht: (2025)
A conclusive remark on linguistic theorizing and language modeling
von: Chesi, Cristiano
Veröffentlicht: (2025)
von: Chesi, Cristiano
Veröffentlicht: (2025)
Linguini: A benchmark for language-agnostic linguistic reasoning
von: Sánchez, Eduardo, et al.
Veröffentlicht: (2024)
von: Sánchez, Eduardo, et al.
Veröffentlicht: (2024)
A blind spot for large language models: Supradiegetic linguistic information
von: Zimmerman, Julia Witte, et al.
Veröffentlicht: (2023)
von: Zimmerman, Julia Witte, et al.
Veröffentlicht: (2023)
Zero-shot World Models Are Developmentally Efficient Learners
von: Aw, Khai Loong, et al.
Veröffentlicht: (2026)
von: Aw, Khai Loong, et al.
Veröffentlicht: (2026)
Evaluating Polish linguistic and cultural competency in large language models
von: Dadas, Sławomir, et al.
Veröffentlicht: (2025)
von: Dadas, Sławomir, et al.
Veröffentlicht: (2025)
Deaf in AI: AI language technologies and the erosion of linguistic rights
von: De Meulder, Maartje
Veröffentlicht: (2025)
von: De Meulder, Maartje
Veröffentlicht: (2025)
VaPR -- Vision-language Preference alignment for Reasoning
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2025)
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2025)
Closing the gap in multimodal medical representation alignment
von: Grassucci, Eleonora, et al.
Veröffentlicht: (2026)
von: Grassucci, Eleonora, et al.
Veröffentlicht: (2026)
Taming generative video models for zero-shot optical flow extraction
von: Kim, Seungwoo, et al.
Veröffentlicht: (2025)
von: Kim, Seungwoo, et al.
Veröffentlicht: (2025)
Are formal and functional linguistic mechanisms dissociated in language models?
von: Hanna, Michael, et al.
Veröffentlicht: (2025)
von: Hanna, Michael, et al.
Veröffentlicht: (2025)
Do language models accommodate their users? A study of linguistic convergence
von: Blevins, Terra, et al.
Veröffentlicht: (2025)
von: Blevins, Terra, et al.
Veröffentlicht: (2025)
Vocabulary embeddings organize linguistic structure early in language model training
von: Papadimitriou, Isabel, et al.
Veröffentlicht: (2025)
von: Papadimitriou, Isabel, et al.
Veröffentlicht: (2025)
RoMemes: A multimodal meme corpus for the Romanian language
von: Păiş, Vasile, et al.
Veröffentlicht: (2024)
von: Păiş, Vasile, et al.
Veröffentlicht: (2024)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
von: Matos, João, et al.
Veröffentlicht: (2024)
von: Matos, João, et al.
Veröffentlicht: (2024)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
Can multimodal representation learning by alignment preserve modality-specific information?
von: Thoreau, Romain, et al.
Veröffentlicht: (2025)
von: Thoreau, Romain, et al.
Veröffentlicht: (2025)
A Glass Half Full: Limitations in ChiLDES Point to Ways Forward for a More Representative Developmental Science. Commentary on Scaff et al. (2025)
von: Virginia A. Marchman, et al.
Veröffentlicht: (2025)
von: Virginia A. Marchman, et al.
Veröffentlicht: (2025)
Gravity Network for end-to-end small lesion detection
von: Russo, Ciro, et al.
Veröffentlicht: (2023)
von: Russo, Ciro, et al.
Veröffentlicht: (2023)
A solution to generalized learning from small training sets found in infant repeated visual experiences of individual objects
von: Ramirez, Frangil, et al.
Veröffentlicht: (2025)
von: Ramirez, Frangil, et al.
Veröffentlicht: (2025)
Phonological distances for linguistic typology and the origin of Indo-European languages
von: Mavridis, Marius, et al.
Veröffentlicht: (2026)
von: Mavridis, Marius, et al.
Veröffentlicht: (2026)
UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
von: Bhattacharya, Antara Raaghavi, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Antara Raaghavi, et al.
Veröffentlicht: (2025)
Human-like object concept representations emerge naturally in multimodal large language models
von: Du, Changde, et al.
Veröffentlicht: (2024)
von: Du, Changde, et al.
Veröffentlicht: (2024)
Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
von: Jacobs, Cassandra L., et al.
Veröffentlicht: (2024)
von: Jacobs, Cassandra L., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Characterizing the visual representation of objects from the child's view
von: Yang, Jane, et al.
Veröffentlicht: (2026) -
The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences
von: Long, Bria, et al.
Veröffentlicht: (2024) -
DevBench: A multimodal developmental benchmark for language learning
von: Tan, Alvin Wei Ming, et al.
Veröffentlicht: (2024) -
Instruction-tuning Aligns LLMs to the Human Brain
von: Aw, Khai Loong, et al.
Veröffentlicht: (2023) -
Unified 3D Scene Understanding Through Physical World Modeling
von: Lee, Wanhee, et al.
Veröffentlicht: (2026)