KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Yiu, Eunice, Qraitem, Maan, Majhi, Anisa Noor, Wong, Charlie, Bai, Yutong, Ginosar, Shiry, Gopnik, Alison, Saenko, Kate |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions
por: Yiu, Eunice, et al.
Publicado: (2025)
por: Yiu, Eunice, et al.
Publicado: (2025)
Breaking the Assistant Mold: Modeling Behavioral Variation in LLM Based Procedural Character Generation
por: Qraitem, Maan, et al.
Publicado: (2026)
por: Qraitem, Maan, et al.
Publicado: (2026)
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
por: Qraitem, Maan, et al.
Publicado: (2023)
por: Qraitem, Maan, et al.
Publicado: (2023)
SLANT: Spurious Logo ANalysis Toolkit
por: Qraitem, Maan, et al.
Publicado: (2024)
por: Qraitem, Maan, et al.
Publicado: (2024)
Web Artifact Attacks Disrupt Vision Language Models
por: Qraitem, Maan, et al.
Publicado: (2025)
por: Qraitem, Maan, et al.
Publicado: (2025)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
por: Qraitem, Maan, et al.
Publicado: (2024)
por: Qraitem, Maan, et al.
Publicado: (2024)
Poly-Autoregressive Prediction for Modeling Interactions
por: Thakkar, Neerja, et al.
Publicado: (2025)
por: Thakkar, Neerja, et al.
Publicado: (2025)
Pose Priors from Language Models
por: Subramanian, Sanjay, et al.
Publicado: (2024)
por: Subramanian, Sanjay, et al.
Publicado: (2024)
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
por: Koepke, A. Sophia, et al.
Publicado: (2026)
por: Koepke, A. Sophia, et al.
Publicado: (2026)
Diffusion Models as Data Mining Tools
por: Siglidis, Ioannis, et al.
Publicado: (2024)
por: Siglidis, Ioannis, et al.
Publicado: (2024)
Forecasting Motion in the Wild
por: Thakkar, Neerja, et al.
Publicado: (2026)
por: Thakkar, Neerja, et al.
Publicado: (2026)
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
por: Mishra, Samarth, et al.
Publicado: (2025)
por: Mishra, Samarth, et al.
Publicado: (2025)
Gaussian Masked Autoencoders
por: Rajasegaran, Jathushan, et al.
Publicado: (2025)
por: Rajasegaran, Jathushan, et al.
Publicado: (2025)
Synergy and Synchrony in Couple Dances
por: Maluleke, Vongani, et al.
Publicado: (2024)
por: Maluleke, Vongani, et al.
Publicado: (2024)
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
por: Petsiuk, Vitali, et al.
Publicado: (2024)
por: Petsiuk, Vitali, et al.
Publicado: (2024)
Tell Me What's Next: Textual Foresight for Generic UI Representations
por: Burns, Andrea, et al.
Publicado: (2024)
por: Burns, Andrea, et al.
Publicado: (2024)
Children's Mental Models of Generative Visual and Text Based AI Models
por: Kosoy, Eliza, et al.
Publicado: (2024)
por: Kosoy, Eliza, et al.
Publicado: (2024)
KiC: Keyword-inspired Cascade for Cost-Efficient Text Generation with LLMs
por: Kim, Woo-Chan, et al.
Publicado: (2025)
por: Kim, Woo-Chan, et al.
Publicado: (2025)
Federated Adversarial Domain Adaptation
por: Peng, Xingchao, et al.
Publicado: (2019)
por: Peng, Xingchao, et al.
Publicado: (2019)
SingaKids: A Multilingual Multimodal Dialogic Tutor for Language Learning
por: Liu, Zhengyuan, et al.
Publicado: (2025)
por: Liu, Zhengyuan, et al.
Publicado: (2025)
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
por: Miller, Kevin, et al.
Publicado: (2025)
por: Miller, Kevin, et al.
Publicado: (2025)
Enhancing immersion in Virtual Reality sports through Physical Interactions
por: Majhi, Arka
Publicado: (2026)
por: Majhi, Arka
Publicado: (2026)
OP-LoRA: The Blessing of Dimensionality
por: Teterwak, Piotr, et al.
Publicado: (2024)
por: Teterwak, Piotr, et al.
Publicado: (2024)
Frozen Forecasting: A Unified Evaluation
por: Walker, Jacob C, et al.
Publicado: (2025)
por: Walker, Jacob C, et al.
Publicado: (2025)
CLAMP: Contrastive LAnguage Model Prompt-tuning
por: Teterwak, Piotr, et al.
Publicado: (2023)
por: Teterwak, Piotr, et al.
Publicado: (2023)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
por: Teterwak, Piotr, et al.
Publicado: (2024)
por: Teterwak, Piotr, et al.
Publicado: (2024)
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
por: Mishra, Samarth, et al.
Publicado: (2023)
por: Mishra, Samarth, et al.
Publicado: (2023)
ERM++: An Improved Baseline for Domain Generalization
por: Teterwak, Piotr, et al.
Publicado: (2023)
por: Teterwak, Piotr, et al.
Publicado: (2023)
An Author in Every Classroom: Kids Connecting with Authors via Skype. It's the next Best Thing to Being There
por: Messner, Kate
Publicado: (2010)
por: Messner, Kate
Publicado: (2010)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
por: Liu, Aoming, et al.
Publicado: (2025)
por: Liu, Aoming, et al.
Publicado: (2025)
Can Multimodal Large Language Model Think Analogically?
por: Guo, Diandian, et al.
Publicado: (2024)
por: Guo, Diandian, et al.
Publicado: (2024)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
por: Dissen, Yehoshua, et al.
Publicado: (2024)
por: Dissen, Yehoshua, et al.
Publicado: (2024)
Cheating the Kids.
por: Meltzer, Bonnie
Publicado: (2000)
por: Meltzer, Bonnie
Publicado: (2000)
LLaVA-Critic: Learning to Evaluate Multimodal Models
por: Xiong, Tianyi, et al.
Publicado: (2024)
por: Xiong, Tianyi, et al.
Publicado: (2024)
Posture Clip: Sit properly or I wont let you work
por: Majhi, Arka, et al.
Publicado: (2026)
por: Majhi, Arka, et al.
Publicado: (2026)
Achieving 3D Attention via Triplet Squeeze and Excitation Block
por: Alhazmi, Maan, et al.
Publicado: (2025)
por: Alhazmi, Maan, et al.
Publicado: (2025)
Adapting Small Language Models to Low-Resource Domains: A Case Study in Hindi Tourism QA
por: Majhi, Sandipan, et al.
Publicado: (2025)
por: Majhi, Sandipan, et al.
Publicado: (2025)
Kids Not Getting the Web Access They Want
por: Minkel, Walter
Publicado: (2004)
por: Minkel, Walter
Publicado: (2004)
Analysis of AWW (Anganwadi Workers) Training Content, ILA (Incremental Learning Approach) Modules Following CDT (Component Display Theory)
por: Majhi, Arka, et al.
Publicado: (2026)
por: Majhi, Arka, et al.
Publicado: (2026)
From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG
por: Chen, Guanhua, et al.
Publicado: (2026)
por: Chen, Guanhua, et al.
Publicado: (2026)
Ejemplares similares
-
Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions
por: Yiu, Eunice, et al.
Publicado: (2025) -
Breaking the Assistant Mold: Modeling Behavioral Variation in LLM Based Procedural Character Generation
por: Qraitem, Maan, et al.
Publicado: (2026) -
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
por: Qraitem, Maan, et al.
Publicado: (2023) -
SLANT: Spurious Logo ANalysis Toolkit
por: Qraitem, Maan, et al.
Publicado: (2024) -
Web Artifact Attacks Disrupt Vision Language Models
por: Qraitem, Maan, et al.
Publicado: (2025)