KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yiu, Eunice, Qraitem, Maan, Majhi, Anisa Noor, Wong, Charlie, Bai, Yutong, Ginosar, Shiry, Gopnik, Alison, Saenko, Kate |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions
von: Yiu, Eunice, et al.
Veröffentlicht: (2025)
von: Yiu, Eunice, et al.
Veröffentlicht: (2025)
Breaking the Assistant Mold: Modeling Behavioral Variation in LLM Based Procedural Character Generation
von: Qraitem, Maan, et al.
Veröffentlicht: (2026)
von: Qraitem, Maan, et al.
Veröffentlicht: (2026)
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
von: Qraitem, Maan, et al.
Veröffentlicht: (2023)
von: Qraitem, Maan, et al.
Veröffentlicht: (2023)
SLANT: Spurious Logo ANalysis Toolkit
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
Web Artifact Attacks Disrupt Vision Language Models
von: Qraitem, Maan, et al.
Veröffentlicht: (2025)
von: Qraitem, Maan, et al.
Veröffentlicht: (2025)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
von: Qraitem, Maan, et al.
Veröffentlicht: (2024)
Poly-Autoregressive Prediction for Modeling Interactions
von: Thakkar, Neerja, et al.
Veröffentlicht: (2025)
von: Thakkar, Neerja, et al.
Veröffentlicht: (2025)
Pose Priors from Language Models
von: Subramanian, Sanjay, et al.
Veröffentlicht: (2024)
von: Subramanian, Sanjay, et al.
Veröffentlicht: (2024)
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
von: Koepke, A. Sophia, et al.
Veröffentlicht: (2026)
von: Koepke, A. Sophia, et al.
Veröffentlicht: (2026)
Diffusion Models as Data Mining Tools
von: Siglidis, Ioannis, et al.
Veröffentlicht: (2024)
von: Siglidis, Ioannis, et al.
Veröffentlicht: (2024)
Forecasting Motion in the Wild
von: Thakkar, Neerja, et al.
Veröffentlicht: (2026)
von: Thakkar, Neerja, et al.
Veröffentlicht: (2026)
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
von: Mishra, Samarth, et al.
Veröffentlicht: (2025)
von: Mishra, Samarth, et al.
Veröffentlicht: (2025)
Gaussian Masked Autoencoders
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
Synergy and Synchrony in Couple Dances
von: Maluleke, Vongani, et al.
Veröffentlicht: (2024)
von: Maluleke, Vongani, et al.
Veröffentlicht: (2024)
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
von: Petsiuk, Vitali, et al.
Veröffentlicht: (2024)
von: Petsiuk, Vitali, et al.
Veröffentlicht: (2024)
Tell Me What's Next: Textual Foresight for Generic UI Representations
von: Burns, Andrea, et al.
Veröffentlicht: (2024)
von: Burns, Andrea, et al.
Veröffentlicht: (2024)
Children's Mental Models of Generative Visual and Text Based AI Models
von: Kosoy, Eliza, et al.
Veröffentlicht: (2024)
von: Kosoy, Eliza, et al.
Veröffentlicht: (2024)
KiC: Keyword-inspired Cascade for Cost-Efficient Text Generation with LLMs
von: Kim, Woo-Chan, et al.
Veröffentlicht: (2025)
von: Kim, Woo-Chan, et al.
Veröffentlicht: (2025)
Federated Adversarial Domain Adaptation
von: Peng, Xingchao, et al.
Veröffentlicht: (2019)
von: Peng, Xingchao, et al.
Veröffentlicht: (2019)
SingaKids: A Multilingual Multimodal Dialogic Tutor for Language Learning
von: Liu, Zhengyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zhengyuan, et al.
Veröffentlicht: (2025)
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
von: Miller, Kevin, et al.
Veröffentlicht: (2025)
von: Miller, Kevin, et al.
Veröffentlicht: (2025)
Enhancing immersion in Virtual Reality sports through Physical Interactions
von: Majhi, Arka
Veröffentlicht: (2026)
von: Majhi, Arka
Veröffentlicht: (2026)
OP-LoRA: The Blessing of Dimensionality
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
Frozen Forecasting: A Unified Evaluation
von: Walker, Jacob C, et al.
Veröffentlicht: (2025)
von: Walker, Jacob C, et al.
Veröffentlicht: (2025)
CLAMP: Contrastive LAnguage Model Prompt-tuning
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023)
Is Large-Scale Pretraining the Secret to Good Domain Generalization?
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2024)
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
von: Mishra, Samarth, et al.
Veröffentlicht: (2023)
von: Mishra, Samarth, et al.
Veröffentlicht: (2023)
ERM++: An Improved Baseline for Domain Generalization
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023)
von: Teterwak, Piotr, et al.
Veröffentlicht: (2023)
An Author in Every Classroom: Kids Connecting with Authors via Skype. It's the next Best Thing to Being There
von: Messner, Kate
Veröffentlicht: (2010)
von: Messner, Kate
Veröffentlicht: (2010)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
von: Liu, Aoming, et al.
Veröffentlicht: (2025)
von: Liu, Aoming, et al.
Veröffentlicht: (2025)
Can Multimodal Large Language Model Think Analogically?
von: Guo, Diandian, et al.
Veröffentlicht: (2024)
von: Guo, Diandian, et al.
Veröffentlicht: (2024)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
von: Dissen, Yehoshua, et al.
Veröffentlicht: (2024)
von: Dissen, Yehoshua, et al.
Veröffentlicht: (2024)
Cheating the Kids.
von: Meltzer, Bonnie
Veröffentlicht: (2000)
von: Meltzer, Bonnie
Veröffentlicht: (2000)
LLaVA-Critic: Learning to Evaluate Multimodal Models
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
Posture Clip: Sit properly or I wont let you work
von: Majhi, Arka, et al.
Veröffentlicht: (2026)
von: Majhi, Arka, et al.
Veröffentlicht: (2026)
Achieving 3D Attention via Triplet Squeeze and Excitation Block
von: Alhazmi, Maan, et al.
Veröffentlicht: (2025)
von: Alhazmi, Maan, et al.
Veröffentlicht: (2025)
Adapting Small Language Models to Low-Resource Domains: A Case Study in Hindi Tourism QA
von: Majhi, Sandipan, et al.
Veröffentlicht: (2025)
von: Majhi, Sandipan, et al.
Veröffentlicht: (2025)
Kids Not Getting the Web Access They Want
von: Minkel, Walter
Veröffentlicht: (2004)
von: Minkel, Walter
Veröffentlicht: (2004)
Analysis of AWW (Anganwadi Workers) Training Content, ILA (Incremental Learning Approach) Modules Following CDT (Component Display Theory)
von: Majhi, Arka, et al.
Veröffentlicht: (2026)
von: Majhi, Arka, et al.
Veröffentlicht: (2026)
From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG
von: Chen, Guanhua, et al.
Veröffentlicht: (2026)
von: Chen, Guanhua, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions
von: Yiu, Eunice, et al.
Veröffentlicht: (2025) -
Breaking the Assistant Mold: Modeling Behavioral Variation in LLM Based Procedural Character Generation
von: Qraitem, Maan, et al.
Veröffentlicht: (2026) -
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
von: Qraitem, Maan, et al.
Veröffentlicht: (2023) -
SLANT: Spurious Logo ANalysis Toolkit
von: Qraitem, Maan, et al.
Veröffentlicht: (2024) -
Web Artifact Attacks Disrupt Vision Language Models
von: Qraitem, Maan, et al.
Veröffentlicht: (2025)