What Makes a Maze Look Like a Maze?
Fuente:
arXiv
Guardado en:
| Autores principales: | Hsu, Joy, Mao, Jiayuan, Tenenbaum, Joshua B., Goodman, Noah D., Wu, Jiajun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Neuro-Symbolic Concepts
por: Mao, Jiayuan, et al.
Publicado: (2025)
por: Mao, Jiayuan, et al.
Publicado: (2025)
CLEVRER-Humans: Describing Physical and Causal Events the Human Way
por: Mao, Jiayuan, et al.
Publicado: (2023)
por: Mao, Jiayuan, et al.
Publicado: (2023)
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
por: Li, Hong, et al.
Publicado: (2024)
por: Li, Hong, et al.
Publicado: (2024)
One-Shot Manipulation Strategy Learning by Making Contact Analogies
por: Liu, Yuyao, et al.
Publicado: (2024)
por: Liu, Yuyao, et al.
Publicado: (2024)
Learning Compositional Behaviors from Demonstration and Language
por: Liu, Weiyu, et al.
Publicado: (2025)
por: Liu, Weiyu, et al.
Publicado: (2025)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
por: Feng, Chun, et al.
Publicado: (2024)
por: Feng, Chun, et al.
Publicado: (2024)
STAR: A Benchmark for Situated Reasoning in Real-World Videos
por: Wu, Bo, et al.
Publicado: (2024)
por: Wu, Bo, et al.
Publicado: (2024)
HAZARD Challenge: Embodied Decision Making in Dynamically Changing Environments
por: Zhou, Qinhong, et al.
Publicado: (2024)
por: Zhou, Qinhong, et al.
Publicado: (2024)
Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
por: Deng, Hokin
Publicado: (2025)
por: Deng, Hokin
Publicado: (2025)
Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
por: Yang, Cheng, et al.
Publicado: (2025)
por: Yang, Cheng, et al.
Publicado: (2025)
Visually Descriptive Language Model for Vector Graphics Reasoning
por: Wang, Zhenhailong, et al.
Publicado: (2024)
por: Wang, Zhenhailong, et al.
Publicado: (2024)
Composable Part-Based Manipulation
por: Liu, Weiyu, et al.
Publicado: (2024)
por: Liu, Weiyu, et al.
Publicado: (2024)
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
por: Hsu, Joy, et al.
Publicado: (2025)
por: Hsu, Joy, et al.
Publicado: (2025)
Building Cooperative Embodied Agents Modularly with Large Language Models
por: Zhang, Hongxin, et al.
Publicado: (2023)
por: Zhang, Hongxin, et al.
Publicado: (2023)
Moral Mazes in the Era of LLMs
por: Nguyen, Dang, et al.
Publicado: (2026)
por: Nguyen, Dang, et al.
Publicado: (2026)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
por: Cherian, Anoop, et al.
Publicado: (2024)
por: Cherian, Anoop, et al.
Publicado: (2024)
Predicate Hierarchies Improve Few-Shot State Classification
por: Jin, Emily, et al.
Publicado: (2025)
por: Jin, Emily, et al.
Publicado: (2025)
Keypoint Abstraction using Large Models for Object-Relative Imitation Learning
por: Fang, Xiaolin, et al.
Publicado: (2024)
por: Fang, Xiaolin, et al.
Publicado: (2024)
Agent S: An Open Agentic Framework that Uses Computers Like a Human
por: Agashe, Saaket, et al.
Publicado: (2024)
por: Agashe, Saaket, et al.
Publicado: (2024)
Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
por: Hong, Haodong, et al.
Publicado: (2024)
por: Hong, Haodong, et al.
Publicado: (2024)
General Scene Adaptation for Vision-and-Language Navigation
por: Hong, Haodong, et al.
Publicado: (2025)
por: Hong, Haodong, et al.
Publicado: (2025)
MMToM-QA: Multimodal Theory of Mind Question Answering
por: Jin, Chuanyang, et al.
Publicado: (2024)
por: Jin, Chuanyang, et al.
Publicado: (2024)
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry
por: Sakib, Syed Nazmus, et al.
Publicado: (2026)
por: Sakib, Syed Nazmus, et al.
Publicado: (2026)
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
por: Morini, Marco, et al.
Publicado: (2026)
por: Morini, Marco, et al.
Publicado: (2026)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
por: Zhang, Jiarui, et al.
Publicado: (2025)
por: Zhang, Jiarui, et al.
Publicado: (2025)
What You See is What You Ask: Evaluating Audio Descriptions
por: Kala, Divy, et al.
Publicado: (2025)
por: Kala, Divy, et al.
Publicado: (2025)
Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
por: Fang, Ye, et al.
Publicado: (2024)
por: Fang, Ye, et al.
Publicado: (2024)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
por: Chen, Jiali, et al.
Publicado: (2024)
por: Chen, Jiali, et al.
Publicado: (2024)
Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
por: Chen, Qiyuan, et al.
Publicado: (2026)
por: Chen, Qiyuan, et al.
Publicado: (2026)
Can Large Language Models Understand Symbolic Graphics Programs?
por: Qiu, Zeju, et al.
Publicado: (2024)
por: Qiu, Zeju, et al.
Publicado: (2024)
Evaluating Automated Radiology Report Quality through Fine-Grained Phrasal Grounding of Clinical Findings
por: Mahmood, Razi, et al.
Publicado: (2024)
por: Mahmood, Razi, et al.
Publicado: (2024)
Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
por: Buckley, Thomas, et al.
Publicado: (2023)
por: Buckley, Thomas, et al.
Publicado: (2023)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
por: Zhou, Yuqi, et al.
Publicado: (2025)
por: Zhou, Yuqi, et al.
Publicado: (2025)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
por: Chen, Qiguang, et al.
Publicado: (2026)
por: Chen, Qiguang, et al.
Publicado: (2026)
Data or Language Supervision: What Makes CLIP Better than DINO?
por: Liu, Yiming, et al.
Publicado: (2025)
por: Liu, Yiming, et al.
Publicado: (2025)
Holistic Evaluation for Interleaved Text-and-Image Generation
por: Liu, Minqian, et al.
Publicado: (2024)
por: Liu, Minqian, et al.
Publicado: (2024)
Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment
por: Zhang, Yue, et al.
Publicado: (2025)
por: Zhang, Yue, et al.
Publicado: (2025)
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
por: Yao, Jihan, et al.
Publicado: (2025)
por: Yao, Jihan, et al.
Publicado: (2025)
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
por: Qin, Libo, et al.
Publicado: (2024)
por: Qin, Libo, et al.
Publicado: (2024)
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
por: Chen, Haonan, et al.
Publicado: (2025)
por: Chen, Haonan, et al.
Publicado: (2025)
Ejemplares similares
-
Neuro-Symbolic Concepts
por: Mao, Jiayuan, et al.
Publicado: (2025) -
CLEVRER-Humans: Describing Physical and Causal Events the Human Way
por: Mao, Jiayuan, et al.
Publicado: (2023) -
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
por: Li, Hong, et al.
Publicado: (2024) -
One-Shot Manipulation Strategy Learning by Making Contact Analogies
por: Liu, Yuyao, et al.
Publicado: (2024) -
Learning Compositional Behaviors from Demonstration and Language
por: Liu, Weiyu, et al.
Publicado: (2025)