GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles
Fuente:
arXiv
Saved in:
| Main Authors: | Shan, Mengyi, Curless, Brian, Kemelmacher-Shlizerman, Ira, Seitz, Steve |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
COMIC: Agentic Sketch Comedy Generation
by: Hong, Susung, et al.
Published: (2026)
by: Hong, Susung, et al.
Published: (2026)
Generating Fit Check Videos with a Handheld Camera
by: Chen, Bowei, et al.
Published: (2025)
by: Chen, Bowei, et al.
Published: (2025)
Total Selfie: Generating Full-Body Selfies
by: Chen, Bowei, et al.
Published: (2023)
by: Chen, Bowei, et al.
Published: (2023)
UltraZoom: Generating Gigapixel Images from Regular Photos
by: Ma, Jingwei, et al.
Published: (2025)
by: Ma, Jingwei, et al.
Published: (2025)
Inverse Painting: Reconstructing The Painting Process
by: Chen, Bowei, et al.
Published: (2024)
by: Chen, Bowei, et al.
Published: (2024)
MusicInfuser: Making Video Diffusion Listen and Dance
by: Hong, Susung, et al.
Published: (2025)
by: Hong, Susung, et al.
Published: (2025)
Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation
by: Wang, Xiaojuan, et al.
Published: (2024)
by: Wang, Xiaojuan, et al.
Published: (2024)
How Animals Dance (When You're Not Looking)
by: Wang, Xiaojuan, et al.
Published: (2025)
by: Wang, Xiaojuan, et al.
Published: (2025)
Generative Powers of Ten
by: Wang, Xiaojuan, et al.
Published: (2023)
by: Wang, Xiaojuan, et al.
Published: (2023)
Don't Look at the Camera: Achieving Perceived Eye Contact
by: Gao, Alice, et al.
Published: (2024)
by: Gao, Alice, et al.
Published: (2024)
Perturb-and-Revise: Flexible 3D Editing with Generative Trajectories
by: Hong, Susung, et al.
Published: (2024)
by: Hong, Susung, et al.
Published: (2024)
FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On
by: Karras, Johanna, et al.
Published: (2026)
by: Karras, Johanna, et al.
Published: (2026)
M&M VTO: Multi-Garment Virtual Try-On and Editing
by: Zhu, Luyang, et al.
Published: (2024)
by: Zhu, Luyang, et al.
Published: (2024)
EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability
by: Wang, Ziyue, et al.
Published: (2025)
by: Wang, Ziyue, et al.
Published: (2025)
HoloGarment: 360° Novel View Synthesis of In-the-Wild Garments
by: Karras, Johanna, et al.
Published: (2025)
by: Karras, Johanna, et al.
Published: (2025)
Infinite Texture: Text-guided High Resolution Diffusion Texture Synthesis
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Fashion-VDM: Video Diffusion Model for Virtual Try-On
by: Karras, Johanna, et al.
Published: (2024)
by: Karras, Johanna, et al.
Published: (2024)
Test-Time Anchoring for Discrete Diffusion Posterior Sampling
by: Rout, Litu, et al.
Published: (2025)
by: Rout, Litu, et al.
Published: (2025)
Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps
by: Mishra, Sandeep, et al.
Published: (2025)
by: Mishra, Sandeep, et al.
Published: (2025)
GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEs
by: Basioti, Kalliopi, et al.
Published: (2025)
by: Basioti, Kalliopi, et al.
Published: (2025)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
by: Wang, Xuehui, et al.
Published: (2025)
by: Wang, Xuehui, et al.
Published: (2025)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
by: Siingh, Shikhhar, et al.
Published: (2025)
by: Siingh, Shikhhar, et al.
Published: (2025)
Evaluating Time Awareness and Cross-modal Active Perception of Large Models via 4D Escape Room Task
by: Dong, Yurui, et al.
Published: (2026)
by: Dong, Yurui, et al.
Published: (2026)
SurvAgent: Hierarchical CoT-Enhanced Case Banking and Dichotomy-Based Multi-Agent System for Multimodal Survival Prediction
by: Huang, Guolin, et al.
Published: (2025)
by: Huang, Guolin, et al.
Published: (2025)
Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
by: Shao, Jie, et al.
Published: (2025)
by: Shao, Jie, et al.
Published: (2025)
AutoGen Driven Multi Agent Framework for Iterative Crime Data Analysis and Prediction
by: Fatima, Syeda Kisaa, et al.
Published: (2025)
by: Fatima, Syeda Kisaa, et al.
Published: (2025)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
by: Kang, Caixin, et al.
Published: (2025)
by: Kang, Caixin, et al.
Published: (2025)
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
by: Gal, Rinon, et al.
Published: (2024)
by: Gal, Rinon, et al.
Published: (2024)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
by: Leng, Jixuan, et al.
Published: (2025)
by: Leng, Jixuan, et al.
Published: (2025)
Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces
by: Hadgi, Souhail, et al.
Published: (2025)
by: Hadgi, Souhail, et al.
Published: (2025)
On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI
by: Restrepo, David, et al.
Published: (2025)
by: Restrepo, David, et al.
Published: (2025)
Medical Context Distorts Decisions in Clinical Vision Language Models
by: Restrepo, David, et al.
Published: (2026)
by: Restrepo, David, et al.
Published: (2026)
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
by: Wang, Junyang, et al.
Published: (2024)
by: Wang, Junyang, et al.
Published: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
by: Wu, Chengyue, et al.
Published: (2024)
by: Wu, Chengyue, et al.
Published: (2024)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
by: Wu, Chenyuan, et al.
Published: (2025)
by: Wu, Chenyuan, et al.
Published: (2025)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
by: Zhao, Xiangyu, et al.
Published: (2023)
by: Zhao, Xiangyu, et al.
Published: (2023)
Similar Items
-
COMIC: Agentic Sketch Comedy Generation
by: Hong, Susung, et al.
Published: (2026) -
Generating Fit Check Videos with a Handheld Camera
by: Chen, Bowei, et al.
Published: (2025) -
Total Selfie: Generating Full-Body Selfies
by: Chen, Bowei, et al.
Published: (2023) -
UltraZoom: Generating Gigapixel Images from Regular Photos
by: Ma, Jingwei, et al.
Published: (2025) -
Inverse Painting: Reconstructing The Painting Process
by: Chen, Bowei, et al.
Published: (2024)