Evaluating Time Awareness and Cross-modal Active Perception of Large Models via 4D Escape Room Task
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Yurui, Wang, Ziyue, Lu, Shuyun, Liu, Dairu, Liu, Xuechen, Luo, Fuwen, Li, Peng, Liu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability
von: Wang, Ziyue, et al.
Veröffentlicht: (2025)
von: Wang, Ziyue, et al.
Veröffentlicht: (2025)
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
von: Wang, Ziyue, et al.
Veröffentlicht: (2024)
von: Wang, Ziyue, et al.
Veröffentlicht: (2024)
Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
von: Liu, Dairu, et al.
Veröffentlicht: (2025)
von: Liu, Dairu, et al.
Veröffentlicht: (2025)
Perspective Transition of Large Language Models for Solving Subjective Tasks
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
von: Wan, Zihao, et al.
Veröffentlicht: (2025)
von: Wan, Zihao, et al.
Veröffentlicht: (2025)
DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms
von: Bi, Xiaojun, et al.
Veröffentlicht: (2025)
von: Bi, Xiaojun, et al.
Veröffentlicht: (2025)
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
von: Luo, Fuwen, et al.
Veröffentlicht: (2025)
von: Luo, Fuwen, et al.
Veröffentlicht: (2025)
Collaborative Cross-modal Fusion with Large Language Model for Recommendation
von: Liu, Zhongzhou, et al.
Veröffentlicht: (2024)
von: Liu, Zhongzhou, et al.
Veröffentlicht: (2024)
Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2023)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2023)
MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
von: Qin, Yiran, et al.
Veröffentlicht: (2023)
von: Qin, Yiran, et al.
Veröffentlicht: (2023)
Region-Aware CAM: High-Resolution Weakly-Supervised Defect Segmentation via Salient Region Perception
von: Dong, Hang-Cheng, et al.
Veröffentlicht: (2025)
von: Dong, Hang-Cheng, et al.
Veröffentlicht: (2025)
VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception
von: Chen, Chuang, et al.
Veröffentlicht: (2024)
von: Chen, Chuang, et al.
Veröffentlicht: (2024)
Training Dynamics-Aware Multi-Factor Curriculum Learning for Target Speaker Extraction
von: Liu, Yun, et al.
Veröffentlicht: (2026)
von: Liu, Yun, et al.
Veröffentlicht: (2026)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
von: Liu, Hugh Xuechen, et al.
Veröffentlicht: (2026)
von: Liu, Hugh Xuechen, et al.
Veröffentlicht: (2026)
Model Composition for Multimodal Large Language Models
von: Chen, Chi, et al.
Veröffentlicht: (2024)
von: Chen, Chi, et al.
Veröffentlicht: (2024)
Browse and Concentrate: Comprehending Multimodal Content via prior-LLM Context Fusion
von: Wang, Ziyue, et al.
Veröffentlicht: (2024)
von: Wang, Ziyue, et al.
Veröffentlicht: (2024)
Cross-modal Active Complementary Learning with Self-refining Correspondence
von: Qin, Yang, et al.
Veröffentlicht: (2023)
von: Qin, Yang, et al.
Veröffentlicht: (2023)
Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
von: Zhang, Jinlu, et al.
Veröffentlicht: (2024)
von: Zhang, Jinlu, et al.
Veröffentlicht: (2024)
GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles
von: Shan, Mengyi, et al.
Veröffentlicht: (2025)
von: Shan, Mengyi, et al.
Veröffentlicht: (2025)
Enabling Stroke-Level Structural Analysis of Hieroglyphic Scripts without Language-Specific Priors
von: Luo, Fuwen, et al.
Veröffentlicht: (2026)
von: Luo, Fuwen, et al.
Veröffentlicht: (2026)
ELMAR: Enhancing LiDAR Detection with 4D Radar Motion Awareness and Cross-modal Uncertainty
von: Peng, Xiangyuan, et al.
Veröffentlicht: (2025)
von: Peng, Xiangyuan, et al.
Veröffentlicht: (2025)
Multi-Modality Distillation via Learning the teacher's modality-level Gram Matrix
von: Liu, Peng
Veröffentlicht: (2021)
von: Liu, Peng
Veröffentlicht: (2021)
Are Escape Rooms an Antidote for Anxiety in Simulation? An Experimental Randomized Study Comparing High‐Fidelity Simulation to an Escape Room Format
von: Aubrey A. Bethel, et al.
Veröffentlicht: (2025)
von: Aubrey A. Bethel, et al.
Veröffentlicht: (2025)
The Solution for Temporal Sound Localisation Task of ICCV 1st Perception Test Challenge 2023
von: Huang, Yurui, et al.
Veröffentlicht: (2024)
von: Huang, Yurui, et al.
Veröffentlicht: (2024)
Zero-Day Audio DeepFake Detection via Retrieval Augmentation and Profile Matching
von: Liu, Xuechen, et al.
Veröffentlicht: (2025)
von: Liu, Xuechen, et al.
Veröffentlicht: (2025)
Monitoring Decoding: Mitigating Hallucination via Evaluating the Factuality of Partial Response during Generation
von: Chang, Yurui, et al.
Veröffentlicht: (2025)
von: Chang, Yurui, et al.
Veröffentlicht: (2025)
SkyNET-scape Room: An Escape Room to Explore Astroparticle Physics
von: Prandini, Elisa, et al.
Veröffentlicht: (2025)
von: Prandini, Elisa, et al.
Veröffentlicht: (2025)
The bidirectional NLS approximation for the one-dimensional Euler-Poisson system
von: Liu, Huimin, et al.
Veröffentlicht: (2025)
von: Liu, Huimin, et al.
Veröffentlicht: (2025)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
Beyond Words: Evaluating and Bridging Epistemic Divergence in User-Agent Interaction via Theory of Mind
von: Ruan, Minyuan, et al.
Veröffentlicht: (2026)
von: Ruan, Minyuan, et al.
Veröffentlicht: (2026)
Design of an Educational Escape Room by Future Teachers
von: Vankúš, Peter, et al.
Veröffentlicht: (2023)
von: Vankúš, Peter, et al.
Veröffentlicht: (2023)
Fusion-Mamba for Cross-modality Object Detection
von: Dong, Wenhao, et al.
Veröffentlicht: (2024)
von: Dong, Wenhao, et al.
Veröffentlicht: (2024)
Escape Room combined with European Board Game Concepts for self-adjusted Challenge Levels: An educational Eurogame Escape Room in Physics
von: Bräuninger, Sascha Albert, et al.
Veröffentlicht: (2024)
von: Bräuninger, Sascha Albert, et al.
Veröffentlicht: (2024)
Generalizing Speaker Verification for Spoof Awareness in the Embedding Space
von: Liu, Xuechen, et al.
Veröffentlicht: (2024)
von: Liu, Xuechen, et al.
Veröffentlicht: (2024)
MemCollab: Cross-Model Memory Collaboration via Contrastive Trajectory Distillation
von: Chang, Yurui, et al.
Veröffentlicht: (2026)
von: Chang, Yurui, et al.
Veröffentlicht: (2026)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025)
von: Yao, Lei, et al.
Veröffentlicht: (2025)
Improving curriculum learning for target speaker extraction with synthetic speakers
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability
von: Wang, Ziyue, et al.
Veröffentlicht: (2025) -
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
von: Wang, Ziyue, et al.
Veröffentlicht: (2024) -
Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
von: Liu, Dairu, et al.
Veröffentlicht: (2025) -
Perspective Transition of Large Language Models for Solving Subjective Tasks
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025) -
Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)