FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rodionov, Fedor, Eldesokey, Abdelrahman, Birsak, Michael, Femiani, John, Ghanem, Bernard, Wonka, Peter |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MatCLIP: Light- and Shape-Insensitive Assignment of PBR Material Models
von: Birsak, Michael, et al.
Veröffentlicht: (2025)
von: Birsak, Michael, et al.
Veröffentlicht: (2025)
Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2025)
NearID: Identity Representation Learning via Near-identity Distractors
von: Cvejic, Aleksandar, et al.
Veröffentlicht: (2026)
von: Cvejic, Aleksandar, et al.
Veröffentlicht: (2026)
LatentMan: Generating Consistent Animated Characters using Image Diffusion Models
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2023)
Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2024)
CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models
von: Alzahrani, Reem, et al.
Veröffentlicht: (2026)
von: Alzahrani, Reem, et al.
Veröffentlicht: (2026)
PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
von: Cvejic, Aleksandar, et al.
Veröffentlicht: (2025)
von: Cvejic, Aleksandar, et al.
Veröffentlicht: (2025)
EditCLIP: Representation Learning for Image Editing
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
von: Shi, Jian, et al.
Veröffentlicht: (2026)
von: Shi, Jian, et al.
Veröffentlicht: (2026)
WinSyn: A High Resolution Testbed for Synthetic Data
von: Kelly, Tom, et al.
Veröffentlicht: (2023)
von: Kelly, Tom, et al.
Veröffentlicht: (2023)
PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes
von: Abdelreheem, Ahmed, et al.
Veröffentlicht: (2025)
von: Abdelreheem, Ahmed, et al.
Veröffentlicht: (2025)
EVGeoQA: Benchmarking LLMs on Dynamic, Multi-Objective Geo-Spatial Exploration
von: Wu, Jianfei, et al.
Veröffentlicht: (2026)
von: Wu, Jianfei, et al.
Veröffentlicht: (2026)
Executable World Models for ARC-AGI-3 in the Era of Coding Agents
von: Rodionov, Sergey
Veröffentlicht: (2026)
von: Rodionov, Sergey
Veröffentlicht: (2026)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
von: Tian, Kexin, et al.
Veröffentlicht: (2025)
CON-QA: Privacy-Preserving QA using cloud LLMs in Contract Domain
von: Singh, Ajeet Kumar, et al.
Veröffentlicht: (2025)
von: Singh, Ajeet Kumar, et al.
Veröffentlicht: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
EasyV2V: A High-quality Instruction-based Video Editing Framework
von: Mai, Jinjie, et al.
Veröffentlicht: (2025)
von: Mai, Jinjie, et al.
Veröffentlicht: (2025)
ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
von: Xu, Rui, et al.
Veröffentlicht: (2025)
von: Xu, Rui, et al.
Veröffentlicht: (2025)
Learning to Correct for QA Reasoning with Black-box LLMs
von: Kim, Jaehyung, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyung, et al.
Veröffentlicht: (2024)
CAPTAIN: Semantic Feature Injection for Memorization Mitigation in Text-to-Image Diffusion Models
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
von: Cai, Yanan, et al.
Veröffentlicht: (2025)
von: Cai, Yanan, et al.
Veröffentlicht: (2025)
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
von: Opsahl-Ong, Krista, et al.
Veröffentlicht: (2026)
von: Opsahl-Ong, Krista, et al.
Veröffentlicht: (2026)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
Advancing Routing-Awareness in Analog ICs Floorplanning
von: Basso, Davide, et al.
Veröffentlicht: (2025)
von: Basso, Davide, et al.
Veröffentlicht: (2025)
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
von: Lu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Lu, Zhiyuan, et al.
Veröffentlicht: (2026)
PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
von: Zhang, Qiran, et al.
Veröffentlicht: (2026)
von: Zhang, Qiran, et al.
Veröffentlicht: (2026)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
Purifying Task Vectors in Knowledge-Aware Subspace for Model Merging
von: An, Bang, et al.
Veröffentlicht: (2025)
von: An, Bang, et al.
Veröffentlicht: (2025)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
Improving Rule-based Reasoning in LLMs using Neurosymbolic Representations
von: Dhanraj, Varun, et al.
Veröffentlicht: (2025)
von: Dhanraj, Varun, et al.
Veröffentlicht: (2025)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
von: Wang, Minzheng, et al.
Veröffentlicht: (2024)
Entropy and Attention Dynamics in Small Language Models: A Trace-Level Structural Analysis on the TruthfulQA Benchmark
von: Adeseye, Adeyemi, et al.
Veröffentlicht: (2026)
von: Adeseye, Adeyemi, et al.
Veröffentlicht: (2026)
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
von: Ma, Jianzhe, et al.
Veröffentlicht: (2026)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
von: Zhou, Zhiyu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhiyu, et al.
Veröffentlicht: (2026)
Enhancing Reinforcement Learning for the Floorplanning of Analog ICs with Beam Search
von: Della Rovere, Sandro Junior, et al.
Veröffentlicht: (2025)
von: Della Rovere, Sandro Junior, et al.
Veröffentlicht: (2025)
LLMs Can Compensate for Deficiencies in Visual Representations
von: Takishita, Sho, et al.
Veröffentlicht: (2025)
von: Takishita, Sho, et al.
Veröffentlicht: (2025)
FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence
von: Arun, Abhinav, et al.
Veröffentlicht: (2025)
von: Arun, Abhinav, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MatCLIP: Light- and Shape-Insensitive Assignment of PBR Material Models
von: Birsak, Michael, et al.
Veröffentlicht: (2025) -
Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2025) -
NearID: Identity Representation Learning via Near-identity Distractors
von: Cvejic, Aleksandar, et al.
Veröffentlicht: (2026) -
LatentMan: Generating Consistent Animated Characters using Image Diffusion Models
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2023) -
Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2024)