STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Agrawal, Palaash, Azaman, Haidi, Tan, Cheston |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Generation of Spatial Relations in Text and Image Generative Models
by: Sim, Shang Hong, et al.
Published: (2024)
by: Sim, Shang Hong, et al.
Published: (2024)
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022)
by: Tan, Clement, et al.
Published: (2022)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
by: Nagar, Aishik, et al.
Published: (2024)
by: Nagar, Aishik, et al.
Published: (2024)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
by: Jaiswal, Shantanu, et al.
Published: (2024)
by: Jaiswal, Shantanu, et al.
Published: (2024)
Human-like compositional learning of visually-grounded concepts using synthetic environments
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
Can LLMs perform structured graph reasoning?
by: Agrawal, Palaash, et al.
Published: (2024)
by: Agrawal, Palaash, et al.
Published: (2024)
Stencil: Subject-Driven Generation with Context Guidance
by: Chen, Gordon, et al.
Published: (2025)
by: Chen, Gordon, et al.
Published: (2025)
Synthetic Art Generation and DeepFake Detection A Study on Jamini Roy Inspired Dataset
by: Agrawal, Kushal, et al.
Published: (2025)
by: Agrawal, Kushal, et al.
Published: (2025)
Visual Agentic AI for Spatial Reasoning with a Dynamic API
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
Global License Plate Dataset
by: Agrawal, Siddharth
Published: (2024)
by: Agrawal, Siddharth
Published: (2024)
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
by: Deng, Nianchen, et al.
Published: (2025)
by: Deng, Nianchen, et al.
Published: (2025)
Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos
by: Jiang, Songtao, et al.
Published: (2026)
by: Jiang, Songtao, et al.
Published: (2026)
CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography
by: Prakash, Eva, et al.
Published: (2026)
by: Prakash, Eva, et al.
Published: (2026)
Temporal-Spatial Object Relations Modeling for Vision-and-Language Navigation
by: Huang, Bowen, et al.
Published: (2024)
by: Huang, Bowen, et al.
Published: (2024)
WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding
by: Kong, Quan, et al.
Published: (2024)
by: Kong, Quan, et al.
Published: (2024)
Learning Multi-View Spatial Reasoning from Cross-View Relations
by: Jeong, Suchae, et al.
Published: (2026)
by: Jeong, Suchae, et al.
Published: (2026)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)
by: Yang, Qian, et al.
Published: (2026)
DivCon: Divide and Conquer for Complex Numerical and Spatial Reasoning in Text-to-Image Generation
by: Jia, Yuhao, et al.
Published: (2024)
by: Jia, Yuhao, et al.
Published: (2024)
STRMs: Spatial Temporal Reasoning Models for Vision-Based Localization Rivaling GPS Precision
by: Lui, Hin Wai, et al.
Published: (2025)
by: Lui, Hin Wai, et al.
Published: (2025)
PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
by: Hao, Yanbin, et al.
Published: (2024)
by: Hao, Yanbin, et al.
Published: (2024)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
by: Liang, Yiming, et al.
Published: (2026)
by: Liang, Yiming, et al.
Published: (2026)
FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing
by: Wu, Bizhu, et al.
Published: (2025)
by: Wu, Bizhu, et al.
Published: (2025)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
by: Hu, Yuhang, et al.
Published: (2025)
by: Hu, Yuhang, et al.
Published: (2025)
WHU-Synthetic: A Synthetic Perception Dataset for 3-D Multitask Model Research
by: Zhou, Jiahao, et al.
Published: (2024)
by: Zhou, Jiahao, et al.
Published: (2024)
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
by: Wu, Xiyang, et al.
Published: (2025)
by: Wu, Xiyang, et al.
Published: (2025)
An Approach to Enriching Surgical Video Datasets for Fine-Grained Spatial-Temporal Understanding of Vision-Language Models
by: Maack, Lennart, et al.
Published: (2026)
by: Maack, Lennart, et al.
Published: (2026)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
Symmetria: A Synthetic Dataset for Learning in Point Clouds
by: Sipiran, Ivan, et al.
Published: (2025)
by: Sipiran, Ivan, et al.
Published: (2025)
Synthetic Fungi Datasets: A Time-Aligned Approach
by: Rani, A., et al.
Published: (2025)
by: Rani, A., et al.
Published: (2025)
Unveiling Synthetic Faces: How Synthetic Datasets Can Expose Real Identities
by: Shahreza, Hatef Otroshi, et al.
Published: (2024)
by: Shahreza, Hatef Otroshi, et al.
Published: (2024)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
by: Liao, Ruotong, et al.
Published: (2024)
by: Liao, Ruotong, et al.
Published: (2024)
AM Flow: Adapters for Temporal Processing in Action Recognition
by: Agrawal, Tanay, et al.
Published: (2024)
by: Agrawal, Tanay, et al.
Published: (2024)
SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
by: Ogezi, Michael, et al.
Published: (2025)
by: Ogezi, Michael, et al.
Published: (2025)
On Applicability of Synthetic Datasets for Facial Expression Recognition
by: Azmoudeh, Ali, et al.
Published: (2026)
by: Azmoudeh, Ali, et al.
Published: (2026)
Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning
by: Muturi, Tanner, et al.
Published: (2025)
by: Muturi, Tanner, et al.
Published: (2025)
Similar Items
-
Evaluating the Generation of Spatial Relations in Text and Image Generative Models
by: Sim, Shang Hong, et al.
Published: (2024) -
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
by: Lin, Zijun, et al.
Published: (2025) -
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022) -
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
by: Nagar, Aishik, et al.
Published: (2024) -
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
by: Jaiswal, Shantanu, et al.
Published: (2024)