Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chetan, Aditya, Cai, Eric, Kushwaha, Peeyush, Kani, Bharath Raj Nagoor, Mall, Utkarsh, Wang, Qianqian, Snavely, Noah, Hariharan, Bharath |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing
von: Kani, Bharath Raj Nagoor, et al.
Veröffentlicht: (2026)
von: Kani, Bharath Raj Nagoor, et al.
Veröffentlicht: (2026)
Learning Feature Descriptors using Camera Pose Supervision
von: Wang, Qianqian, et al.
Veröffentlicht: (2020)
von: Wang, Qianqian, et al.
Veröffentlicht: (2020)
UpFusion: Novel View Diffusion from Unposed Sparse View Observations
von: Kani, Bharath Raj Nagoor, et al.
Veröffentlicht: (2023)
von: Kani, Bharath Raj Nagoor, et al.
Veröffentlicht: (2023)
C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
von: Huang, Kuan Wei, et al.
Veröffentlicht: (2025)
von: Huang, Kuan Wei, et al.
Veröffentlicht: (2025)
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
von: Revankar, Shreelekha, et al.
Veröffentlicht: (2025)
von: Revankar, Shreelekha, et al.
Veröffentlicht: (2025)
Scale-Aware Recognition in Satellite Images under Resource Constraints
von: Revankar, Shreelekha, et al.
Veröffentlicht: (2024)
von: Revankar, Shreelekha, et al.
Veröffentlicht: (2024)
DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery
von: Mall, Utkarsh, et al.
Veröffentlicht: (2025)
von: Mall, Utkarsh, et al.
Veröffentlicht: (2025)
AllClear: A Comprehensive Dataset and Benchmark for Cloud Removal in Satellite Imagery
von: Zhou, Hangyu, et al.
Veröffentlicht: (2024)
von: Zhou, Hangyu, et al.
Veröffentlicht: (2024)
Accurate Differential Operators for Hybrid Neural Fields
von: Chetan, Aditya, et al.
Veröffentlicht: (2023)
von: Chetan, Aditya, et al.
Veröffentlicht: (2023)
MOD-UV: Learning Mobile Object Detectors from Unlabeled Videos
von: Sun, Yihong, et al.
Veröffentlicht: (2024)
von: Sun, Yihong, et al.
Veröffentlicht: (2024)
ObjectCarver: Semi-automatic segmentation, reconstruction and separation of 3D objects
von: Hassena, Gemmechu, et al.
Veröffentlicht: (2024)
von: Hassena, Gemmechu, et al.
Veröffentlicht: (2024)
FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution
von: Chou, Gene, et al.
Veröffentlicht: (2025)
von: Chou, Gene, et al.
Veröffentlicht: (2025)
KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos
von: Chou, Gene, et al.
Veröffentlicht: (2024)
von: Chou, Gene, et al.
Veröffentlicht: (2024)
MegaScenes: Scene-Level View Synthesis at Scale
von: Tung, Joseph, et al.
Veröffentlicht: (2024)
von: Tung, Joseph, et al.
Veröffentlicht: (2024)
Counter-Current Learning: A Biologically Plausible Dual Network Approach for Deep Learning
von: Kao, Chia-Hsiang, et al.
Veröffentlicht: (2024)
von: Kao, Chia-Hsiang, et al.
Veröffentlicht: (2024)
Tracking and Understanding Object Transformations
von: Sun, Yihong, et al.
Veröffentlicht: (2025)
von: Sun, Yihong, et al.
Veröffentlicht: (2025)
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
von: Chou, Gene, et al.
Veröffentlicht: (2026)
von: Chou, Gene, et al.
Veröffentlicht: (2026)
Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
von: Peng, Wenxuan, et al.
Veröffentlicht: (2026)
von: Peng, Wenxuan, et al.
Veröffentlicht: (2026)
Towards Real-Time 2D Mapping: Harnessing Drones, AI, and Computer Vision for Advanced Insights
von: Agnur, Bharath Kumar
Veröffentlicht: (2024)
von: Agnur, Bharath Kumar
Veröffentlicht: (2024)
Sports Re-ID: Improving Re-Identification Of Players In Broadcast Videos Of Team Sports
von: Comandur, Bharath
Veröffentlicht: (2022)
von: Comandur, Bharath
Veröffentlicht: (2022)
Color Bind: Exploring Color Perception in Text-to-Image Models
von: Shomer-Chai, Shay, et al.
Veröffentlicht: (2025)
von: Shomer-Chai, Shay, et al.
Veröffentlicht: (2025)
MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
von: Shaar, Shaden, et al.
Veröffentlicht: (2026)
von: Shaar, Shaden, et al.
Veröffentlicht: (2026)
Collaborative Perception in Multi-Robot Systems: Case Studies in Household Cleaning and Warehouse Operations
von: Nair, Bharath Rajiv
Veröffentlicht: (2024)
von: Nair, Bharath Rajiv
Veröffentlicht: (2024)
Live Interactive Training for Video Segmentation
von: Yang, Xinyu, et al.
Veröffentlicht: (2026)
von: Yang, Xinyu, et al.
Veröffentlicht: (2026)
DOOMGAN:High-Fidelity Dynamic Identity Obfuscation Ocular Generative Morphing
von: Krishnamurthy, Bharath, et al.
Veröffentlicht: (2025)
von: Krishnamurthy, Bharath, et al.
Veröffentlicht: (2025)
Evolving Interpretable Visual Classifiers with Large Language Models
von: Chiquier, Mia, et al.
Veröffentlicht: (2024)
von: Chiquier, Mia, et al.
Veröffentlicht: (2024)
When Every Token Counts: Optimal Segmentation for Low-Resource Language Models
von: Raj, Bharath, et al.
Veröffentlicht: (2024)
von: Raj, Bharath, et al.
Veröffentlicht: (2024)
MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation
von: Krishnamurthy, Bharath, et al.
Veröffentlicht: (2026)
von: Krishnamurthy, Bharath, et al.
Veröffentlicht: (2026)
Semantic Labeling of Large-Area Geographic Regions Using Multi-View and Multi-Date Satellite Images and Noisy OSM Training Labels
von: Comandur, Bharath, et al.
Veröffentlicht: (2020)
von: Comandur, Bharath, et al.
Veröffentlicht: (2020)
How Video Meetings Change Your Expression
von: Sarin, Sumit, et al.
Veröffentlicht: (2024)
von: Sarin, Sumit, et al.
Veröffentlicht: (2024)
Wide-Baseline Relative Camera Pose Estimation with Directional Learning
von: Chen, Kefan, et al.
Veröffentlicht: (2021)
von: Chen, Kefan, et al.
Veröffentlicht: (2021)
SatFlow: Generative model based framework for producing High Resolution Gap Free Remote Sensing Imagery
von: Irigireddy, Bharath, et al.
Veröffentlicht: (2025)
von: Irigireddy, Bharath, et al.
Veröffentlicht: (2025)
Open Source Infrastructure for Automatic Cell Segmentation
von: Menezes, Aaron Rock, et al.
Veröffentlicht: (2024)
von: Menezes, Aaron Rock, et al.
Veröffentlicht: (2024)
Animated Public Furniture as an Interaction Mediator: Engaging Passersby In-the-Wild with Robotic Benches
von: Yu, Xinyan, et al.
Veröffentlicht: (2026)
von: Yu, Xinyan, et al.
Veröffentlicht: (2026)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
von: Ma, Jingtian, et al.
Veröffentlicht: (2025)
Do Vision-Language Models Understand Compound Nouns?
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
Generative Image Dynamics
von: Li, Zhengqi, et al.
Veröffentlicht: (2023)
von: Li, Zhengqi, et al.
Veröffentlicht: (2023)
ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
von: Chen, Hanyu, et al.
Veröffentlicht: (2026)
von: Chen, Hanyu, et al.
Veröffentlicht: (2026)
Seeing a Rose in Five Thousand Ways
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2022)
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2022)
Honey, I Shrunk the Arc de Triomphe!
von: Xiangli, Yuanbo, et al.
Veröffentlicht: (2026)
von: Xiangli, Yuanbo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing
von: Kani, Bharath Raj Nagoor, et al.
Veröffentlicht: (2026) -
Learning Feature Descriptors using Camera Pose Supervision
von: Wang, Qianqian, et al.
Veröffentlicht: (2020) -
UpFusion: Novel View Diffusion from Unposed Sparse View Observations
von: Kani, Bharath Raj Nagoor, et al.
Veröffentlicht: (2023) -
C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
von: Huang, Kuan Wei, et al.
Veröffentlicht: (2025) -
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
von: Revankar, Shreelekha, et al.
Veröffentlicht: (2025)