SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ray, Arijit, Duan, Jiafei, Brown, Ellis, Tan, Reuben, Bashkirova, Dina, Hendrix, Rose, Ehsani, Kiana, Kembhavi, Aniruddha, Plummer, Bryan A., Krishna, Ranjay, Zeng, Kuo-Hao, Saenko, Kate |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Harmonic Mobile Manipulation
by: Yang, Ruihan, et al.
Published: (2023)
by: Yang, Ruihan, et al.
Published: (2023)
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024)
by: Eftekhar, Ainaz, et al.
Published: (2024)
FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
by: Hu, Jiaheng, et al.
Published: (2024)
by: Hu, Jiaheng, et al.
Published: (2024)
PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
by: Zeng, Kuo-Hao, et al.
Published: (2024)
by: Zeng, Kuo-Hao, et al.
Published: (2024)
Mull-Tokens: Modality-Agnostic Latent Thinking
by: Ray, Arijit, et al.
Published: (2025)
by: Ray, Arijit, et al.
Published: (2025)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)
by: Eftekhar, Ainaz, et al.
Published: (2023)
GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
by: Deshpande, Abhay, et al.
Published: (2025)
by: Deshpande, Abhay, et al.
Published: (2025)
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024)
by: Zheng, Chenhao, et al.
Published: (2024)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
PaintBench: Deterministic Evaluation of Precise Visual Editing
by: Xu, Kai, et al.
Published: (2026)
by: Xu, Kai, et al.
Published: (2026)
Line Segment Clipping using Quadrilateral Concavity and Convexity
by: Ray, Bimal Kumar
Published: (2026)
by: Ray, Bimal Kumar
Published: (2026)
SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
by: Ehsani, Kiana, et al.
Published: (2023)
by: Ehsani, Kiana, et al.
Published: (2023)
Breaking the Assistant Mold: Modeling Behavioral Variation in LLM Based Procedural Character Generation
by: Qraitem, Maan, et al.
Published: (2026)
by: Qraitem, Maan, et al.
Published: (2026)
Tell Me What's Next: Textual Foresight for Generic UI Representations
by: Burns, Andrea, et al.
Published: (2024)
by: Burns, Andrea, et al.
Published: (2024)
From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition
by: Qraitem, Maan, et al.
Published: (2023)
by: Qraitem, Maan, et al.
Published: (2023)
Reinforcement Learning-Enhanced Procedural Generation for Dynamic Narrative-Driven AR Experiences
by: Joshi, Aniruddha Srinivas
Published: (2025)
by: Joshi, Aniruddha Srinivas
Published: (2025)
Winding Number Features for Vector Sketch Colorization
by: Daniel Scrivener, et al.
Published: (2024)
by: Daniel Scrivener, et al.
Published: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
GS-Share: Enabling High-fidelity Map Sharing with Incremental Gaussian Splatting
by: Zhang, Xinran, et al.
Published: (2025)
by: Zhang, Xinran, et al.
Published: (2025)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
by: Zheng, Chenhao, et al.
Published: (2025)
by: Zheng, Chenhao, et al.
Published: (2025)
Modeling and Measuring the Chart Communication Recall Process
by: A. Arunkumar, et al.
Published: (2025)
by: A. Arunkumar, et al.
Published: (2025)
Engagement vs. Understanding: Comparing Immersive Virtual Reality and Desktop Displays for Climate Data Visualization
by: A. Arunkumar, et al.
Published: (2026)
by: A. Arunkumar, et al.
Published: (2026)
SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
by: Brown, Ellis, et al.
Published: (2025)
by: Brown, Ellis, et al.
Published: (2025)
More Bang For Your Buck(et): Fast and Space-efficient Hardware-accelerated Coarse-granular Indexing on GPUs
by: Henneberg, Justus, et al.
Published: (2024)
by: Henneberg, Justus, et al.
Published: (2024)
Hybrid Explicit Representation for Ultra-Realistic Head Avatars
by: Cai, Hongrui, et al.
Published: (2024)
by: Cai, Hongrui, et al.
Published: (2024)
ReverBERT: A State Space Model for Efficient Text-Driven Speech Style Transfer
by: Brown, Michael, et al.
Published: (2025)
by: Brown, Michael, et al.
Published: (2025)
SLANT: Spurious Logo ANalysis Toolkit
by: Qraitem, Maan, et al.
Published: (2024)
by: Qraitem, Maan, et al.
Published: (2024)
Web Artifact Attacks Disrupt Vision Language Models
by: Qraitem, Maan, et al.
Published: (2025)
by: Qraitem, Maan, et al.
Published: (2025)
Multi-axis Analysis of Image Manipulation Localization
by: Nichols, Keanu, et al.
Published: (2026)
by: Nichols, Keanu, et al.
Published: (2026)
ConfEviSurrogate: A Conformalized Evidential Surrogate Model for Uncertainty Quantification
by: Duan, Yuhan, et al.
Published: (2025)
by: Duan, Yuhan, et al.
Published: (2025)
Negative Token Merging: Image-based Adversarial Feature Guidance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
On the Content Bias in Fréchet Video Distance
by: Ge, Songwei, et al.
Published: (2024)
by: Ge, Songwei, et al.
Published: (2024)
A Fast Unsupervised Scheme for Polygonal Approximation
by: Ray, Bimal Kumar
Published: (2025)
by: Ray, Bimal Kumar
Published: (2025)
SCARF: Scalable Continual Learning Framework for Memory‐efficient Multiple Neural Radiance Fields
by: Yuze Wang, et al.
Published: (2024)
by: Yuze Wang, et al.
Published: (2024)
High contrast holography through dual modulation
by: Kabuli, Leyla, et al.
Published: (2024)
by: Kabuli, Leyla, et al.
Published: (2024)
OP-LoRA: The Blessing of Dimensionality
by: Teterwak, Piotr, et al.
Published: (2024)
by: Teterwak, Piotr, et al.
Published: (2024)
Neural Product Importance Sampling via Warp Composition
by: Litalien, Joey, et al.
Published: (2024)
by: Litalien, Joey, et al.
Published: (2024)
Choreographing the Digital Canvas: A Machine Learning Approach to Artistic Performance
by: Peng, Siyuan, et al.
Published: (2024)
by: Peng, Siyuan, et al.
Published: (2024)
A Synthetic Benchmarking Pipeline to Compare Camera Calibration Algorithms
by: Ray, Lala Shakti Swarup, et al.
Published: (2023)
by: Ray, Lala Shakti Swarup, et al.
Published: (2023)
Monocular Facial Appearance Capture in the Wild
by: Xu, Yingyan, et al.
Published: (2024)
by: Xu, Yingyan, et al.
Published: (2024)
Similar Items
-
Harmonic Mobile Manipulation
by: Yang, Ruihan, et al.
Published: (2023) -
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024) -
FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
by: Hu, Jiaheng, et al.
Published: (2024) -
PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators
by: Zeng, Kuo-Hao, et al.
Published: (2024) -
Mull-Tokens: Modality-Agnostic Latent Thinking
by: Ray, Arijit, et al.
Published: (2025)