SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Brown, Ellis, Ray, Arijit, Krishna, Ranjay, Girshick, Ross, Fergus, Rob, Xie, Saining |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
von: Brown, Ellis, et al.
Veröffentlicht: (2025)
Cambrian-S: Towards Spatial Supersensing in Video
von: Yang, Shusheng, et al.
Veröffentlicht: (2025)
von: Yang, Shusheng, et al.
Veröffentlicht: (2025)
V-IRL: Grounding Virtual Intelligence in Real Life
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
von: Tong, Shengbang, et al.
Veröffentlicht: (2026)
von: Tong, Shengbang, et al.
Veröffentlicht: (2026)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
von: Seyfioglu, Mehmet Saygin, et al.
Veröffentlicht: (2023)
von: Seyfioglu, Mehmet Saygin, et al.
Veröffentlicht: (2023)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
von: Li, Linjie, et al.
Veröffentlicht: (2025)
von: Li, Linjie, et al.
Veröffentlicht: (2025)
Cambrian-P: Pose-Grounded Video Understanding
von: Yang, Jihan, et al.
Veröffentlicht: (2026)
von: Yang, Jihan, et al.
Veröffentlicht: (2026)
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
Fast Encoding and Decoding for Implicit Video Representation
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Semi-Supervised Learning with Context-Conditional Generative Adversarial Networks
von: Denton, Remi, et al.
Veröffentlicht: (2016)
von: Denton, Remi, et al.
Veröffentlicht: (2016)
TrajTok: Learning Trajectory Tokens enables better Video Understanding
von: Zheng, Chenhao, et al.
Veröffentlicht: (2026)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2026)
Streaming Video Instruction Tuning
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
von: Fan, Xiang, et al.
Veröffentlicht: (2024)
von: Fan, Xiang, et al.
Veröffentlicht: (2024)
InstructionBench: An Instructional Video Understanding Benchmark
von: Wei, Haiwan, et al.
Veröffentlicht: (2025)
von: Wei, Haiwan, et al.
Veröffentlicht: (2025)
Stochastic Video Generation with a Learned Prior
von: Denton, Remi, et al.
Veröffentlicht: (2018)
von: Denton, Remi, et al.
Veröffentlicht: (2018)
Mull-Tokens: Modality-Agnostic Latent Thinking
von: Ray, Arijit, et al.
Veröffentlicht: (2025)
von: Ray, Arijit, et al.
Veröffentlicht: (2025)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
von: Xing, Bohao, et al.
Veröffentlicht: (2024)
von: Xing, Bohao, et al.
Veröffentlicht: (2024)
Osprey: Pixel Understanding with Visual Instruction Tuning
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
von: Hsieh, Cheng-Yu, et al.
Veröffentlicht: (2025)
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
LongViTU: Instruction Tuning for Long-Form Video Understanding
von: Wu, Rujie, et al.
Veröffentlicht: (2025)
von: Wu, Rujie, et al.
Veröffentlicht: (2025)
MindCube: Spatial Mental Modeling from Limited Views
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
RefTok: Reference-Based Tokenization for Video Generation
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
von: Fan, Xiang, et al.
Veröffentlicht: (2026)
von: Fan, Xiang, et al.
Veröffentlicht: (2026)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
von: Fei, Yang, et al.
Veröffentlicht: (2025)
von: Fei, Yang, et al.
Veröffentlicht: (2025)
Iterated Learning Improves Compositionality in Large Vision-Language Models
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
von: Liu, Ruyang, et al.
Veröffentlicht: (2023)
von: Liu, Ruyang, et al.
Veröffentlicht: (2023)
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
von: Li, Chenyu, et al.
Veröffentlicht: (2025)
von: Li, Chenyu, et al.
Veröffentlicht: (2025)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
Beyond Language Modeling: An Exploration of Multimodal Pretraining
von: Tong, Shengbang, et al.
Veröffentlicht: (2026)
von: Tong, Shengbang, et al.
Veröffentlicht: (2026)
REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
von: Leng, Xingjian, et al.
Veröffentlicht: (2025)
von: Leng, Xingjian, et al.
Veröffentlicht: (2025)
A Framework for Generating Semantically Ambiguous Images to Probe Human and Machine Perception
von: Hu, Yuqi, et al.
Veröffentlicht: (2026)
von: Hu, Yuqi, et al.
Veröffentlicht: (2026)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Self-Refining Video Sampling
von: Jang, Sangwon, et al.
Veröffentlicht: (2026)
von: Jang, Sangwon, et al.
Veröffentlicht: (2026)
Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning
von: Liu, Xixi, et al.
Veröffentlicht: (2026)
von: Liu, Xixi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
von: Brown, Ellis, et al.
Veröffentlicht: (2025) -
Cambrian-S: Towards Spatial Supersensing in Video
von: Yang, Shusheng, et al.
Veröffentlicht: (2025) -
V-IRL: Grounding Virtual Intelligence in Real Life
von: Yang, Jihan, et al.
Veröffentlicht: (2024) -
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
von: Tong, Shengbang, et al.
Veröffentlicht: (2026) -
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)