ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Ko, Dohwan, Kim, Sihyeon, Suh, Yumin, G, Vijay Kumar B., Yoon, Minseo, Chandraker, Manmohan, Kim, Hyunwoo J. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
di: Aich, Abhishek, et al.
Pubblicazione: (2024)
di: Aich, Abhishek, et al.
Pubblicazione: (2024)
LLM-Assist: Enhancing Closed-Loop Planning with Language-Based Reasoning
di: Sharan, S P, et al.
Pubblicazione: (2023)
di: Sharan, S P, et al.
Pubblicazione: (2023)
Generating Enhanced Negatives for Training Language-Based Object Detectors
di: Zhao, Shiyu, et al.
Pubblicazione: (2023)
di: Zhao, Shiyu, et al.
Pubblicazione: (2023)
Tuned Contrastive Learning
di: Animesh, Chaitanya, et al.
Pubblicazione: (2023)
di: Animesh, Chaitanya, et al.
Pubblicazione: (2023)
LLaMo: Large Language Model-based Molecular Graph Assistant
di: Park, Jinyoung, et al.
Pubblicazione: (2024)
di: Park, Jinyoung, et al.
Pubblicazione: (2024)
Taming Self-Training for Open-Vocabulary Object Detection
di: Zhao, Shiyu, et al.
Pubblicazione: (2023)
di: Zhao, Shiyu, et al.
Pubblicazione: (2023)
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
di: Ko, Dohwan, et al.
Pubblicazione: (2026)
di: Ko, Dohwan, et al.
Pubblicazione: (2026)
Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation
di: Yao, Manyi, et al.
Pubblicazione: (2024)
di: Yao, Manyi, et al.
Pubblicazione: (2024)
Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
di: Ko, Dohwan, et al.
Pubblicazione: (2025)
di: Ko, Dohwan, et al.
Pubblicazione: (2025)
ST-LINK: Spatially-Aware Large Language Models for Spatio-Temporal Forecasting
di: Jeon, Hyotaek, et al.
Pubblicazione: (2025)
di: Jeon, Hyotaek, et al.
Pubblicazione: (2025)
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning
di: Ke, Fucai, et al.
Pubblicazione: (2025)
di: Ke, Fucai, et al.
Pubblicazione: (2025)
DDMI: Domain-Agnostic Latent Diffusion Models for Synthesizing High-Quality Implicit Neural Representations
di: Park, Dogyun, et al.
Pubblicazione: (2024)
di: Park, Dogyun, et al.
Pubblicazione: (2024)
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
di: Choi, Joonmyung, et al.
Pubblicazione: (2026)
di: Choi, Joonmyung, et al.
Pubblicazione: (2026)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
di: Kalluri, Tarun, et al.
Pubblicazione: (2024)
di: Kalluri, Tarun, et al.
Pubblicazione: (2024)
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
di: Aich, Abhishek, et al.
Pubblicazione: (2026)
di: Aich, Abhishek, et al.
Pubblicazione: (2026)
Kinodynamic Task and Motion Planning using VLM-guided and Interleaved Sampling
di: Kwon, Minseo, et al.
Pubblicazione: (2025)
di: Kwon, Minseo, et al.
Pubblicazione: (2025)
UDA-Bench: Revisiting Common Assumptions in Unsupervised Domain Adaptation Using a Standardized Framework
di: Kalluri, Tarun, et al.
Pubblicazione: (2024)
di: Kalluri, Tarun, et al.
Pubblicazione: (2024)
Locally Orderless Images for Optimization in Differentiable Rendering
di: Mehta, Ishit, et al.
Pubblicazione: (2025)
di: Mehta, Ishit, et al.
Pubblicazione: (2025)
RAD-LAD: Rule and Language Grounded Autonomous Driving in Real-Time
di: Ghosh, Anurag, et al.
Pubblicazione: (2026)
di: Ghosh, Anurag, et al.
Pubblicazione: (2026)
Natural Language Declarative Prompting (NLD-P): A Modular Governance Method for Prompt Design Under Model Drift
di: Kim, Hyunwoo, et al.
Pubblicazione: (2026)
di: Kim, Hyunwoo, et al.
Pubblicazione: (2026)
Latent Bayesian Optimization via Autoregressive Normalizing Flows
di: Lee, Seunghun, et al.
Pubblicazione: (2025)
di: Lee, Seunghun, et al.
Pubblicazione: (2025)
Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action Localization
di: Lim, Geuntaek, et al.
Pubblicazione: (2024)
di: Lim, Geuntaek, et al.
Pubblicazione: (2024)
Latent Preference Modeling for Cross-Session Personalized Tool Calling
di: Yoon, Yejin, et al.
Pubblicazione: (2026)
di: Yoon, Yejin, et al.
Pubblicazione: (2026)
Spatio-Temporal Graphs Beyond Grids: Benchmark for Maritime Anomaly Detection
di: Kim, Jeehong, et al.
Pubblicazione: (2025)
di: Kim, Jeehong, et al.
Pubblicazione: (2025)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
di: Park, Jinho, et al.
Pubblicazione: (2026)
di: Park, Jinho, et al.
Pubblicazione: (2026)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
di: Nguyen-Nhu, Tinh-Anh, et al.
Pubblicazione: (2025)
di: Nguyen-Nhu, Tinh-Anh, et al.
Pubblicazione: (2025)
Constant Acceleration Flow
di: Park, Dogyun, et al.
Pubblicazione: (2024)
di: Park, Dogyun, et al.
Pubblicazione: (2024)
iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
di: Yao, Manyi, et al.
Pubblicazione: (2025)
di: Yao, Manyi, et al.
Pubblicazione: (2025)
DISPATCH: Distilling Selective Patches for Speech Enhancement
di: Kim, Dohwan, et al.
Pubblicazione: (2025)
di: Kim, Dohwan, et al.
Pubblicazione: (2025)
NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
di: Jain, Seemandhar, et al.
Pubblicazione: (2026)
di: Jain, Seemandhar, et al.
Pubblicazione: (2026)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
di: Li, Haodong, et al.
Pubblicazione: (2026)
di: Li, Haodong, et al.
Pubblicazione: (2026)
PhyCo: Learning Controllable Physical Priors for Generative Motion
di: Narayanan, Sriram, et al.
Pubblicazione: (2026)
di: Narayanan, Sriram, et al.
Pubblicazione: (2026)
VideoMamba: Spatio-Temporal Selective State Space Model
di: Park, Jinyoung, et al.
Pubblicazione: (2024)
di: Park, Jinyoung, et al.
Pubblicazione: (2024)
LANGTRAJ: Diffusion Model and Dataset for Language-Conditioned Trajectory Simulation
di: Chang, Wei-Jer, et al.
Pubblicazione: (2025)
di: Chang, Wei-Jer, et al.
Pubblicazione: (2025)
SLIP & ETHICS: Graduated Intervention for AI Emotional Companions
di: Kim, Minseo
Pubblicazione: (2026)
di: Kim, Minseo
Pubblicazione: (2026)
Small Language Models Learn Enhanced Reasoning Skills from Medical Textbooks
di: Kim, Hyunjae, et al.
Pubblicazione: (2024)
di: Kim, Hyunjae, et al.
Pubblicazione: (2024)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
di: Kim, Jongha, et al.
Pubblicazione: (2026)
di: Kim, Jongha, et al.
Pubblicazione: (2026)
Instantaneous Perception of Moving Objects in 3D
di: Liu, Di, et al.
Pubblicazione: (2024)
di: Liu, Di, et al.
Pubblicazione: (2024)
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
di: Yue, Lu, et al.
Pubblicazione: (2025)
di: Yue, Lu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
di: Aich, Abhishek, et al.
Pubblicazione: (2024) -
LLM-Assist: Enhancing Closed-Loop Planning with Language-Based Reasoning
di: Sharan, S P, et al.
Pubblicazione: (2023) -
Generating Enhanced Negatives for Training Language-Based Object Detectors
di: Zhao, Shiyu, et al.
Pubblicazione: (2023) -
Tuned Contrastive Learning
di: Animesh, Chaitanya, et al.
Pubblicazione: (2023) -
LLaMo: Large Language Model-based Molecular Graph Assistant
di: Park, Jinyoung, et al.
Pubblicazione: (2024)