STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Bahrami, Emad, Zatsarynna, Olga, Pathak, Parth, Sengupta, Sunando, Gall, Juergen, Fayyaz, Mohsen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Generalizing Temporal Action Segmentation to Unseen Views
by: Bahrami, Emad, et al.
Published: (2025)
by: Bahrami, Emad, et al.
Published: (2025)
MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2025)
by: Zatsarynna, Olga, et al.
Published: (2025)
Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2024)
by: Zatsarynna, Olga, et al.
Published: (2024)
Looking into the Unknown: Exploring Action Discovery for Segmentation of Known and Unknown Actions
by: Spurio, Federico, et al.
Published: (2025)
by: Spurio, Federico, et al.
Published: (2025)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
by: Hannan, Tanveer, et al.
Published: (2025)
by: Hannan, Tanveer, et al.
Published: (2025)
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
Hierarchical Vector Quantization for Unsupervised Action Segmentation
by: Spurio, Federico, et al.
Published: (2024)
by: Spurio, Federico, et al.
Published: (2024)
Privacy-Preserving Semantic Segmentation from Ultra-Low-Resolution RGB Inputs
by: Huang, Xuying, et al.
Published: (2025)
by: Huang, Xuying, et al.
Published: (2025)
MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
by: Wasim, Syed Talal, et al.
Published: (2025)
by: Wasim, Syed Talal, et al.
Published: (2025)
Sequence-Adaptive Video Prediction in Continuous Streams using Diffusion Noise Optimization
by: Azar, Sina Mokhtarzadeh, et al.
Published: (2025)
by: Azar, Sina Mokhtarzadeh, et al.
Published: (2025)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
Video Panels for Long Video Understanding
by: Doorenbos, Lars, et al.
Published: (2025)
by: Doorenbos, Lars, et al.
Published: (2025)
CamC2V: Context-aware Controllable Video Generation
by: Denninger, Luis, et al.
Published: (2025)
by: Denninger, Luis, et al.
Published: (2025)
Latent Directions: A Simple Pathway to Bias Mitigation in Generative AI
by: Olmos, Carolina Lopez, et al.
Published: (2024)
by: Olmos, Carolina Lopez, et al.
Published: (2024)
LC-SLab -- An Object-based Deep Learning Framework for Large-scale Land Cover Classification from Satellite Imagery and Sparse In-situ Labels
by: Leonhardt, Johannes, et al.
Published: (2025)
by: Leonhardt, Johannes, et al.
Published: (2025)
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
by: Suleman, Hamid, et al.
Published: (2024)
by: Suleman, Hamid, et al.
Published: (2024)
FlowNar: Scalable Streaming Narration for Long-Form Videos
by: Zhong, Zeyun, et al.
Published: (2026)
by: Zhong, Zeyun, et al.
Published: (2026)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
by: Frahm, Noah, et al.
Published: (2025)
by: Frahm, Noah, et al.
Published: (2025)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Occlusion Handling in 3D Human Pose Estimation with Perturbed Positional Encoding
by: Azizi, Niloofar, et al.
Published: (2024)
by: Azizi, Niloofar, et al.
Published: (2024)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
by: Ma, Haodi, et al.
Published: (2025)
by: Ma, Haodi, et al.
Published: (2025)
Identifying Spatio-Temporal Drivers of Extreme Events
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
Learning a Neural Association Network for Self-supervised Multi-Object Tracking
by: Li, Shuai, et al.
Published: (2024)
by: Li, Shuai, et al.
Published: (2024)
Enhancing Video-Based Robot Failure Detection Using Task Knowledge
by: Thoduka, Santosh, et al.
Published: (2025)
by: Thoduka, Santosh, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Agentic Keyframe Search for Video Question Answering
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
by: Rasekh, Ali, et al.
Published: (2025)
by: Rasekh, Ali, et al.
Published: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
Leveraging LLMs with Iterative Loop Structure for Enhanced Social Intelligence in Video Question Answering
by: Mori, Erika, et al.
Published: (2025)
by: Mori, Erika, et al.
Published: (2025)
FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
by: Zemskova, Tatiana, et al.
Published: (2026)
by: Zemskova, Tatiana, et al.
Published: (2026)
STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes
by: Ishihara, Keishi, et al.
Published: (2025)
by: Ishihara, Keishi, et al.
Published: (2025)
TerraQ: Spatiotemporal Question-Answering on Satellite Image Archives
by: Kefalidis, Sergios-Anestis, et al.
Published: (2025)
by: Kefalidis, Sergios-Anestis, et al.
Published: (2025)
Narrative Aligned Long Form Video Question Answering
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
ViLA: Efficient Video-Language Alignment for Video Question Answering
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
by: Oh, Ju-Young
Published: (2025)
by: Oh, Ju-Young
Published: (2025)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
by: Song, Enxin, et al.
Published: (2024)
by: Song, Enxin, et al.
Published: (2024)
Similar Items
-
Towards Generalizing Temporal Action Segmentation to Unseen Views
by: Bahrami, Emad, et al.
Published: (2025) -
MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2025) -
Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2024) -
Looking into the Unknown: Exploring Action Discovery for Segmentation of Known and Unknown Actions
by: Spurio, Federico, et al.
Published: (2025) -
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
by: Hannan, Tanveer, et al.
Published: (2025)