StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
Fuente:
arXiv
Saved in:
| Main Authors: | Siddiqui, Nyle, Gupta, Rohit, Swetha, Sirnam, Shah, Mubarak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TimeLogic: A Temporal Logic Benchmark for Video QA
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
by: Narnaware, Vishal, et al.
Published: (2025)
by: Narnaware, Vishal, et al.
Published: (2025)
The Telephone Game: Evaluating Semantic Drift in Unified Models
by: Mollah, Sabbir, et al.
Published: (2025)
by: Mollah, Sabbir, et al.
Published: (2025)
Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
by: Kim, Younggun, et al.
Published: (2025)
by: Kim, Younggun, et al.
Published: (2025)
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
by: Shatwell, David G., et al.
Published: (2026)
by: Shatwell, David G., et al.
Published: (2026)
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
by: Shatwell, David G., et al.
Published: (2025)
by: Shatwell, David G., et al.
Published: (2025)
DLCR: A Generative Data Expansion Framework via Diffusion for Clothes-Changing Person Re-ID
by: Siddiqui, Nyle, et al.
Published: (2024)
by: Siddiqui, Nyle, et al.
Published: (2024)
VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2026)
by: Kulkarni, Parth Parag, et al.
Published: (2026)
Cross-View Open-Vocabulary Object Detection in Aerial Imagery
by: Kini, Jyoti, et al.
Published: (2025)
by: Kini, Jyoti, et al.
Published: (2025)
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
by: Dave, Ishan Rajendrakumar, et al.
Published: (2024)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
by: Swetha, Sirnam, et al.
Published: (2024)
by: Swetha, Sirnam, et al.
Published: (2024)
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
by: Qiu, Yicheng, et al.
Published: (2026)
by: Qiu, Yicheng, et al.
Published: (2026)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023)
by: Rajendiran, Ramanathan, et al.
Published: (2023)
ALBAR: Adversarial Learning approach to mitigate Biases in Action Recognition
by: Fioresi, Joseph, et al.
Published: (2025)
by: Fioresi, Joseph, et al.
Published: (2025)
Flexible and Efficient Spatio-Temporal Transformer for Sequential Visual Place Recognition
by: Kiu, Yu, et al.
Published: (2025)
by: Kiu, Yu, et al.
Published: (2025)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
by: Gupta, Animesh, et al.
Published: (2025)
by: Gupta, Animesh, et al.
Published: (2025)
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
by: So, Yerim, et al.
Published: (2026)
by: So, Yerim, et al.
Published: (2026)
Minimalistic Video Saliency Prediction via Efficient Decoder & Spatio Temporal Action Cues
by: Girmaji, Rohit, et al.
Published: (2025)
by: Girmaji, Rohit, et al.
Published: (2025)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
ViLL-E: Video LLM Embeddings for Retrieval
by: Gupta, Rohit, et al.
Published: (2026)
by: Gupta, Rohit, et al.
Published: (2026)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
by: Kim, Yehna, et al.
Published: (2025)
by: Kim, Yehna, et al.
Published: (2025)
UniSTFormer: Unified Spatio-Temporal Lightweight Transformer for Efficient Skeleton-Based Action Recognition
by: Wu, Wenhan, et al.
Published: (2025)
by: Wu, Wenhan, et al.
Published: (2025)
UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video Modeling
by: Li, Peiming, et al.
Published: (2025)
by: Li, Peiming, et al.
Published: (2025)
Open-Vocabulary Spatio-Temporal Action Detection
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
by: Li, Kunyang, et al.
Published: (2026)
by: Li, Kunyang, et al.
Published: (2026)
D$^2$ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition
by: Pei, Wenjie, et al.
Published: (2023)
by: Pei, Wenjie, et al.
Published: (2023)
DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition
by: Ullah, Hayat, et al.
Published: (2025)
by: Ullah, Hayat, et al.
Published: (2025)
Micro-DualNet: Dual-Path Spatio-Temporal Network for Micro-Action Recognition
by: Chappa, Naga VS Raviteja, et al.
Published: (2026)
by: Chappa, Naga VS Raviteja, et al.
Published: (2026)
MSSTNet: A Multi-Scale Spatio-Temporal CNN-Transformer Network for Dynamic Facial Expression Recognition
by: Wang, Linhuang, et al.
Published: (2024)
by: Wang, Linhuang, et al.
Published: (2024)
CityGuessr: City-Level Video Geo-Localization on a Global Scale
by: Kulkarni, Parth Parag, et al.
Published: (2024)
by: Kulkarni, Parth Parag, et al.
Published: (2024)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
by: Miao, Bo, et al.
Published: (2024)
by: Miao, Bo, et al.
Published: (2024)
Leveraging Pre-Trained Visual Models for AI-Generated Video Detection
by: Veeramachaneni, Keerthi, et al.
Published: (2025)
by: Veeramachaneni, Keerthi, et al.
Published: (2025)
SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action Recognition
by: Huang, Wenbo, et al.
Published: (2024)
by: Huang, Wenbo, et al.
Published: (2024)
Surgical Triplet Recognition via Diffusion Model
by: Liu, Daochang, et al.
Published: (2024)
by: Liu, Daochang, et al.
Published: (2024)
Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
by: Qu, Hongyu, et al.
Published: (2026)
by: Qu, Hongyu, et al.
Published: (2026)
Friends Across Time: Multi-Scale Action Segmentation Transformer for Surgical Phase Recognition
by: Zhang, Bokai, et al.
Published: (2024)
by: Zhang, Bokai, et al.
Published: (2024)
Open Vocabulary Multi-Label Video Classification
by: Gupta, Rohit, et al.
Published: (2024)
by: Gupta, Rohit, et al.
Published: (2024)
Leveraging Temporal Contextualization for Video Action Recognition
by: Kim, Minji, et al.
Published: (2024)
by: Kim, Minji, et al.
Published: (2024)
Similar Items
-
TimeLogic: A Temporal Logic Benchmark for Video QA
by: Swetha, Sirnam, et al.
Published: (2025) -
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025) -
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
by: Narnaware, Vishal, et al.
Published: (2025) -
The Telephone Game: Evaluating Semantic Drift in Unified Models
by: Mollah, Sabbir, et al.
Published: (2025) -
Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
by: Kim, Younggun, et al.
Published: (2025)