MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sinha, Arkaprava, Raj, Monish Soundar, Wang, Pu, Helmy, Ahmed, Le, Hieu, Das, Srijan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping
von: Bondurant, Weston, et al.
Veröffentlicht: (2025)
von: Bondurant, Weston, et al.
Veröffentlicht: (2025)
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025)
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025)
Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier
von: Howlader, Prantik, et al.
Veröffentlicht: (2024)
von: Howlader, Prantik, et al.
Veröffentlicht: (2024)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2024)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2024)
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
von: Feng, Yue, et al.
Veröffentlicht: (2026)
von: Feng, Yue, et al.
Veröffentlicht: (2026)
Mamba-OTR: a Mamba-based Solution for Online Take and Release Detection from Untrimmed Egocentric Video
von: Catinello, Alessandro Sebastiano, et al.
Veröffentlicht: (2025)
von: Catinello, Alessandro Sebastiano, et al.
Veröffentlicht: (2025)
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
von: Reilly, Dominick, et al.
Veröffentlicht: (2024)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
von: Peng, Liyang, et al.
Veröffentlicht: (2025)
von: Peng, Liyang, et al.
Veröffentlicht: (2025)
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
von: Tang, Yin, et al.
Veröffentlicht: (2024)
von: Tang, Yin, et al.
Veröffentlicht: (2024)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
von: Chen, Brian, et al.
Veröffentlicht: (2023)
von: Chen, Brian, et al.
Veröffentlicht: (2023)
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
von: Feng, Yue, et al.
Veröffentlicht: (2025)
von: Feng, Yue, et al.
Veröffentlicht: (2025)
Multi-Stage Boundary-Aware Transformer Network for Action Segmentation in Untrimmed Surgical Videos
von: Shuvo, Rezowan, et al.
Veröffentlicht: (2025)
von: Shuvo, Rezowan, et al.
Veröffentlicht: (2025)
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
von: Yang, Min, et al.
Veröffentlicht: (2023)
von: Yang, Min, et al.
Veröffentlicht: (2023)
Cutup and Detect: Human Fall Detection on Cutup Untrimmed Videos Using a Large Foundational Video Understanding Model
von: Grutschus, Till, et al.
Veröffentlicht: (2024)
von: Grutschus, Till, et al.
Veröffentlicht: (2024)
MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild
von: Saleem, Muhammad Usama, et al.
Veröffentlicht: (2024)
von: Saleem, Muhammad Usama, et al.
Veröffentlicht: (2024)
MonoMSK: Monocular 3D Musculoskeletal Dynamics Estimation
von: Koleini, Farnoosh, et al.
Veröffentlicht: (2025)
von: Koleini, Farnoosh, et al.
Veröffentlicht: (2025)
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
von: Howlader, Prantik, et al.
Veröffentlicht: (2025)
von: Howlader, Prantik, et al.
Veröffentlicht: (2025)
Localizing Moments of Actions in Untrimmed Videos of Infants with Autism Spectrum Disorder
von: Helvaci, Halil Ismail, et al.
Veröffentlicht: (2024)
von: Helvaci, Halil Ismail, et al.
Veröffentlicht: (2024)
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
von: Nguyen, Hong, et al.
Veröffentlicht: (2025)
von: Nguyen, Hong, et al.
Veröffentlicht: (2025)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
von: Fateh, Fawad Javed, et al.
Veröffentlicht: (2024)
von: Fateh, Fawad Javed, et al.
Veröffentlicht: (2024)
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
von: Ying, Xinru, et al.
Veröffentlicht: (2025)
von: Ying, Xinru, et al.
Veröffentlicht: (2025)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
von: An, Joungbin, et al.
Veröffentlicht: (2025)
von: An, Joungbin, et al.
Veröffentlicht: (2025)
Hallucination Mitigation Prompts Long-term Video Understanding
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
MLP: Motion Label Prior for Temporal Sentence Localization in Untrimmed 3D Human Motions
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
von: Ye, Jinhui, et al.
Veröffentlicht: (2025)
von: Ye, Jinhui, et al.
Veröffentlicht: (2025)
DensePercept-NCSSD: Vision Mamba towards Real-time Dense Visual Perception with Non-Causal State Space Duality
von: Anand, Tushar, et al.
Veröffentlicht: (2025)
von: Anand, Tushar, et al.
Veröffentlicht: (2025)
Fusion-SSAT: Unleashing the Potential of Self-supervised Auxiliary Task by Feature Fusion for Generalized Deepfake Detection
von: Reddy, Shukesh, et al.
Veröffentlicht: (2026)
von: Reddy, Shukesh, et al.
Veröffentlicht: (2026)
VideoMamba: State Space Model for Efficient Video Understanding
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
Learning Human Motion with Temporally Conditional Mamba
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang, et al.
Veröffentlicht: (2025)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
von: Truong, Chau, et al.
Veröffentlicht: (2025)
von: Truong, Chau, et al.
Veröffentlicht: (2025)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
BioPose: Biomechanically-accurate 3D Pose Estimation from Monocular Videos
von: Koleini, Farnoosh, et al.
Veröffentlicht: (2025)
von: Koleini, Farnoosh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DiffSwap++: 3D Latent-Controlled Diffusion for Identity-Preserving Face Swapping
von: Bondurant, Weston, et al.
Veröffentlicht: (2025) -
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
von: Sinha, Arkaprava, et al.
Veröffentlicht: (2025) -
Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier
von: Howlader, Prantik, et al.
Veröffentlicht: (2024) -
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2024) -
Explainable Forensics of Manipulated Segments in Untrimmed Long Videos
von: Feng, Yue, et al.
Veröffentlicht: (2026)