MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
Fuente:
arXiv
Saved in:
| Main Authors: | Wasim, Syed Talal, Suleman, Hamid, Zatsarynna, Olga, Naseer, Muzammal, Gall, Juergen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
by: Suleman, Hamid, et al.
Published: (2024)
by: Suleman, Hamid, et al.
Published: (2024)
Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2024)
by: Zatsarynna, Olga, et al.
Published: (2024)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2025)
by: Zatsarynna, Olga, et al.
Published: (2025)
Towards Generalizing Temporal Action Segmentation to Unseen Views
by: Bahrami, Emad, et al.
Published: (2025)
by: Bahrami, Emad, et al.
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
Looking into the Unknown: Exploring Action Discovery for Segmentation of Known and Unknown Actions
by: Spurio, Federico, et al.
Published: (2025)
by: Spurio, Federico, et al.
Published: (2025)
GroupMamba: Efficient Group-Based Visual State Space Model
by: Shaker, Abdelrahman, et al.
Published: (2024)
by: Shaker, Abdelrahman, et al.
Published: (2024)
Privacy-Preserving Semantic Segmentation from Ultra-Low-Resolution RGB Inputs
by: Huang, Xuying, et al.
Published: (2025)
by: Huang, Xuying, et al.
Published: (2025)
SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction
by: Pallotta, Enrico, et al.
Published: (2025)
by: Pallotta, Enrico, et al.
Published: (2025)
A Survey on Deep Learning Techniques for Action Anticipation
by: Zhong, Zeyun, et al.
Published: (2023)
by: Zhong, Zeyun, et al.
Published: (2023)
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
by: Bahrami, Emad, et al.
Published: (2026)
by: Bahrami, Emad, et al.
Published: (2026)
STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
by: Velayudhan, Divya, et al.
Published: (2025)
by: Velayudhan, Divya, et al.
Published: (2025)
Rethinking Memory Design in SAM-Based Visual Object Tracking
by: Alansari, Mohamad, et al.
Published: (2025)
by: Alansari, Mohamad, et al.
Published: (2025)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
by: Shaker, Abdelrahman, et al.
Published: (2024)
by: Shaker, Abdelrahman, et al.
Published: (2024)
Hierarchical Vector Quantization for Unsupervised Action Segmentation
by: Spurio, Federico, et al.
Published: (2024)
by: Spurio, Federico, et al.
Published: (2024)
Makeup-Guided Facial Privacy Protection via Untrained Neural Network Priors
by: Shamshad, Fahad, et al.
Published: (2024)
by: Shamshad, Fahad, et al.
Published: (2024)
Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
by: Gökay, Uzay, et al.
Published: (2025)
by: Gökay, Uzay, et al.
Published: (2025)
Analysing the Robustness of Vision-Language-Models to Common Corruptions
by: Usama, Muhammad, et al.
Published: (2025)
by: Usama, Muhammad, et al.
Published: (2025)
Language Guided Domain Generalized Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2024)
by: Kunhimon, Shahina, et al.
Published: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
by: Dharmasiri, Amaya, et al.
Published: (2024)
by: Dharmasiri, Amaya, et al.
Published: (2024)
DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation
by: Assefa, Maregu, et al.
Published: (2025)
by: Assefa, Maregu, et al.
Published: (2025)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
by: Hussein, Noor, et al.
Published: (2024)
by: Hussein, Noor, et al.
Published: (2024)
Action-Guided Attention for Video Action Anticipation
by: Tai, Tsung-Ming, et al.
Published: (2026)
by: Tai, Tsung-Ming, et al.
Published: (2026)
LC-SLab -- An Object-based Deep Learning Framework for Large-scale Land Cover Classification from Satellite Imagery and Sparse In-situ Labels
by: Leonhardt, Johannes, et al.
Published: (2025)
by: Leonhardt, Johannes, et al.
Published: (2025)
Identifying Spatio-Temporal Drivers of Extreme Events
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
QueryMamba: A Mamba-Based Encoder-Decoder Architecture with a Statistical Verb-Noun Interaction Module for Video Action Forecasting @ Ego4D Long-Term Action Anticipation Challenge 2024
by: Zhong, Zeyun, et al.
Published: (2024)
by: Zhong, Zeyun, et al.
Published: (2024)
MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
by: Gani, Hanan, et al.
Published: (2024)
by: Gani, Hanan, et al.
Published: (2024)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
by: Mahmood, Ahmad, et al.
Published: (2024)
by: Mahmood, Ahmad, et al.
Published: (2024)
LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts
by: Gani, Hanan, et al.
Published: (2023)
by: Gani, Hanan, et al.
Published: (2023)
STEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion Models
by: Srivatsan, Koushik, et al.
Published: (2024)
by: Srivatsan, Koushik, et al.
Published: (2024)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
by: Bharadwaj, Rohit, et al.
Published: (2024)
by: Bharadwaj, Rohit, et al.
Published: (2024)
Y-CA-Net: A Convolutional Attention Based Network for Volumetric Medical Image Segmentation
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
CamC2V: Context-aware Controllable Video Generation
by: Denninger, Luis, et al.
Published: (2025)
by: Denninger, Luis, et al.
Published: (2025)
LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs
by: Bozorgtabar, Behzad, et al.
Published: (2026)
by: Bozorgtabar, Behzad, et al.
Published: (2026)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
by: Malik, Hashmat Shadab, et al.
Published: (2025)
by: Malik, Hashmat Shadab, et al.
Published: (2025)
Multimodal Large Models Are Effective Action Anticipators
by: Wang, Binglu, et al.
Published: (2025)
by: Wang, Binglu, et al.
Published: (2025)
Video Panels for Long Video Understanding
by: Doorenbos, Lars, et al.
Published: (2025)
by: Doorenbos, Lars, et al.
Published: (2025)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
by: Watawana, Hasindri, et al.
Published: (2024)
by: Watawana, Hasindri, et al.
Published: (2024)
Similar Items
-
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
by: Suleman, Hamid, et al.
Published: (2024) -
Gated Temporal Diffusion for Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2024) -
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
by: Yi, Jinhui, et al.
Published: (2024) -
MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Anticipation
by: Zatsarynna, Olga, et al.
Published: (2025) -
Towards Generalizing Temporal Action Segmentation to Unseen Views
by: Bahrami, Emad, et al.
Published: (2025)