GroupMamba: Efficient Group-Based Visual State Space Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Shaker, Abdelrahman, Wasim, Syed Talal, Khan, Salman, Gall, Juergen, Khan, Fahad Shahbaz |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
por: Shaker, Abdelrahman, et al.
Publicado: (2024)
por: Shaker, Abdelrahman, et al.
Publicado: (2024)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
por: Wasim, Syed Talal, et al.
Publicado: (2023)
por: Wasim, Syed Talal, et al.
Publicado: (2023)
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
por: Suleman, Hamid, et al.
Publicado: (2024)
por: Suleman, Hamid, et al.
Publicado: (2024)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
por: Shaker, Abdelrahman, et al.
Publicado: (2025)
por: Shaker, Abdelrahman, et al.
Publicado: (2025)
UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
por: Shaker, Abdelrahman, et al.
Publicado: (2022)
por: Shaker, Abdelrahman, et al.
Publicado: (2022)
Towards Evaluating the Robustness of Visual State Space Models
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
por: Kunhimon, Shahina, et al.
Publicado: (2023)
por: Kunhimon, Shahina, et al.
Publicado: (2023)
WorldCache: Content-Aware Caching for Accelerated Video World Models
por: Nawaz, Umair, et al.
Publicado: (2026)
por: Nawaz, Umair, et al.
Publicado: (2026)
Diversity Has Always Been There in Your Visual Autoregressive Models
por: Wang, Tong, et al.
Publicado: (2025)
por: Wang, Tong, et al.
Publicado: (2025)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
por: Yi, Jinhui, et al.
Publicado: (2024)
por: Yi, Jinhui, et al.
Publicado: (2024)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
por: Rasheed, Hanoona, et al.
Publicado: (2025)
por: Rasheed, Hanoona, et al.
Publicado: (2025)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
por: Ahmad, Ghazi Shazan, et al.
Publicado: (2025)
por: Ahmad, Ghazi Shazan, et al.
Publicado: (2025)
MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
por: Wasim, Syed Talal, et al.
Publicado: (2025)
por: Wasim, Syed Talal, et al.
Publicado: (2025)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
por: Thawakar, Omkar, et al.
Publicado: (2023)
por: Thawakar, Omkar, et al.
Publicado: (2023)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
por: Demidov, Dmitry, et al.
Publicado: (2025)
por: Demidov, Dmitry, et al.
Publicado: (2025)
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
por: Heakl, Ahmed, et al.
Publicado: (2026)
por: Heakl, Ahmed, et al.
Publicado: (2026)
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
por: Ashraf, Tajamul, et al.
Publicado: (2025)
por: Ashraf, Tajamul, et al.
Publicado: (2025)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
por: Chen, Shiming, et al.
Publicado: (2025)
por: Chen, Shiming, et al.
Publicado: (2025)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
por: Maaz, Muhammad, et al.
Publicado: (2025)
por: Maaz, Muhammad, et al.
Publicado: (2025)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
por: Maaz, Muhammad, et al.
Publicado: (2023)
por: Maaz, Muhammad, et al.
Publicado: (2023)
Dual Hyperspectral Mamba for Efficient Spectral Compressive Imaging
por: Dong, Jiahua, et al.
Publicado: (2024)
por: Dong, Jiahua, et al.
Publicado: (2024)
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
por: Li, Senmao, et al.
Publicado: (2025)
por: Li, Senmao, et al.
Publicado: (2025)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
por: Noman, Mubashir, et al.
Publicado: (2024)
por: Noman, Mubashir, et al.
Publicado: (2024)
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
por: Dudhane, Akshay, et al.
Publicado: (2024)
por: Dudhane, Akshay, et al.
Publicado: (2024)
Language Guided Domain Generalized Medical Image Segmentation
por: Kunhimon, Shahina, et al.
Publicado: (2024)
por: Kunhimon, Shahina, et al.
Publicado: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
por: Dharmasiri, Amaya, et al.
Publicado: (2024)
por: Dharmasiri, Amaya, et al.
Publicado: (2024)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
por: Thawakar, Omkar, et al.
Publicado: (2025)
por: Thawakar, Omkar, et al.
Publicado: (2025)
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
por: Ding, Bonan, et al.
Publicado: (2026)
por: Ding, Bonan, et al.
Publicado: (2026)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
por: Kumar, Komal, et al.
Publicado: (2025)
por: Kumar, Komal, et al.
Publicado: (2025)
AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation
por: Cui, Yuning, et al.
Publicado: (2024)
por: Cui, Yuning, et al.
Publicado: (2024)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
por: Munasinghe, Shehan, et al.
Publicado: (2024)
por: Munasinghe, Shehan, et al.
Publicado: (2024)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
por: Malik, Hashmat Shadab, et al.
Publicado: (2025)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
por: Chen, Shiming, et al.
Publicado: (2024)
por: Chen, Shiming, et al.
Publicado: (2024)
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
por: Chen, Shiming, et al.
Publicado: (2025)
por: Chen, Shiming, et al.
Publicado: (2025)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
Enhancing Novel Object Detection via Cooperative Foundational Models
por: Bharadwaj, Rohit, et al.
Publicado: (2023)
por: Bharadwaj, Rohit, et al.
Publicado: (2023)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
por: Mahmood, Ahmad, et al.
Publicado: (2024)
por: Mahmood, Ahmad, et al.
Publicado: (2024)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
por: Bharadwaj, Rohit, et al.
Publicado: (2024)
por: Bharadwaj, Rohit, et al.
Publicado: (2024)
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
por: Shaker, Abdelrahman, et al.
Publicado: (2026)
por: Shaker, Abdelrahman, et al.
Publicado: (2026)
PALO: A Polyglot Large Multimodal Model for 5B People
por: Maaz, Muhammad, et al.
Publicado: (2024)
por: Maaz, Muhammad, et al.
Publicado: (2024)
Ejemplares similares
-
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
por: Shaker, Abdelrahman, et al.
Publicado: (2024) -
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
por: Wasim, Syed Talal, et al.
Publicado: (2023) -
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
por: Suleman, Hamid, et al.
Publicado: (2024) -
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
por: Shaker, Abdelrahman, et al.
Publicado: (2025) -
UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
por: Shaker, Abdelrahman, et al.
Publicado: (2022)