Can Graphs Help Vision SSMs See Better?
Fuente:
arXiv
Saved in:
| Main Authors: | Parikh, Dhruv, Ramachandran, Anvitha, Fan, Haoyang, Munir, Mustafa, Kannan, Rajgopal, Prasanna, Viktor |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
by: Ramachandran, Anvitha, et al.
Published: (2026)
by: Ramachandran, Anvitha, et al.
Published: (2026)
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
Latent Denoising Improves Visual Alignment in Large Multimodal Models
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
by: Parikh, Dhruv, et al.
Published: (2025)
by: Parikh, Dhruv, et al.
Published: (2025)
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025)
by: Ramachandran, Anvitha, et al.
Published: (2025)
VTR: An Optimized Vision Transformer for SAR ATR Acceleration on FPGA
by: Wickramasinghe, Sachini, et al.
Published: (2024)
by: Wickramasinghe, Sachini, et al.
Published: (2024)
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
by: Parikh, Dhruv, et al.
Published: (2024)
by: Parikh, Dhruv, et al.
Published: (2024)
Uncertainty-Aware SAR ATR: Defending Against Adversarial Attacks via Bayesian Neural Networks
by: Ye, Tian, et al.
Published: (2024)
by: Ye, Tian, et al.
Published: (2024)
ImageHD: Energy-Efficient On-Device Continual Learning of Visual Representations via Hyperdimensional Computing
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
Adversarial Training in Low-Label Regimes with Margin-Based Interpolation
by: Ye, Tian, et al.
Published: (2024)
by: Ye, Tian, et al.
Published: (2024)
Studying the Effects of Self-Attention on SAR Automatic Target Recognition
by: Fein-Ashley, Jacob, et al.
Published: (2024)
by: Fein-Ashley, Jacob, et al.
Published: (2024)
GCV-Turbo: End-to-end Acceleration of GNN-based Computer Vision Tasks on FPGA
by: Zhang, Bingyi, et al.
Published: (2024)
by: Zhang, Bingyi, et al.
Published: (2024)
FACTUAL: A Novel Framework for Contrastive Learning Based Robust SAR Image Classification
by: Wang, Xu, et al.
Published: (2024)
by: Wang, Xu, et al.
Published: (2024)
Primitive-Driven Acceleration of Hyperdimensional Computing for Real-Time Image Classification
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
A Single Graph Convolution Is All You Need: Efficient Grayscale Image Classification
by: Fein-Ashley, Jacob, et al.
Published: (2024)
by: Fein-Ashley, Jacob, et al.
Published: (2024)
PAHD: Perception-Action based Human Decision Making using Explainable Graph Neural Networks on SAR Images
by: Wijeratne, Sasindu, et al.
Published: (2024)
by: Wijeratne, Sasindu, et al.
Published: (2024)
Diffusion Feedback Helps CLIP See Better
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Benchmarking Deep Learning Classifiers for SAR Automatic Target Recognition
by: Fein-Ashley, Jacob, et al.
Published: (2023)
by: Fein-Ashley, Jacob, et al.
Published: (2023)
Scaling Graph Convolutions for Mobile Vision
by: Avery, William, et al.
Published: (2024)
by: Avery, William, et al.
Published: (2024)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
by: Nath, Mriganka, et al.
Published: (2026)
by: Nath, Mriganka, et al.
Published: (2026)
Arcee: Differentiable Recurrent State Chain for Generative Vision Modeling with Mamba SSMs
by: Chavan, Jitesh, et al.
Published: (2025)
by: Chavan, Jitesh, et al.
Published: (2025)
Vision-Language Models Can't See the Obvious
by: Dahou, Yasser, et al.
Published: (2025)
by: Dahou, Yasser, et al.
Published: (2025)
OCTOPUS: Enhancing the Spatial-Awareness of Vision SSMs with Multi-Dimensional Scans and Traversal Selection
by: Mahatha, Kunal, et al.
Published: (2026)
by: Mahatha, Kunal, et al.
Published: (2026)
Information Extraction from Unstructured data using Augmented-AI and Computer Vision
by: Parikh, Aditya
Published: (2023)
by: Parikh, Aditya
Published: (2023)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
GreedyViG: Dynamic Axial Graph Construction for Efficient Vision GNNs
by: Munir, Mustafa, et al.
Published: (2024)
by: Munir, Mustafa, et al.
Published: (2024)
EDmamba: Rethinking Efficient Event Denoising with Spatiotemporal Decoupled SSMs
by: Ruan, Ciyu, et al.
Published: (2025)
by: Ruan, Ciyu, et al.
Published: (2025)
Aligning What EEG Can See: Structural Representations for Brain-Vision Matching
by: Tang, Jingyi, et al.
Published: (2026)
by: Tang, Jingyi, et al.
Published: (2026)
Mixup Helps Understanding Multimodal Video Better
by: Ma, Xiaoyu, et al.
Published: (2025)
by: Ma, Xiaoyu, et al.
Published: (2025)
StableMamba: Distillation-free Scaling of Large SSMs for Images and Videos
by: Suleman, Hamid, et al.
Published: (2024)
by: Suleman, Hamid, et al.
Published: (2024)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
Multi-Scale High-Resolution Logarithmic Grapher Module for Efficient Vision GNNs
by: Munir, Mustafa, et al.
Published: (2025)
by: Munir, Mustafa, et al.
Published: (2025)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
by: Ding, Ning, et al.
Published: (2023)
by: Ding, Ning, et al.
Published: (2023)
TabConv: Low-Computation CNN Inference via Table Lookups
by: Gupta, Neelesh, et al.
Published: (2024)
by: Gupta, Neelesh, et al.
Published: (2024)
Bridging the Geometry Mismatch: Frequency-Aware Anisotropic Serialization for Thin-Structure SSMs
by: Bai, Jin, et al.
Published: (2026)
by: Bai, Jin, et al.
Published: (2026)
MambaCSR: Dual-Interleaved Scanning for Compressed Image Super-Resolution With SSMs
by: Ren, Yulin, et al.
Published: (2024)
by: Ren, Yulin, et al.
Published: (2024)
AdaptViG: Adaptive Vision GNN with Exponential Decay Gating
by: Munir, Mustafa, et al.
Published: (2025)
by: Munir, Mustafa, et al.
Published: (2025)
Multimodal Language Models See Better When They Look Shallower
by: Chen, Haoran, et al.
Published: (2025)
by: Chen, Haoran, et al.
Published: (2025)
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
Similar Items
-
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
by: Ramachandran, Anvitha, et al.
Published: (2026) -
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models
by: Parikh, Dhruv, et al.
Published: (2026) -
Latent Denoising Improves Visual Alignment in Large Multimodal Models
by: Parikh, Dhruv, et al.
Published: (2026) -
ClusterViG: Efficient Globally Aware Vision GNNs via Image Partitioning
by: Parikh, Dhruv, et al.
Published: (2025) -
Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
by: Ramachandran, Anvitha, et al.
Published: (2025)