ClassifyViStA:WCE Classification with Visual understanding through Segmentation and Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Balasubramanian, S., Abhishek, Ammu, Krishna, Yedu, Gera, Darshan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
S2IL: Structurally Stable Incremental Learning
by: Balasubramanian, S, et al.
Published: (2025)
by: Balasubramanian, S, et al.
Published: (2025)
EXACFS -- A CIL Method to mitigate Catastrophic Forgetting
by: Balasubramanian, S, et al.
Published: (2024)
by: Balasubramanian, S, et al.
Published: (2024)
Learn From Zoom: Decoupled Supervised Contrastive Learning For WCE Image Classification
by: Qiu, Kunpeng, et al.
Published: (2024)
by: Qiu, Kunpeng, et al.
Published: (2024)
Multi-scale fMRI time series analysis for understanding neurodegeneration in MCI
by: R., Ammu, et al.
Published: (2024)
by: R., Ammu, et al.
Published: (2024)
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
by: Bi, Hanbo, et al.
Published: (2025)
by: Bi, Hanbo, et al.
Published: (2025)
Predicting Visual Attention in Graphic Design Documents
by: Chakraborty, Souradeep, et al.
Published: (2024)
by: Chakraborty, Souradeep, et al.
Published: (2024)
ToaSt: Token Channel Selection and Structured Pruning for Efficient ViT
by: Moon, Hyunchan, et al.
Published: (2026)
by: Moon, Hyunchan, et al.
Published: (2026)
Your ViT is Secretly an Image Segmentation Model
by: Kerssies, Tommie, et al.
Published: (2025)
by: Kerssies, Tommie, et al.
Published: (2025)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
ViM-UNet: Vision Mamba for Biomedical Segmentation
by: Archit, Anwai, et al.
Published: (2024)
by: Archit, Anwai, et al.
Published: (2024)
Applying ViT in Generalized Few-shot Semantic Segmentation
by: Geng, Liyuan, et al.
Published: (2024)
by: Geng, Liyuan, et al.
Published: (2024)
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
by: Puy, Gilles, et al.
Published: (2026)
by: Puy, Gilles, et al.
Published: (2026)
Diabetic Retinopathy Lesion Segmentation through Attention Mechanisms
by: Jithesh, Aruna, et al.
Published: (2026)
by: Jithesh, Aruna, et al.
Published: (2026)
CanViT: Toward Active-Vision Foundation Models
by: Berreby, Yohaï-Eliel, et al.
Published: (2026)
by: Berreby, Yohaï-Eliel, et al.
Published: (2026)
SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention
by: Dhake, Shreyas C., et al.
Published: (2025)
by: Dhake, Shreyas C., et al.
Published: (2025)
Sequence-Preserving Dual-FoV Defense for Traffic Sign and Light Recognition in Autonomous Vehicles
by: Joshi, Abhishek, et al.
Published: (2025)
by: Joshi, Abhishek, et al.
Published: (2025)
Active Learning via Classifier Impact and Greedy Selection for Interactive Image Retrieval
by: Bar, Leah, et al.
Published: (2024)
by: Bar, Leah, et al.
Published: (2024)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
CellViT++: Energy-Efficient and Adaptive Cell Segmentation and Classification Using Foundation Models
by: Hörst, Fabian, et al.
Published: (2025)
by: Hörst, Fabian, et al.
Published: (2025)
CarGait: Cross-Attention based Re-ranking for Gait recognition
by: Habib, Gavriel, et al.
Published: (2025)
by: Habib, Gavriel, et al.
Published: (2025)
ViUniT: Visual Unit Tests for More Robust Visual Programming
by: Panagopoulou, Artemis, et al.
Published: (2024)
by: Panagopoulou, Artemis, et al.
Published: (2024)
QMaxViT-Unet+: A Query-Based MaxViT-Unet with Edge Enhancement for Scribble-Supervised Segmentation of Medical Images
by: Nguyen-Tat, Thien B., et al.
Published: (2025)
by: Nguyen-Tat, Thien B., et al.
Published: (2025)
StAR: Segment Anything Reasoner
by: Yun, Seokju, et al.
Published: (2026)
by: Yun, Seokju, et al.
Published: (2026)
ViLLa: Video Reasoning Segmentation with Large Language Model
by: Zheng, Rongkun, et al.
Published: (2024)
by: Zheng, Rongkun, et al.
Published: (2024)
RepViT-SAM: Towards Real-Time Segmenting Anything
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
ViSpeak: Visual Instruction Feedback in Streaming Videos
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
LoopViT: Scaling Visual ARC with Looped Transformers
by: Shu, Wen-Jie, et al.
Published: (2026)
by: Shu, Wen-Jie, et al.
Published: (2026)
Creating Ensembles of Classifiers through UMDA for Aerial Scene Classification
by: Faria, Fabio A., et al.
Published: (2023)
by: Faria, Fabio A., et al.
Published: (2023)
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
Alias-Free ViT: Fractional Shift Invariance via Linear Attention
by: Michaeli, Hagay, et al.
Published: (2025)
by: Michaeli, Hagay, et al.
Published: (2025)
MoViAD: A Modular Library for Visual Anomaly Detection
by: Barusco, Manuel, et al.
Published: (2025)
by: Barusco, Manuel, et al.
Published: (2025)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
by: Xu, Lijian, et al.
Published: (2024)
by: Xu, Lijian, et al.
Published: (2024)
A Technique for Classifying Static Gestures Using UWB Radar
by: Sebastian, Abhishek, et al.
Published: (2023)
by: Sebastian, Abhishek, et al.
Published: (2023)
MicroViT: A Vision Transformer with Low Complexity Self Attention for Edge Device
by: Setyawan, Novendra, et al.
Published: (2025)
by: Setyawan, Novendra, et al.
Published: (2025)
ViSTa Dataset: Do vision-language models understand sequential tasks?
by: Wybitul, Evžen, et al.
Published: (2024)
by: Wybitul, Evžen, et al.
Published: (2024)
MC-ViViT: Multi-branch Classifier-ViViT to detect Mild Cognitive Impairment in older adults using facial videos
by: Sun, Jian, et al.
Published: (2023)
by: Sun, Jian, et al.
Published: (2023)
SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
by: Kombol, Naomi, et al.
Published: (2026)
by: Kombol, Naomi, et al.
Published: (2026)
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images
by: Tosato, Lucrezia, et al.
Published: (2024)
by: Tosato, Lucrezia, et al.
Published: (2024)
BEVANet: Bilateral Efficient Visual Attention Network for Real-Time Semantic Segmentation
by: Huang, Ping-Mao, et al.
Published: (2025)
by: Huang, Ping-Mao, et al.
Published: (2025)
FilterViT and DropoutViT
by: Sun, Bohang
Published: (2024)
by: Sun, Bohang
Published: (2024)
Similar Items
-
S2IL: Structurally Stable Incremental Learning
by: Balasubramanian, S, et al.
Published: (2025) -
EXACFS -- A CIL Method to mitigate Catastrophic Forgetting
by: Balasubramanian, S, et al.
Published: (2024) -
Learn From Zoom: Decoupled Supervised Contrastive Learning For WCE Image Classification
by: Qiu, Kunpeng, et al.
Published: (2024) -
Multi-scale fMRI time series analysis for understanding neurodegeneration in MCI
by: R., Ammu, et al.
Published: (2024) -
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
by: Bi, Hanbo, et al.
Published: (2025)