ClustViT: Clustering-based Token Merging for Semantic Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Montello, Fabio, Güldenring, Ronja, Nalpantidis, Lazaros |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
von: Montello, Fabio, et al.
Veröffentlicht: (2025)
von: Montello, Fabio, et al.
Veröffentlicht: (2025)
PlaneSAM: Multimodal Plane Instance Segmentation Using the Segment Anything Model
von: Deng, Zhongchen, et al.
Veröffentlicht: (2024)
von: Deng, Zhongchen, et al.
Veröffentlicht: (2024)
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
von: Zhang, Yecheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yecheng, et al.
Veröffentlicht: (2026)
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025)
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025)
Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
von: Zhang, Lintong, et al.
Veröffentlicht: (2025)
von: Zhang, Lintong, et al.
Veröffentlicht: (2025)
Deep Learning Approaches for Human Action Recognition in Video Data
von: Xie, Yufei
Veröffentlicht: (2024)
von: Xie, Yufei
Veröffentlicht: (2024)
LatentForensics: Towards frugal deepfake detection in the StyleGAN latent space
von: Delmas, Matthieu, et al.
Veröffentlicht: (2023)
von: Delmas, Matthieu, et al.
Veröffentlicht: (2023)
Synthetic Industrial Object Detection: GenAI vs. Feature-Based Methods
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2025)
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2025)
One-to-Normal: Anomaly Personalization for Few-shot Anomaly Detection
von: Li, Yiyue, et al.
Veröffentlicht: (2025)
von: Li, Yiyue, et al.
Veröffentlicht: (2025)
Zero-Shot Multi-Criteria Visual Quality Inspection for Semi-Controlled Industrial Environments via Real-Time 3D Digital Twin Simulation
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2025)
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2025)
Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
von: Tian, Jie, et al.
Veröffentlicht: (2025)
von: Tian, Jie, et al.
Veröffentlicht: (2025)
SynthRender and IRIS: Open-Source Framework and Dataset for Bidirectional Sim-Real Transfer in Industrial Object Perception
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2026)
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2026)
DiffYOLO: Object Detection for Anti-Noise via YOLO and Diffusion Models
von: Liu, Yichen, et al.
Veröffentlicht: (2024)
von: Liu, Yichen, et al.
Veröffentlicht: (2024)
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
von: Su, Qile, et al.
Veröffentlicht: (2026)
von: Su, Qile, et al.
Veröffentlicht: (2026)
Instance Segmentation for Point Sets
von: Talwar, Abhimanyu, et al.
Veröffentlicht: (2025)
von: Talwar, Abhimanyu, et al.
Veröffentlicht: (2025)
Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
von: Chen, Jingkun, et al.
Veröffentlicht: (2025)
von: Chen, Jingkun, et al.
Veröffentlicht: (2025)
Image Reconstruction as a Tool for Feature Analysis
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025)
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
Sequence Matters: Harnessing Video Models in 3D Super-Resolution
von: Ko, Hyun-kyu, et al.
Veröffentlicht: (2024)
von: Ko, Hyun-kyu, et al.
Veröffentlicht: (2024)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
Learning Association via Track-Detection Matching for Multi-Object Tracking
von: Adžemović, Momir
Veröffentlicht: (2025)
von: Adžemović, Momir
Veröffentlicht: (2025)
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
von: Jin, Hang, et al.
Veröffentlicht: (2025)
von: Jin, Hang, et al.
Veröffentlicht: (2025)
VDPP: Video Depth Post-Processing for Speed and Scalability
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
Parking Space Detection in the City of Granada
von: Luis, Crespo-Orti, et al.
Veröffentlicht: (2025)
von: Luis, Crespo-Orti, et al.
Veröffentlicht: (2025)
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
von: Mohammad, Noor Islam S., et al.
Veröffentlicht: (2025)
von: Mohammad, Noor Islam S., et al.
Veröffentlicht: (2025)
Hierarchical Spatial Algorithms for High-Resolution Image Quantization and Feature Extraction
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
Creating Realistic Anterior Segment Optical Coherence Tomography Images using Generative Adversarial Networks
von: Assaf, Jad F., et al.
Veröffentlicht: (2023)
von: Assaf, Jad F., et al.
Veröffentlicht: (2023)
Segmenting the Complex and Irregular in Two-Phase Flows: A Real-World Empirical Study with SAM2
von: Küçük, Semanur, et al.
Veröffentlicht: (2025)
von: Küçük, Semanur, et al.
Veröffentlicht: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
LVP-CLIP:Revisiting CLIP for Continual Learning with Label Vector Pool
von: Ma, Yue, et al.
Veröffentlicht: (2024)
von: Ma, Yue, et al.
Veröffentlicht: (2024)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
Addressing Issues with Working Memory in Video Object Segmentation
von: Bromley, Clayton, et al.
Veröffentlicht: (2024)
von: Bromley, Clayton, et al.
Veröffentlicht: (2024)
Revisiting Energy-Based Model for Out-of-Distribution Detection
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
Differentiable Hierarchical Visual Tokenization
von: Aasan, Marius, et al.
Veröffentlicht: (2025)
von: Aasan, Marius, et al.
Veröffentlicht: (2025)
3D Reconstruction from Sketches
von: Talwar, Abhimanyu, et al.
Veröffentlicht: (2025)
von: Talwar, Abhimanyu, et al.
Veröffentlicht: (2025)
AI-Enhanced Precision in Sport Taekwondo: Increasing Fairness, Speed, and Trust in Competition (FST.ai)
von: Shariatmadar, Keivan, et al.
Veröffentlicht: (2025)
von: Shariatmadar, Keivan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
von: Montello, Fabio, et al.
Veröffentlicht: (2025) -
PlaneSAM: Multimodal Plane Instance Segmentation Using the Segment Anything Model
von: Deng, Zhongchen, et al.
Veröffentlicht: (2024) -
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
von: Zhang, Yecheng, et al.
Veröffentlicht: (2026) -
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025) -
Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
von: Zhang, Lintong, et al.
Veröffentlicht: (2025)