Agglomerative Token Clustering
Fuente:
arXiv
Saved in:
| Main Authors: | Haurum, Joakim Bruslund, Escalera, Sergio, Taylor, Graham W., Moeslund, Thomas B. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Sample High Quality 3D Fractals for Action Recognition Pre-Training?
by: Putak, Marko, et al.
Published: (2026)
by: Putak, Marko, et al.
Published: (2026)
An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
by: Lowe, Scott C., et al.
Published: (2024)
by: Lowe, Scott C., et al.
Published: (2024)
Generating Synthetic Stereo Datasets using 3D Gaussian Splatting and Expert Knowledge Transfer
by: Slezak, Filip, et al.
Published: (2025)
by: Slezak, Filip, et al.
Published: (2025)
CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at Scale
by: Gong, ZeMing, et al.
Published: (2024)
by: Gong, ZeMing, et al.
Published: (2024)
A Hyperbolic Perspective on Hierarchical Structure in Object-Centric Scene Representations
by: Madan, Neelu, et al.
Published: (2026)
by: Madan, Neelu, et al.
Published: (2026)
T-DEED: Temporal-Discriminability Enhancer Encoder-Decoder for Precise Event Spotting in Sports Videos
by: Xarles, Artur, et al.
Published: (2024)
by: Xarles, Artur, et al.
Published: (2024)
ASTRA: An Action Spotting TRAnsformer for Soccer Videos
by: Xarles, Artur, et al.
Published: (2024)
by: Xarles, Artur, et al.
Published: (2024)
AdaSpot: Spend Resolution Where It Matters for Precise Event Spotting
by: Xarles, Artur, et al.
Published: (2026)
by: Xarles, Artur, et al.
Published: (2026)
Action Valuation in Sports: A Survey
by: Xarles, Artur, et al.
Published: (2025)
by: Xarles, Artur, et al.
Published: (2025)
A noisy elephant in the room: Is your out-of-distribution detector robust to label noise?
by: Humblot-Renaux, Galadrielle, et al.
Published: (2024)
by: Humblot-Renaux, Galadrielle, et al.
Published: (2024)
Hyperbolic Multimodal Representation Learning for Biological Taxonomies
by: Gong, ZeMing, et al.
Published: (2025)
by: Gong, ZeMing, et al.
Published: (2025)
Comparing Euclidean and Hyperbolic K-Means for Generalized Category Discovery
by: Dalal, Mohamad, et al.
Published: (2026)
by: Dalal, Mohamad, et al.
Published: (2026)
COOkeD: Ensemble-based OOD detection in the era of zero-shot CLIP
by: Humblot-Renaux, Galadrielle, et al.
Published: (2025)
by: Humblot-Renaux, Galadrielle, et al.
Published: (2025)
Point Cloud Segmentation of Agricultural Vehicles using 3D Gaussian Splatting
by: Christiansen, Alfred T., et al.
Published: (2025)
by: Christiansen, Alfred T., et al.
Published: (2025)
Machine Unlearning in Hyperbolic vs. Euclidean Multimodal Contrastive Learning: Adapting Alignment Calibration to MERU
by: Vidal, Àlex Pujol, et al.
Published: (2025)
by: Vidal, Àlex Pujol, et al.
Published: (2025)
From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition
by: Aslam, Nazia, et al.
Published: (2026)
by: Aslam, Nazia, et al.
Published: (2026)
Multimodal classification of forest biodiversity potential from 2D orthophotos and 3D airborne laser scanning point clouds
by: Jensen, Simon B., et al.
Published: (2025)
by: Jensen, Simon B., et al.
Published: (2025)
BIOSCAN-5M: A Multimodal Dataset for Insect Biodiversity
by: Gharaee, Zahra, et al.
Published: (2024)
by: Gharaee, Zahra, et al.
Published: (2024)
Bounding Boxes and Probabilistic Graphical Models: Video Anomaly Detection Simplified
by: Siemon, Mia, et al.
Published: (2024)
by: Siemon, Mia, et al.
Published: (2024)
Video Anomaly Detection with Contours -- A Study
by: Siemon, Mia, et al.
Published: (2025)
by: Siemon, Mia, et al.
Published: (2025)
A Transformer Model for Boundary Detection in Continuous Sign Language
by: Rastgoo, Razieh, et al.
Published: (2024)
by: Rastgoo, Razieh, et al.
Published: (2024)
Seamless Human Motion Composition with Blended Positional Encodings
by: Barquero, German, et al.
Published: (2024)
by: Barquero, German, et al.
Published: (2024)
L-SWAG: Layer-Sample Wise Activation with Gradients information for Zero-Shot NAS on Vision Transformers
by: Casarin, Sofia, et al.
Published: (2025)
by: Casarin, Sofia, et al.
Published: (2025)
Foundation Models for Video Understanding: A Survey
by: Madan, Neelu, et al.
Published: (2024)
by: Madan, Neelu, et al.
Published: (2024)
Sparse-Dense Side-Tuner for efficient Video Temporal Grounding
by: Pujol-Perich, David, et al.
Published: (2025)
by: Pujol-Perich, David, et al.
Published: (2025)
What Matters in Virtual Try-Off? Dual-UNet Diffusion Model For Garment Reconstruction
by: Truong, Loc-Phat, et al.
Published: (2026)
by: Truong, Loc-Phat, et al.
Published: (2026)
SADA: Semantic adversarial unsupervised domain adaptation for Temporal Action Localization
by: Pujol-Perich, David, et al.
Published: (2023)
by: Pujol-Perich, David, et al.
Published: (2023)
SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models
by: Rabasseda, Oriol, et al.
Published: (2026)
by: Rabasseda, Oriol, et al.
Published: (2026)
A Generative Multi-Resolution Pyramid and Normal-Conditioning 3D Cloth Draping
by: Laczkó, Hunor, et al.
Published: (2023)
by: Laczkó, Hunor, et al.
Published: (2023)
GRASP-GCN: Graph-Shape Prioritization for Neural Architecture Search under Distribution Shifts
by: Casarin, Sofia, et al.
Published: (2024)
by: Casarin, Sofia, et al.
Published: (2024)
YOLO11-JDE: Fast and Accurate Multi-Object Tracking with Self-Supervised Re-ID
by: Erregue, Iñaki, et al.
Published: (2025)
by: Erregue, Iñaki, et al.
Published: (2025)
PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
by: Erregue, Iñaki, et al.
Published: (2026)
by: Erregue, Iñaki, et al.
Published: (2026)
AutoFish: Dataset and Benchmark for Fine-grained Analysis of Fish
by: Bengtson, Stefan Hein, et al.
Published: (2025)
by: Bengtson, Stefan Hein, et al.
Published: (2025)
OpenTrench3D: A Photogrammetric 3D Point Cloud Dataset for Semantic Segmentation of Underground Utilities
by: Hansen, Lasse H., et al.
Published: (2024)
by: Hansen, Lasse H., et al.
Published: (2024)
Action Anticipation from SoccerNet Football Video Broadcasts
by: Dalal, Mohamad, et al.
Published: (2025)
by: Dalal, Mohamad, et al.
Published: (2025)
Rank-Aware Agglomeration of Foundation Models for Immunohistochemistry Image Cell Counting
by: Huang, Zuqi, et al.
Published: (2025)
by: Huang, Zuqi, et al.
Published: (2025)
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
by: Fan, Qihang, et al.
Published: (2024)
by: Fan, Qihang, et al.
Published: (2024)
Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining
by: Li, Yuheng, et al.
Published: (2026)
by: Li, Yuheng, et al.
Published: (2026)
AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One
by: Ranzinger, Mike, et al.
Published: (2023)
by: Ranzinger, Mike, et al.
Published: (2023)
Sea-ing Through Scattered Rays: Revisiting the Image Formation Model for Realistic Underwater Image Generation
by: Ismiroglou, Vasiliki, et al.
Published: (2025)
by: Ismiroglou, Vasiliki, et al.
Published: (2025)
Similar Items
-
How to Sample High Quality 3D Fractals for Action Recognition Pre-Training?
by: Putak, Marko, et al.
Published: (2026) -
An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
by: Lowe, Scott C., et al.
Published: (2024) -
Generating Synthetic Stereo Datasets using 3D Gaussian Splatting and Expert Knowledge Transfer
by: Slezak, Filip, et al.
Published: (2025) -
CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at Scale
by: Gong, ZeMing, et al.
Published: (2024) -
A Hyperbolic Perspective on Hierarchical Structure in Object-Centric Scene Representations
by: Madan, Neelu, et al.
Published: (2026)