ENACT: Entropy-based Clustering of Attention Input for Reducing the Computational Needs of Object Detection Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Savathrakis, Giorgos, Argyros, Antonis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OCCAM: Class-Agnostic, Training-Free, Prior-Free and Multi-Class Object Counting
von: Spanakis, Michail, et al.
Veröffentlicht: (2026)
von: Spanakis, Michail, et al.
Veröffentlicht: (2026)
Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
von: Karvounas, Giorgos, et al.
Veröffentlicht: (2025)
von: Karvounas, Giorgos, et al.
Veröffentlicht: (2025)
Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
von: Bacharidis, Konstantinos, et al.
Veröffentlicht: (2025)
von: Bacharidis, Konstantinos, et al.
Veröffentlicht: (2025)
D-PoSE: Depth as an Intermediate Representation for 3D Human Pose and Shape Estimation
von: Vasilikopoulos, Nikolaos, et al.
Veröffentlicht: (2024)
von: Vasilikopoulos, Nikolaos, et al.
Veröffentlicht: (2024)
Recognizing Unseen States of Unknown Objects by Leveraging Knowledge Graphs
von: Gouidis, Filipos, et al.
Veröffentlicht: (2023)
von: Gouidis, Filipos, et al.
Veröffentlicht: (2023)
Anticipating Object State Changes in Long Procedural Videos
von: Manousaki, Victoria, et al.
Veröffentlicht: (2024)
von: Manousaki, Victoria, et al.
Veröffentlicht: (2024)
Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images
von: Qammaz, Ammar, et al.
Veröffentlicht: (2024)
von: Qammaz, Ammar, et al.
Veröffentlicht: (2024)
Combining Facial Videos and Biosignals for Stress Estimation During Driving
von: Valergaki, Paraskevi, et al.
Veröffentlicht: (2026)
von: Valergaki, Paraskevi, et al.
Veröffentlicht: (2026)
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
von: Benavent-Lledo, Manuel, et al.
Veröffentlicht: (2026)
von: Benavent-Lledo, Manuel, et al.
Veröffentlicht: (2026)
Fusion Transformer with Object Mask Guidance for Image Forgery Analysis
von: Karageorgiou, Dimitrios, et al.
Veröffentlicht: (2024)
von: Karageorgiou, Dimitrios, et al.
Veröffentlicht: (2024)
Adversarial Attention Perturbations for Large Object Detection Transformers
von: Yahn, Zachary, et al.
Veröffentlicht: (2025)
von: Yahn, Zachary, et al.
Veröffentlicht: (2025)
Fusing Domain-Specific Content from Large Language Models into Knowledge Graphs for Enhanced Zero Shot Object State Classification
von: Gouidis, Filippos, et al.
Veröffentlicht: (2024)
von: Gouidis, Filippos, et al.
Veröffentlicht: (2024)
Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
von: Benavent-Lledo, Manuel, et al.
Veröffentlicht: (2025)
von: Benavent-Lledo, Manuel, et al.
Veröffentlicht: (2025)
Quantization Robustness to Input Degradations for Object Detection
von: Karimov, Toghrul, et al.
Veröffentlicht: (2025)
von: Karimov, Toghrul, et al.
Veröffentlicht: (2025)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
Hierarchical Graph Interaction Transformer with Dynamic Token Clustering for Camouflaged Object Detection
von: Yao, Siyuan, et al.
Veröffentlicht: (2024)
von: Yao, Siyuan, et al.
Veröffentlicht: (2024)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
von: Song, Jiale, et al.
Veröffentlicht: (2026)
von: Song, Jiale, et al.
Veröffentlicht: (2026)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
von: Benavent-Lledo, Manuel, et al.
Veröffentlicht: (2024)
von: Benavent-Lledo, Manuel, et al.
Veröffentlicht: (2024)
Scene Adaptive Sparse Transformer for Event-based Object Detection
von: Peng, Yansong, et al.
Veröffentlicht: (2024)
von: Peng, Yansong, et al.
Veröffentlicht: (2024)
Tri-Modal Fusion Transformers for UAV-based Object Detection
von: Iaboni, Craig, et al.
Veröffentlicht: (2026)
von: Iaboni, Craig, et al.
Veröffentlicht: (2026)
DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination
von: Gong, Xuan, et al.
Veröffentlicht: (2024)
von: Gong, Xuan, et al.
Veröffentlicht: (2024)
Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
von: Hao, Haiqing, et al.
Veröffentlicht: (2026)
von: Hao, Haiqing, et al.
Veröffentlicht: (2026)
You Only Need Less Attention at Each Stage in Vision Transformers
von: Zhang, Shuoxi, et al.
Veröffentlicht: (2024)
von: Zhang, Shuoxi, et al.
Veröffentlicht: (2024)
EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy
von: Ge, Xuanyu, et al.
Veröffentlicht: (2026)
von: Ge, Xuanyu, et al.
Veröffentlicht: (2026)
Distilling Vision Transformers for Distortion-Robust Representation Learning
von: Alexis, Konstantinos, et al.
Veröffentlicht: (2026)
von: Alexis, Konstantinos, et al.
Veröffentlicht: (2026)
Index-Aligned Query Distillation for Transformer-based Incremental Object Detection
von: Ma, Mingxiao, et al.
Veröffentlicht: (2025)
von: Ma, Mingxiao, et al.
Veröffentlicht: (2025)
StereoDETR: Stereo-based Transformer for 3D Object Detection
von: Mu, Shiyi, et al.
Veröffentlicht: (2025)
von: Mu, Shiyi, et al.
Veröffentlicht: (2025)
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
von: Luo, Jiayi, et al.
Veröffentlicht: (2026)
OAT: Object-Level Attention Transformer for Gaze Scanpath Prediction
von: Fang, Yini, et al.
Veröffentlicht: (2024)
von: Fang, Yini, et al.
Veröffentlicht: (2024)
Dynamic Object Queries for Transformer-based Incremental Object Detection
von: Zhang, Jichuan, et al.
Veröffentlicht: (2024)
von: Zhang, Jichuan, et al.
Veröffentlicht: (2024)
A Simple yet Effective Network based on Vision Transformer for Camouflaged Object and Salient Object Detection
von: Hao, Chao, et al.
Veröffentlicht: (2024)
von: Hao, Chao, et al.
Veröffentlicht: (2024)
SeaDATE: Remedy Dual-Attention Transformer with Semantic Alignment via Contrast Learning for Multimodal Object Detection
von: Dong, Shuhan, et al.
Veröffentlicht: (2024)
von: Dong, Shuhan, et al.
Veröffentlicht: (2024)
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
von: Lee, Joohyeon, et al.
Veröffentlicht: (2025)
von: Lee, Joohyeon, et al.
Veröffentlicht: (2025)
Transformer based Multitask Learning for Image Captioning and Object Detection
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection
von: Wang, Hui, et al.
Veröffentlicht: (2026)
von: Wang, Hui, et al.
Veröffentlicht: (2026)
DPFT: Dual Perspective Fusion Transformer for Camera-Radar-based Object Detection
von: Fent, Felix, et al.
Veröffentlicht: (2024)
von: Fent, Felix, et al.
Veröffentlicht: (2024)
LAM-YOLO: Drones-based Small Object Detection on Lighting-Occlusion Attention Mechanism YOLO
von: Zheng, Yuchen, et al.
Veröffentlicht: (2024)
von: Zheng, Yuchen, et al.
Veröffentlicht: (2024)
MonoCLUE : Object-Aware Clustering Enhances Monocular 3D Object Detection
von: Yang, Sunghun, et al.
Veröffentlicht: (2025)
von: Yang, Sunghun, et al.
Veröffentlicht: (2025)
Small Object Detection for Birds with Swin Transformer
von: Huo, Da, et al.
Veröffentlicht: (2025)
von: Huo, Da, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OCCAM: Class-Agnostic, Training-Free, Prior-Free and Multi-Class Object Counting
von: Spanakis, Michail, et al.
Veröffentlicht: (2026) -
Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
von: Karvounas, Giorgos, et al.
Veröffentlicht: (2025) -
Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
von: Bacharidis, Konstantinos, et al.
Veröffentlicht: (2025) -
D-PoSE: Depth as an Intermediate Representation for 3D Human Pose and Shape Estimation
von: Vasilikopoulos, Nikolaos, et al.
Veröffentlicht: (2024) -
Recognizing Unseen States of Unknown Objects by Leveraging Knowledge Graphs
von: Gouidis, Filipos, et al.
Veröffentlicht: (2023)