Make Your ViT-based Multi-view 3D Detectors Faster via Token Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Dingyuan, Liang, Dingkang, Tan, Zichang, Ye, Xiaoqing, Zhang, Cheng, Wang, Jingdong, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
von: Ji, Mingqian, et al.
Veröffentlicht: (2026)
von: Ji, Mingqian, et al.
Veröffentlicht: (2026)
SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2023)
Token Cropr: Faster ViTs for Quite a Few Tasks
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2026)
von: Zhou, Xin, et al.
Veröffentlicht: (2026)
OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection
von: Hou, Jinghua, et al.
Veröffentlicht: (2024)
von: Hou, Jinghua, et al.
Veröffentlicht: (2024)
ViT Registers and Fractal ViT
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2026)
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2026)
You Only Look Bottom-Up for Monocular 3D Object Detection
von: Xiong, Kaixin, et al.
Veröffentlicht: (2024)
von: Xiong, Kaixin, et al.
Veröffentlicht: (2024)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
Your ViT is Secretly an Image Segmentation Model
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
von: Xu, Xuwei, et al.
Veröffentlicht: (2023)
von: Xu, Xuwei, et al.
Veröffentlicht: (2023)
DiffPoint: Single and Multi-view Point Cloud Reconstruction with ViT Based Diffusion Model
von: Feng, Yu, et al.
Veröffentlicht: (2024)
von: Feng, Yu, et al.
Veröffentlicht: (2024)
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
von: Han, Donghoon, et al.
Veröffentlicht: (2023)
von: Han, Donghoon, et al.
Veröffentlicht: (2023)
UniFuture: A 4D Driving World Model for Future Generation and Perception
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-based Roadside 3D Object Detection
von: Wang, Wenjie, et al.
Veröffentlicht: (2024)
von: Wang, Wenjie, et al.
Veröffentlicht: (2024)
MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition
von: Kim, Ye-eun, et al.
Veröffentlicht: (2025)
von: Kim, Ye-eun, et al.
Veröffentlicht: (2025)
PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference
von: Li, Ye, et al.
Veröffentlicht: (2024)
von: Li, Ye, et al.
Veröffentlicht: (2024)
Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
von: Bar-Shalom, Guy, et al.
Veröffentlicht: (2025)
von: Bar-Shalom, Guy, et al.
Veröffentlicht: (2025)
TFS-ViT: Token-Level Feature Stylization for Domain Generalization
von: Noori, Mehrdad, et al.
Veröffentlicht: (2023)
von: Noori, Mehrdad, et al.
Veröffentlicht: (2023)
TPC-ViT: Token Propagation Controller for Efficient Vision Transformer
von: Zhu, Wentao
Veröffentlicht: (2024)
von: Zhu, Wentao
Veröffentlicht: (2024)
ViTCAE: ViT-based Class-conditioned Autoencoder
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
von: Alvetreti, Federico, et al.
Veröffentlicht: (2025)
von: Alvetreti, Federico, et al.
Veröffentlicht: (2025)
SOOD++: Leveraging Unlabeled Data to Boost Oriented Object Detection
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
Refining Datapath for Microscaling ViTs
von: Xiao, Can, et al.
Veröffentlicht: (2025)
von: Xiao, Can, et al.
Veröffentlicht: (2025)
AE-ViT: Token Enhancement for Vision Transformers via CNN-Based Autoencoder Ensembles
von: AIRCC
Veröffentlicht: (2025)
von: AIRCC
Veröffentlicht: (2025)
SEED: A Simple and Effective 3D DETR in Point Clouds
von: Liu, Zhe, et al.
Veröffentlicht: (2024)
von: Liu, Zhe, et al.
Veröffentlicht: (2024)
ToaSt: Token Channel Selection and Structured Pruning for Efficient ViT
von: Moon, Hyunchan, et al.
Veröffentlicht: (2026)
von: Moon, Hyunchan, et al.
Veröffentlicht: (2026)
CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
von: Gong, Wenyi, et al.
Veröffentlicht: (2025)
von: Gong, Wenyi, et al.
Veröffentlicht: (2025)
SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
von: Papais, Sandro, et al.
Veröffentlicht: (2026)
von: Papais, Sandro, et al.
Veröffentlicht: (2026)
Pneumonia Image Classification Based on Lightweight Mobile ViT Networks
von: Zhiqiang Zheng, et al.
Veröffentlicht: (2025)
von: Zhiqiang Zheng, et al.
Veröffentlicht: (2025)
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
von: Zhao, Zongchuang, et al.
Veröffentlicht: (2025)
von: Zhao, Zongchuang, et al.
Veröffentlicht: (2025)
PointMamba: A Simple State Space Model for Point Cloud Analysis
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
von: You, Haoran, et al.
Veröffentlicht: (2022)
von: You, Haoran, et al.
Veröffentlicht: (2022)
FFM ‐ ViT : an efficient fish species classification method based on deep features and transformers
von: Yuwei Gao, et al.
Veröffentlicht: (2025)
von: Yuwei Gao, et al.
Veröffentlicht: (2025)
Improve Contrastive Clustering Performance by Multiple Fusing-Augmenting ViT Blocks
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
von: Liu, Jiani, et al.
Veröffentlicht: (2025)
von: Liu, Jiani, et al.
Veröffentlicht: (2025)
Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense Predictions
von: Zhang, Jingdong, et al.
Veröffentlicht: (2024)
von: Zhang, Jingdong, et al.
Veröffentlicht: (2024)
Deeper Inside Deep ViT
von: Hong, Sungrae
Veröffentlicht: (2025)
von: Hong, Sungrae
Veröffentlicht: (2025)
Exploring Plain ViT Reconstruction for Multi-class Unsupervised Anomaly Detection
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
von: Ji, Mingqian, et al.
Veröffentlicht: (2026) -
SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2023) -
Token Cropr: Faster ViTs for Quite a Few Tasks
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024) -
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2026) -
OPEN: Object-wise Position Embedding for Multi-view 3D Object Detection
von: Hou, Jinghua, et al.
Veröffentlicht: (2024)