ScatterFormer: Efficient Voxel Transformer with Scattered Linear Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Chenhang, Li, Ruihuang, Zhang, Guowen, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object Detection
von: Zhang, Guowen, et al.
Veröffentlicht: (2024)
von: Zhang, Guowen, et al.
Veröffentlicht: (2024)
BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection
von: Zhang, Guowen, et al.
Veröffentlicht: (2025)
von: Zhang, Guowen, et al.
Veröffentlicht: (2025)
Fast Multi-view Consistent 3D Editing with Video Priors
von: Chen, Liyi, et al.
Veröffentlicht: (2025)
von: Chen, Liyi, et al.
Veröffentlicht: (2025)
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
FreCaS: Efficient Higher-Resolution Image Generation via Frequency-aware Cascaded Sampling
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2024)
SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection
von: Son, Hyeongseok, et al.
Veröffentlicht: (2025)
von: Son, Hyeongseok, et al.
Veröffentlicht: (2025)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
von: Meng, Weikang, et al.
Veröffentlicht: (2025)
von: Meng, Weikang, et al.
Veröffentlicht: (2025)
Source Prompt Disentangled Inversion for Boosting Image Editability with Diffusion Models
von: Li, Ruibin, et al.
Veröffentlicht: (2024)
von: Li, Ruibin, et al.
Veröffentlicht: (2024)
Efficient Point Clouds Upsampling via Flow Matching
von: Liu, Zhi-Song, et al.
Veröffentlicht: (2025)
von: Liu, Zhi-Song, et al.
Veröffentlicht: (2025)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
von: Shan, Jiquan, et al.
Veröffentlicht: (2025)
von: Shan, Jiquan, et al.
Veröffentlicht: (2025)
Motion-Guided Latent Diffusion for Temporally Consistent Real-world Video Super-resolution
von: Yang, Xi, et al.
Veröffentlicht: (2023)
von: Yang, Xi, et al.
Veröffentlicht: (2023)
HeightFormer: Learning Height Prediction in Voxel Features for Roadside Vision Centric 3D Object Detection via Transformer
von: Zhang, Zhang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhang, et al.
Veröffentlicht: (2025)
VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI
von: Le, Chenqian, et al.
Veröffentlicht: (2025)
von: Le, Chenqian, et al.
Veröffentlicht: (2025)
TMP: Temporal Motion Propagation for Online Video Super-Resolution
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2023)
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2023)
LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
von: Feng, Zhe, et al.
Veröffentlicht: (2026)
von: Feng, Zhe, et al.
Veröffentlicht: (2026)
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
von: Tang, Xiaoya, et al.
Veröffentlicht: (2025)
von: Tang, Xiaoya, et al.
Veröffentlicht: (2025)
Point-DAE: Denoising Autoencoders for Self-supervised Point Cloud Learning
von: Zhang, Yabin, et al.
Veröffentlicht: (2022)
von: Zhang, Yabin, et al.
Veröffentlicht: (2022)
Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
von: Huang, You, et al.
Veröffentlicht: (2025)
von: Huang, You, et al.
Veröffentlicht: (2025)
SyncNoise: Geometrically Consistent Noise Prediction for Text-based 3D Scene Editing
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
Voxel-Mesh Hybrid Representation for Real-Time View Synthesis
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
HyLiFormer: Hyperbolic Linear Attention for Skeleton-based Human Action Recognition
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
von: Long, Nguyen Huu Bao, et al.
Veröffentlicht: (2024)
von: Long, Nguyen Huu Bao, et al.
Veröffentlicht: (2024)
PointVoxelFormer -- Reviving point cloud networks for 3D medical imaging
von: Heinrich, Mattias Paul
Veröffentlicht: (2024)
von: Heinrich, Mattias Paul
Veröffentlicht: (2024)
LAPT: Label-driven Automated Prompt Tuning for OOD Detection with Vision-Language Models
von: Zhang, Yabin, et al.
Veröffentlicht: (2024)
von: Zhang, Yabin, et al.
Veröffentlicht: (2024)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
von: Xiao, Chaodong, et al.
Veröffentlicht: (2026)
von: Xiao, Chaodong, et al.
Veröffentlicht: (2026)
MB-TaylorFormer V2: Improved Multi-branch Linear Transformer Expanded by Taylor Formula for Image Restoration
von: Jin, Zhi, et al.
Veröffentlicht: (2025)
von: Jin, Zhi, et al.
Veröffentlicht: (2025)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
von: Fang, Tongcheng, et al.
Veröffentlicht: (2026)
von: Fang, Tongcheng, et al.
Veröffentlicht: (2026)
LiteVoxel: Low-memory Intelligent Thresholding for Efficient Voxel Rasterization
von: Lee, Jee Won, et al.
Veröffentlicht: (2025)
von: Lee, Jee Won, et al.
Veröffentlicht: (2025)
Geometrical Cross-Attention and Nonvoid Voxelization for Efficient 3D Medical Image Segmentation
von: Yuan, Chenxin, et al.
Veröffentlicht: (2026)
von: Yuan, Chenxin, et al.
Veröffentlicht: (2026)
WidthFormer: Toward Efficient Transformer-based BEV View Transformation
von: Yang, Chenhongyi, et al.
Veröffentlicht: (2024)
von: Yang, Chenhongyi, et al.
Veröffentlicht: (2024)
MedFormer: Hierarchical Medical Vision Transformer with Content-Aware Dual Sparse Selection Attention
von: Xia, Zunhui, et al.
Veröffentlicht: (2025)
von: Xia, Zunhui, et al.
Veröffentlicht: (2025)
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
von: Ding, Hangliang, et al.
Veröffentlicht: (2025)
von: Ding, Hangliang, et al.
Veröffentlicht: (2025)
RouteWinFormer: A Route-Window Transformer for Middle-range Attention in Image Restoration
von: Li, Qifan, et al.
Veröffentlicht: (2025)
von: Li, Qifan, et al.
Veröffentlicht: (2025)
PVTransformer: Point-to-Voxel Transformer for Scalable 3D Object Detection
von: Leng, Zhaoqi, et al.
Veröffentlicht: (2024)
von: Leng, Zhaoqi, et al.
Veröffentlicht: (2024)
Context and Geometry Aware Voxel Transformer for Semantic Scene Completion
von: Yu, Zhu, et al.
Veröffentlicht: (2024)
von: Yu, Zhu, et al.
Veröffentlicht: (2024)
AesFormer: Transform Everyday Photos into Beautiful Memories
von: Du, Tianxiang, et al.
Veröffentlicht: (2026)
von: Du, Tianxiang, et al.
Veröffentlicht: (2026)
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
von: Ghafoorian, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghafoorian, Mohsen, et al.
Veröffentlicht: (2025)
LowFormer: Hardware Efficient Design for Convolutional Transformer Backbones
von: Nottebaum, Moritz, et al.
Veröffentlicht: (2024)
von: Nottebaum, Moritz, et al.
Veröffentlicht: (2024)
General Geometry-aware Weakly Supervised 3D Object Detection
von: Zhang, Guowen, et al.
Veröffentlicht: (2024)
von: Zhang, Guowen, et al.
Veröffentlicht: (2024)
MixFormerV2: Efficient Fully Transformer Tracking
von: Cui, Yutao, et al.
Veröffentlicht: (2023)
von: Cui, Yutao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object Detection
von: Zhang, Guowen, et al.
Veröffentlicht: (2024) -
BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection
von: Zhang, Guowen, et al.
Veröffentlicht: (2025) -
Fast Multi-view Consistent 3D Editing with Video Priors
von: Chen, Liyi, et al.
Veröffentlicht: (2025) -
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
von: Li, Ruihuang, et al.
Veröffentlicht: (2024) -
FreCaS: Efficient Higher-Resolution Image Generation via Frequency-aware Cascaded Sampling
von: Zhang, Zhengqiang, et al.
Veröffentlicht: (2024)