Point Transformer V3: Simpler, Faster, Stronger
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Xiaoyang, Jiang, Li, Wang, Peng-Shuai, Liu, Zhijian, Liu, Xihui, Qiao, Yu, Ouyang, Wanli, He, Tong, Zhao, Hengshuang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Point Transformer V3 Extreme: 1st Place Solution for 2024 Waymo Open Dataset Challenge in Semantic Segmentation
by: Wu, Xiaoyang, et al.
Published: (2024)
by: Wu, Xiaoyang, et al.
Published: (2024)
Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training
by: Wu, Xiaoyang, et al.
Published: (2023)
by: Wu, Xiaoyang, et al.
Published: (2023)
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
by: Zhu, Haoyi, et al.
Published: (2023)
by: Zhu, Haoyi, et al.
Published: (2023)
PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection
by: Ren, Botao, et al.
Published: (2024)
by: Ren, Botao, et al.
Published: (2024)
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
by: Yang, Yunhan, et al.
Published: (2025)
by: Yang, Yunhan, et al.
Published: (2025)
Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images
by: Qi, Zhangyang, et al.
Published: (2024)
by: Qi, Zhangyang, et al.
Published: (2024)
DreamComposer: Controllable 3D Object Generation via Multi-View Conditions
by: Yang, Yunhan, et al.
Published: (2023)
by: Yang, Yunhan, et al.
Published: (2023)
FiT: Flexible Vision Transformer for Diffusion Model
by: Lu, Zeyu, et al.
Published: (2024)
by: Lu, Zeyu, et al.
Published: (2024)
GPT4Point: A Unified Framework for Point-Language Understanding and Generation
by: Qi, Zhangyang, et al.
Published: (2023)
by: Qi, Zhangyang, et al.
Published: (2023)
UniPAD: A Universal Pre-training Paradigm for Autonomous Driving
by: Yang, Honghui, et al.
Published: (2023)
by: Yang, Honghui, et al.
Published: (2023)
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
by: Chu, Xiangxiang, et al.
Published: (2024)
by: Chu, Xiangxiang, et al.
Published: (2024)
OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding
by: Wang, Chengyao, et al.
Published: (2024)
by: Wang, Chengyao, et al.
Published: (2024)
PredBench: Benchmarking Spatio-Temporal Prediction across Diverse Disciplines
by: Wang, ZiDong, et al.
Published: (2024)
by: Wang, ZiDong, et al.
Published: (2024)
LitePT: Lighter Yet Stronger Point Transformer
by: Yue, Yuanwen, et al.
Published: (2025)
by: Yue, Yuanwen, et al.
Published: (2025)
OpenIns3D: Snap and Lookup for 3D Open-vocabulary Instance Segmentation
by: Huang, Zhening, et al.
Published: (2023)
by: Huang, Zhening, et al.
Published: (2023)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
by: Fan, Xianzhe, et al.
Published: (2026)
by: Fan, Xianzhe, et al.
Published: (2026)
Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
by: Liu, Jiazhen, et al.
Published: (2025)
by: Liu, Jiazhen, et al.
Published: (2025)
Unleashing Degradation-Carrying Features in Symmetric U-Net: Simpler and Stronger Baselines for All-in-One Image Restoration
by: Jiao, Wenlong, et al.
Published: (2025)
by: Jiao, Wenlong, et al.
Published: (2025)
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
VST++: Efficient and Stronger Visual Saliency Transformer
by: Liu, Nian, et al.
Published: (2023)
by: Liu, Nian, et al.
Published: (2023)
Utonia: Toward One Encoder for All Point Clouds
by: Zhang, Yujia, et al.
Published: (2026)
by: Zhang, Yujia, et al.
Published: (2026)
Pixel-GS: Density Control with Pixel-aware Gradient for 3D Gaussian Splatting
by: Zhang, Zheng, et al.
Published: (2024)
by: Zhang, Zheng, et al.
Published: (2024)
UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Cut2Next: Generating Next Shot via In-Context Tuning
by: He, Jingwen, et al.
Published: (2025)
by: He, Jingwen, et al.
Published: (2025)
One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
GigaGS: Scaling up Planar-Based 3D Gaussians for Large Scene Surface Reconstruction
by: Chen, Junyi, et al.
Published: (2024)
by: Chen, Junyi, et al.
Published: (2024)
LION: Linear Group RNN for 3D Object Detection in Point Clouds
by: Liu, Zhe, et al.
Published: (2024)
by: Liu, Zhe, et al.
Published: (2024)
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
by: Zhang, Yujia, et al.
Published: (2025)
by: Zhang, Yujia, et al.
Published: (2025)
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025)
by: Yue, Xiaoyu, et al.
Published: (2025)
Tracking Meets LoRA: Faster Training, Larger Model, Stronger Performance
by: Lin, Liting, et al.
Published: (2024)
by: Lin, Liting, et al.
Published: (2024)
CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
by: Karaev, Nikita, et al.
Published: (2024)
by: Karaev, Nikita, et al.
Published: (2024)
Sonata: Self-Supervised Learning of Reliable Point Representations
by: Wu, Xiaoyang, et al.
Published: (2025)
by: Wu, Xiaoyang, et al.
Published: (2025)
LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D Scans
by: Huang, Zhening, et al.
Published: (2025)
by: Huang, Zhening, et al.
Published: (2025)
Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
by: Chen, Junyi, et al.
Published: (2024)
by: Chen, Junyi, et al.
Published: (2024)
GVGEN: Text-to-3D Generation with Volumetric Representation
by: He, Xianglong, et al.
Published: (2024)
by: He, Xianglong, et al.
Published: (2024)
Stronger Normalization-Free Transformers
by: Chen, Mingzhi, et al.
Published: (2025)
by: Chen, Mingzhi, et al.
Published: (2025)
Similar Items
-
Point Transformer V3 Extreme: 1st Place Solution for 2024 Waymo Open Dataset Challenge in Semantic Segmentation
by: Wu, Xiaoyang, et al.
Published: (2024) -
Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training
by: Wu, Xiaoyang, et al.
Published: (2023) -
PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
by: Zhu, Haoyi, et al.
Published: (2023) -
PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection
by: Ren, Botao, et al.
Published: (2024) -
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024)