Building a Strong Pre-Training Baseline for Universal 3D Large-Scale Perception
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Haoming, Zhang, Zhizhong, Qu, Yanyun, Zhang, Ruixin, Tan, Xin, Xie, Yuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring the Untouched Sweeps for Conflict-Aware 3D Segmentation Pretraining
por: Sun, Tianfang, et al.
Publicado: (2024)
por: Sun, Tianfang, et al.
Publicado: (2024)
Beyond the Label Itself: Latent Labels Enhance Semi-supervised Point Cloud Panoptic Segmentation
por: Chen, Yujun, et al.
Publicado: (2023)
por: Chen, Yujun, et al.
Publicado: (2023)
COTR: Compact Occupancy TRansformer for Vision-based 3D Occupancy Prediction
por: Ma, Qihang, et al.
Publicado: (2023)
por: Ma, Qihang, et al.
Publicado: (2023)
Mutual Information Guided Optimal Transport for Unsupervised Visible-Infrared Person Re-identification
por: Zhang, Zhizhong, et al.
Publicado: (2024)
por: Zhang, Zhizhong, et al.
Publicado: (2024)
PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection
por: Li, Xiaofan, et al.
Publicado: (2024)
por: Li, Xiaofan, et al.
Publicado: (2024)
One-for-More: Continual Diffusion Model for Anomaly Detection
por: Li, Xiaofan, et al.
Publicado: (2025)
por: Li, Xiaofan, et al.
Publicado: (2025)
YouTube-Occ: Learning Indoor 3D Semantic Occupancy Prediction from YouTube Videos
por: Chen, Haoming, et al.
Publicado: (2025)
por: Chen, Haoming, et al.
Publicado: (2025)
Learning Commonality, Divergence and Variety for Unsupervised Visible-Infrared Person Re-identification
por: Shi, Jiangming, et al.
Publicado: (2024)
por: Shi, Jiangming, et al.
Publicado: (2024)
Multi-Memory Matching for Unsupervised Visible-Infrared Person Re-Identification
por: Shi, Jiangming, et al.
Publicado: (2024)
por: Shi, Jiangming, et al.
Publicado: (2024)
Robust Pseudo-label Learning with Neighbor Relation for Unsupervised Visible-Infrared Person Re-Identification
por: Yin, Xiangbo, et al.
Publicado: (2024)
por: Yin, Xiangbo, et al.
Publicado: (2024)
From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning
por: Wang, Sen, et al.
Publicado: (2025)
por: Wang, Sen, et al.
Publicado: (2025)
PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation
por: Tan, Wenbin, et al.
Publicado: (2026)
por: Tan, Wenbin, et al.
Publicado: (2026)
Fusion-then-Distillation: Toward Cross-modal Positive Distillation for Domain Adaptive 3D Semantic Segmentation
por: Wu, Yao, et al.
Publicado: (2024)
por: Wu, Yao, et al.
Publicado: (2024)
GSCompleter: A Distillation-Free Plugin for Metric-Aware 3D Gaussian Splatting Completion in Seconds
por: Gao, Ao, et al.
Publicado: (2026)
por: Gao, Ao, et al.
Publicado: (2026)
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
por: Jingyu, Gong, et al.
Publicado: (2025)
por: Jingyu, Gong, et al.
Publicado: (2025)
DEMOS: Dynamic Environment Motion Synthesis in 3D Scenes via Local Spherical-BEV Perception
por: Gong, Jingyu, et al.
Publicado: (2024)
por: Gong, Jingyu, et al.
Publicado: (2024)
Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation
por: Li, Jiahao, et al.
Publicado: (2026)
por: Li, Jiahao, et al.
Publicado: (2026)
Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline
por: Bai, Weikang, et al.
Publicado: (2025)
por: Bai, Weikang, et al.
Publicado: (2025)
3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation
por: Chen, Wenxin, et al.
Publicado: (2025)
por: Chen, Wenxin, et al.
Publicado: (2025)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
por: Lin, Jiawen, et al.
Publicado: (2025)
por: Lin, Jiawen, et al.
Publicado: (2025)
GEOcc: Geometrically Enhanced 3D Occupancy Network with Implicit-Explicit Depth Fusion and Contextual Self-Supervision
por: Tan, Xin, et al.
Publicado: (2024)
por: Tan, Xin, et al.
Publicado: (2024)
VTGaussian-SLAM: RGBD SLAM for Large Scale Scenes with Splatting View-Tied 3D Gaussians
por: Hu, Pengchong, et al.
Publicado: (2025)
por: Hu, Pengchong, et al.
Publicado: (2025)
MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training
por: He, Xingyi, et al.
Publicado: (2025)
por: He, Xingyi, et al.
Publicado: (2025)
Fast-BEV: A Fast and Strong Bird's-Eye View Perception Baseline
por: Li, Yangguang, et al.
Publicado: (2023)
por: Li, Yangguang, et al.
Publicado: (2023)
S2GS: Streaming Semantic Gaussian Splatting for Online Scene Understanding and Reconstruction
por: Zhang, Renhe, et al.
Publicado: (2026)
por: Zhang, Renhe, et al.
Publicado: (2026)
Multi-Modal Building Change Detection for Large-Scale Small Changes: Benchmark and Baseline
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Data-free Distillation with Degradation-prompt Diffusion for Multi-weather Image Restoration
por: Wang, Pei, et al.
Publicado: (2024)
por: Wang, Pei, et al.
Publicado: (2024)
Failure Cases Are Better Learned But Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial Training
por: Wang, Yanyun, et al.
Publicado: (2025)
por: Wang, Yanyun, et al.
Publicado: (2025)
Dual-Scale Transformer for Large-Scale Single-Pixel Imaging
por: Qu, Gang, et al.
Publicado: (2024)
por: Qu, Gang, et al.
Publicado: (2024)
Monocular 3D Lane Detection via Structure Uncertainty-Aware Network with Curve-Point Queries
por: Liu, Ruixin, et al.
Publicado: (2025)
por: Liu, Ruixin, et al.
Publicado: (2025)
Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection
por: Qiu, Mei, et al.
Publicado: (2026)
por: Qiu, Mei, et al.
Publicado: (2026)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
por: Lin, Nie, et al.
Publicado: (2024)
por: Lin, Nie, et al.
Publicado: (2024)
Towards Student Actions in Classroom Scenes: New Dataset and Baseline
por: Tan, Zhuolin, et al.
Publicado: (2024)
por: Tan, Zhuolin, et al.
Publicado: (2024)
Switchable Token-Specific Codebook Quantization For Face Image Compression
por: Wang, Yongbo, et al.
Publicado: (2025)
por: Wang, Yongbo, et al.
Publicado: (2025)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
por: Xu, Mingze, et al.
Publicado: (2024)
por: Xu, Mingze, et al.
Publicado: (2024)
Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation
por: Li, Jiahao, et al.
Publicado: (2025)
por: Li, Jiahao, et al.
Publicado: (2025)
Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective
por: Li, Jiahao, et al.
Publicado: (2025)
por: Li, Jiahao, et al.
Publicado: (2025)
Training-Free Anomaly Generation via Dual-Attention Enhancement in Diffusion Model
por: Zuo, Zuo, et al.
Publicado: (2025)
por: Zuo, Zuo, et al.
Publicado: (2025)
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
por: Zhang, Jiaxin, et al.
Publicado: (2026)
por: Zhang, Jiaxin, et al.
Publicado: (2026)
FastLGS: Speeding up Language Embedded Gaussians with Feature Grid Mapping
por: Ji, Yuzhou, et al.
Publicado: (2024)
por: Ji, Yuzhou, et al.
Publicado: (2024)
Ejemplares similares
-
Exploring the Untouched Sweeps for Conflict-Aware 3D Segmentation Pretraining
por: Sun, Tianfang, et al.
Publicado: (2024) -
Beyond the Label Itself: Latent Labels Enhance Semi-supervised Point Cloud Panoptic Segmentation
por: Chen, Yujun, et al.
Publicado: (2023) -
COTR: Compact Occupancy TRansformer for Vision-based 3D Occupancy Prediction
por: Ma, Qihang, et al.
Publicado: (2023) -
Mutual Information Guided Optimal Transport for Unsupervised Visible-Infrared Person Re-identification
por: Zhang, Zhizhong, et al.
Publicado: (2024) -
PromptAD: Learning Prompts with only Normal Samples for Few-Shot Anomaly Detection
por: Li, Xiaofan, et al.
Publicado: (2024)