Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Wenbo, Zhang, Lu, Hu, Ping, Ma, Liqian, Zhuge, Yunzhi, Lu, Huchuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling
by: Feng, Xiyan, et al.
Published: (2026)
by: Feng, Xiyan, et al.
Published: (2026)
DreamMix: Decoupling Object Attributes for Enhanced Editability in Customized Image Inpainting
by: Yang, Yicheng, et al.
Published: (2024)
by: Yang, Yicheng, et al.
Published: (2024)
StableIdentity: Inserting Anybody into Anywhere at First Sight
by: Wang, Qinghe, et al.
Published: (2024)
by: Wang, Qinghe, et al.
Published: (2024)
Complementary and Contrastive Learning for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
by: Zhuge, Yunzhi, et al.
Published: (2025)
by: Zhuge, Yunzhi, et al.
Published: (2025)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
by: Zhuo, Dong, et al.
Published: (2026)
by: Zhuo, Dong, et al.
Published: (2026)
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
by: Zhang, Lu, et al.
Published: (2025)
by: Zhang, Lu, et al.
Published: (2025)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding
by: Jiang, Minchao, et al.
Published: (2025)
by: Jiang, Minchao, et al.
Published: (2025)
Parameter Aware Mamba Model for Multi-task Dense Prediction
by: Yu, Xinzhuo, et al.
Published: (2025)
by: Yu, Xinzhuo, et al.
Published: (2025)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
Contrastive Gaussian Clustering: Weakly Supervised 3D Scene Segmentation
by: Silva, Myrna C., et al.
Published: (2024)
by: Silva, Myrna C., et al.
Published: (2024)
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
TESGNN: Temporal Equivariant Scene Graph Neural Networks for Efficient and Robust Multi-View 3D Scene Understanding
by: Pham, Quang P. M., et al.
Published: (2024)
by: Pham, Quang P. M., et al.
Published: (2024)
Phase-Consistent Magnetic Spectral Learning for Multi-View Clustering
by: Lu, Mingdong, et al.
Published: (2026)
by: Lu, Mingdong, et al.
Published: (2026)
VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
by: Li, Baolu, et al.
Published: (2025)
by: Li, Baolu, et al.
Published: (2025)
Learning Universal Features for Generalizable Image Forgery Localization
by: Zhao, Hengrun, et al.
Published: (2025)
by: Zhao, Hengrun, et al.
Published: (2025)
PixelGaussian: Generalizable 3D Gaussian Reconstruction from Arbitrary Views
by: Fei, Xin, et al.
Published: (2024)
by: Fei, Xin, et al.
Published: (2024)
Learning with Noisy Ground Truth: From 2D Classification to 3D Reconstruction
by: Lu, Yangdi, et al.
Published: (2024)
by: Lu, Yangdi, et al.
Published: (2024)
TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting
by: Tan, Yuyang, et al.
Published: (2026)
by: Tan, Yuyang, et al.
Published: (2026)
FedCVU: Federated Learning for Cross-View Video Understanding
by: Zhang, Shenghan, et al.
Published: (2026)
by: Zhang, Shenghan, et al.
Published: (2026)
ESGNN: Towards Equivariant Scene Graph Neural Network for 3D Scene Understanding
by: Pham, Quang P. M., et al.
Published: (2024)
by: Pham, Quang P. M., et al.
Published: (2024)
Calib3D: Calibrating Model Preferences for Reliable 3D Scene Understanding
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
Bootstrap Deep Spectral Clustering with Optimal Transport
by: Guo, Wengang, et al.
Published: (2025)
by: Guo, Wengang, et al.
Published: (2025)
SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning
by: Diao, Haiwen, et al.
Published: (2024)
by: Diao, Haiwen, et al.
Published: (2024)
Sampling 3D Gaussian Scenes in Seconds with Latent Diffusion Models
by: Henderson, Paul, et al.
Published: (2024)
by: Henderson, Paul, et al.
Published: (2024)
3D Gaussian Inpainting with Depth-Guided Cross-View Consistency
by: Huang, Sheng-Yu, et al.
Published: (2025)
by: Huang, Sheng-Yu, et al.
Published: (2025)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
by: Liu, Qing'an, et al.
Published: (2026)
by: Liu, Qing'an, et al.
Published: (2026)
Mitigating Noisy Supervision Using Synthetic Samples with Soft Labels
by: Lu, Yangdi, et al.
Published: (2024)
by: Lu, Yangdi, et al.
Published: (2024)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
BEV-IO: Enhancing Bird's-Eye-View 3D Detection with Instance Occupancy
by: Zhang, Zaibin, et al.
Published: (2023)
by: Zhang, Zaibin, et al.
Published: (2023)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
Towards Large Model Feature Coding
by: Pang, Youwei, et al.
Published: (2026)
by: Pang, Youwei, et al.
Published: (2026)
Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding
by: Mao, Ye, et al.
Published: (2026)
by: Mao, Ye, et al.
Published: (2026)
SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation
by: Thengane, Vishal, et al.
Published: (2026)
by: Thengane, Vishal, et al.
Published: (2026)
Spider: A Unified Framework for Context-dependent Concept Segmentation
by: Zhao, Xiaoqi, et al.
Published: (2024)
by: Zhao, Xiaoqi, et al.
Published: (2024)
Similar Items
-
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
by: Xiong, Haomiao, et al.
Published: (2025) -
Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling
by: Feng, Xiyan, et al.
Published: (2026) -
DreamMix: Decoupling Object Attributes for Enhanced Editability in Customized Image Inpainting
by: Yang, Yicheng, et al.
Published: (2024) -
StableIdentity: Inserting Anybody into Anywhere at First Sight
by: Wang, Qinghe, et al.
Published: (2024) -
Complementary and Contrastive Learning for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)