Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Pham, Duc-Hai, Nguyen, Duc-Dung, Pham, Anh, Ho, Tuan, Nguyen, Phong, Nguyen, Khoi, Nguyen, Rang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation
by: Pham, Duc-Hai, et al.
Published: (2024)
by: Pham, Duc-Hai, et al.
Published: (2024)
Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance
by: Nguyen, Phuc D. A., et al.
Published: (2023)
by: Nguyen, Phuc D. A., et al.
Published: (2023)
Semise: Semi-supervised learning for severity representation in medical image
by: Tran, Dung T., et al.
Published: (2025)
by: Tran, Dung T., et al.
Published: (2025)
InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting
by: Vu, Duc, et al.
Published: (2026)
by: Vu, Duc, et al.
Published: (2026)
Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Tracking
by: Nguyen, Phuc, et al.
Published: (2024)
by: Nguyen, Phuc, et al.
Published: (2024)
VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal
by: Do, Pham Khai Nguyen, et al.
Published: (2025)
by: Do, Pham Khai Nguyen, et al.
Published: (2025)
FlexEdit: Flexible and Controllable Diffusion-based Object-centric Image Editing
by: Nguyen, Trong-Tung, et al.
Published: (2024)
by: Nguyen, Trong-Tung, et al.
Published: (2024)
SepFormer: Coarse-to-fine Separator Regression Network for Table Structure Recognition
by: Nguyen, Nam Quan, et al.
Published: (2025)
by: Nguyen, Nam Quan, et al.
Published: (2025)
OE3DIS: Open-Ended 3D Point Cloud Instance Segmentation
by: Nguyen, Phuc D. A., et al.
Published: (2024)
by: Nguyen, Phuc D. A., et al.
Published: (2024)
Blur2Blur: Blur Conversion for Unsupervised Image Deblurring on Unknown Domains
by: Pham, Bang-Dang, et al.
Published: (2024)
by: Pham, Bang-Dang, et al.
Published: (2024)
SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion
by: Nguyen, Trong-Tung, et al.
Published: (2024)
by: Nguyen, Trong-Tung, et al.
Published: (2024)
SwiftBrush v2: Make Your One-step Diffusion Model Better Than Its Teacher
by: Dao, Trung, et al.
Published: (2024)
by: Dao, Trung, et al.
Published: (2024)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
by: Nguyen, Viet, et al.
Published: (2024)
by: Nguyen, Viet, et al.
Published: (2024)
Cycle Training with Semi-Supervised Domain Adaptation: Bridging Accuracy and Efficiency for Real-Time Mobile Scene Detection
by: Phan-Nguyen, Huu-Phong, et al.
Published: (2025)
by: Phan-Nguyen, Huu-Phong, et al.
Published: (2025)
BALM: A Model-Agnostic Framework for Balanced Multimodal Learning under Imbalanced Missing Rates
by: Nguyen, Phuong-Anh, et al.
Published: (2026)
by: Nguyen, Phuong-Anh, et al.
Published: (2026)
Enhanced Generative Data Augmentation for Semantic Segmentation via Stronger Guidance
by: Che, Quang-Huy, et al.
Published: (2024)
by: Che, Quang-Huy, et al.
Published: (2024)
Stable Messenger: Steganography for Message-Concealed Image Generation
by: Nguyen, Quang, et al.
Published: (2023)
by: Nguyen, Quang, et al.
Published: (2023)
SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
by: Nguyen, Hung, et al.
Published: (2024)
by: Nguyen, Hung, et al.
Published: (2024)
MissBench: Benchmarking Multimodal Affective Analysis under Imbalanced Missing Modalities
by: Pham, Tien Anh, et al.
Published: (2026)
by: Pham, Tien Anh, et al.
Published: (2026)
Mono3DV: Monocular 3D Object Detection with 3D-Aware Bipartite Matching and Variational Query DeNoising
by: Vu, Kiet Dang, et al.
Published: (2026)
by: Vu, Kiet Dang, et al.
Published: (2026)
MMAP: A Multi-Magnification and Prototype-Aware Architecture for Predicting Spatial Gene Expression
by: Nguyen, Hai Dang, et al.
Published: (2025)
by: Nguyen, Hai Dang, et al.
Published: (2025)
Hierarchical Neural Collapse Detection Transformer for Class Incremental Object Detection
by: Pham, Duc Thanh, et al.
Published: (2025)
by: Pham, Duc Thanh, et al.
Published: (2025)
LiftRefine: Progressively Refined View Synthesis from 3D Lifting with Volume-Triplane Representations
by: Do, Tung, et al.
Published: (2024)
by: Do, Tung, et al.
Published: (2024)
Virtual Fusion with Contrastive Learning for Single Sensor-based Activity Recognition
by: Nguyen, Duc-Anh, et al.
Published: (2023)
by: Nguyen, Duc-Anh, et al.
Published: (2023)
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
by: Nguyen, Huu Tien, et al.
Published: (2025)
by: Nguyen, Huu Tien, et al.
Published: (2025)
LP-OVOD: Open-Vocabulary Object Detection by Linear Probing
by: Pham, Chau, et al.
Published: (2023)
by: Pham, Chau, et al.
Published: (2023)
DiverseDream: Diverse Text-to-3D Synthesis with Augmented Text Embedding
by: Tran, Uy Dieu, et al.
Published: (2023)
by: Tran, Uy Dieu, et al.
Published: (2023)
MonoVQD: Monocular 3D Object Detection with Variational Query Denoising and Self-Distillation
by: Vu, Kiet Dang, et al.
Published: (2025)
by: Vu, Kiet Dang, et al.
Published: (2025)
Synergizing Deep Learning and Biological Heuristics for Extreme Long-Tail White Blood Cell Classification
by: Nguyen, Duc T., et al.
Published: (2026)
by: Nguyen, Duc T., et al.
Published: (2026)
ModeDreamer: Mode Guiding Score Distillation for Text-to-3D Generation using Reference Image Prompts
by: Tran, Uy Dieu, et al.
Published: (2024)
by: Tran, Uy Dieu, et al.
Published: (2024)
RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses
by: Nguyen, Minh Anh, et al.
Published: (2026)
by: Nguyen, Minh Anh, et al.
Published: (2026)
Adaptive Cache Enhancement for Test-Time Adaptation of Vision-Language Models
by: Nguyen, Khanh-Binh, et al.
Published: (2025)
by: Nguyen, Khanh-Binh, et al.
Published: (2025)
Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
SwiftPie: Lightning-fast Subject-driven Image Personalization via One step Diffusion
by: Duong, Huy, et al.
Published: (2026)
by: Duong, Huy, et al.
Published: (2026)
Overthinking Causes Hallucination: Tracing Confounder Propagation in Vision Language Models
by: Shoby, Abin, et al.
Published: (2026)
by: Shoby, Abin, et al.
Published: (2026)
Domain Generalization through Spatial Relation Induction over Visual Primitives
by: Nguyen, Dat, et al.
Published: (2026)
by: Nguyen, Dat, et al.
Published: (2026)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
ESGNN: Towards Equivariant Scene Graph Neural Network for 3D Scene Understanding
by: Pham, Quang P. M., et al.
Published: (2024)
by: Pham, Quang P. M., et al.
Published: (2024)
SwiftTailor: Efficient 3D Garment Generation with Geometry Image Representation
by: Pham, Phuc, et al.
Published: (2026)
by: Pham, Phuc, et al.
Published: (2026)
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
by: Tran, Quoc-Khang, et al.
Published: (2026)
by: Tran, Quoc-Khang, et al.
Published: (2026)
Similar Items
-
SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation
by: Pham, Duc-Hai, et al.
Published: (2024) -
Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask Guidance
by: Nguyen, Phuc D. A., et al.
Published: (2023) -
Semise: Semi-supervised learning for severity representation in medical image
by: Tran, Dung T., et al.
Published: (2025) -
InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting
by: Vu, Duc, et al.
Published: (2026) -
Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Tracking
by: Nguyen, Phuc, et al.
Published: (2024)