Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Jerry, Sun, Haowen, Gudovskiy, Denis, Nakata, Yohei, Okuno, Tomoyuki, Keutzer, Kurt, Zheng, Wenzhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ODE$_t$(ODE$_l$): Shortcutting the Time and the Length in Diffusion and Flow Models for Faster Sampling
von: Gudovskiy, Denis, et al.
Veröffentlicht: (2025)
von: Gudovskiy, Denis, et al.
Veröffentlicht: (2025)
ContextFlow++: Generalist-Specialist Flow-based Generative Models with Mixed-Variable Context Encoding
von: Gudovskiy, Denis, et al.
Veröffentlicht: (2024)
von: Gudovskiy, Denis, et al.
Veröffentlicht: (2024)
DFM: Interpolant-free Dual Flow Matching
von: Gudovskiy, Denis, et al.
Veröffentlicht: (2024)
von: Gudovskiy, Denis, et al.
Veröffentlicht: (2024)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
Split-Ensemble: Efficient OOD-aware Ensemble via Task and Model Splitting
von: Chen, Anthony, et al.
Veröffentlicht: (2023)
von: Chen, Anthony, et al.
Veröffentlicht: (2023)
Fisher-aware Quantization for DETR Detectors with Critical-category Objectives
von: Yang, Huanrui, et al.
Veröffentlicht: (2024)
von: Yang, Huanrui, et al.
Veröffentlicht: (2024)
VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness
von: Zhang, Rongyu, et al.
Veröffentlicht: (2024)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2024)
PixelGaussian: Generalizable 3D Gaussian Reconstruction from Arbitrary Views
von: Fei, Xin, et al.
Veröffentlicht: (2024)
von: Fei, Xin, et al.
Veröffentlicht: (2024)
Driv3R: Learning Dense 4D Reconstruction for Autonomous Driving
von: Fei, Xin, et al.
Veröffentlicht: (2024)
von: Fei, Xin, et al.
Veröffentlicht: (2024)
UniDrive: Towards Universal Driving Perception Across Camera Configurations
von: Li, Ye, et al.
Veröffentlicht: (2024)
von: Li, Ye, et al.
Veröffentlicht: (2024)
R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation
von: Ljungbergh, William, et al.
Veröffentlicht: (2025)
von: Ljungbergh, William, et al.
Veröffentlicht: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
von: Chen, Anthony, et al.
Veröffentlicht: (2025)
von: Chen, Anthony, et al.
Veröffentlicht: (2025)
GaussianFormer: Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
von: Huang, Yuanhui, et al.
Veröffentlicht: (2024)
von: Huang, Yuanhui, et al.
Veröffentlicht: (2024)
Instruct Large Language Models to Drive like Humans
von: Zhang, Ruijun, et al.
Veröffentlicht: (2024)
von: Zhang, Ruijun, et al.
Veröffentlicht: (2024)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
von: Nehrdich, Sebastian, et al.
Veröffentlicht: (2026)
von: Nehrdich, Sebastian, et al.
Veröffentlicht: (2026)
SliceOcc: Indoor 3D Semantic Occupancy Prediction with Vertical Slice Representation
von: Li, Jianing, et al.
Veröffentlicht: (2025)
von: Li, Jianing, et al.
Veröffentlicht: (2025)
DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving
von: Lu, Hao, et al.
Veröffentlicht: (2024)
von: Lu, Hao, et al.
Veröffentlicht: (2024)
A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision
von: Peng, Chensheng, et al.
Veröffentlicht: (2024)
von: Peng, Chensheng, et al.
Veröffentlicht: (2024)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
$\textit{S}^3$Gaussian: Self-Supervised Street Gaussians for Autonomous Driving
von: Huang, Nan, et al.
Veröffentlicht: (2024)
von: Huang, Nan, et al.
Veröffentlicht: (2024)
QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction
von: Zuo, Sicheng, et al.
Veröffentlicht: (2025)
von: Zuo, Sicheng, et al.
Veröffentlicht: (2025)
DeSiRe-GS: 4D Street Gaussians for Static-Dynamic Decomposition and Surface Reconstruction for Urban Driving Scenes
von: Peng, Chensheng, et al.
Veröffentlicht: (2024)
von: Peng, Chensheng, et al.
Veröffentlicht: (2024)
Vision-based 3D Semantic Scene Completion via Capture Dynamic Representations
von: Wang, Meng, et al.
Veröffentlicht: (2025)
von: Wang, Meng, et al.
Veröffentlicht: (2025)
Segment Any Motion in Videos
von: Huang, Nan, et al.
Veröffentlicht: (2025)
von: Huang, Nan, et al.
Veröffentlicht: (2025)
HPR3D: Hierarchical Proxy Representation for High-Fidelity 3D Reconstruction and Controllable Editing
von: Wang, Tielong, et al.
Veröffentlicht: (2025)
von: Wang, Tielong, et al.
Veröffentlicht: (2025)
3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection
von: Zhu, Haowen, et al.
Veröffentlicht: (2026)
von: Zhu, Haowen, et al.
Veröffentlicht: (2026)
ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
von: Wu, Yuqi, et al.
Veröffentlicht: (2025)
von: Wu, Yuqi, et al.
Veröffentlicht: (2025)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
von: Wu, Yuqi, et al.
Veröffentlicht: (2024)
von: Wu, Yuqi, et al.
Veröffentlicht: (2024)
Hyper3D: Efficient 3D Representation via Hybrid Triplane and Octree Feature for Enhanced 3D Shape Variational Auto-Encoders
von: Guo, Jingyu, et al.
Veröffentlicht: (2025)
von: Guo, Jingyu, et al.
Veröffentlicht: (2025)
Dual-Domain Representation Alignment: Bridging 2D and 3D Vision via Geometry-Aware Architecture Search
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction
von: Huang, Yuanhui, et al.
Veröffentlicht: (2024)
von: Huang, Yuanhui, et al.
Veröffentlicht: (2024)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
OmniIndoor3D: Comprehensive Indoor 3D Reconstruction
von: Wei, Xiaobao, et al.
Veröffentlicht: (2025)
von: Wei, Xiaobao, et al.
Veröffentlicht: (2025)
GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
von: Sun, Lin, et al.
Veröffentlicht: (2025)
von: Sun, Lin, et al.
Veröffentlicht: (2025)
Hardness-Aware Scene Synthesis for Semi-Supervised 3D Object Detection
von: Zeng, Shuai, et al.
Veröffentlicht: (2024)
von: Zeng, Shuai, et al.
Veröffentlicht: (2024)
Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
von: Bai, Weimin, et al.
Veröffentlicht: (2025)
von: Bai, Weimin, et al.
Veröffentlicht: (2025)
BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
von: Li, Peiyan, et al.
Veröffentlicht: (2025)
von: Li, Peiyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ODE$_t$(ODE$_l$): Shortcutting the Time and the Length in Diffusion and Flow Models for Faster Sampling
von: Gudovskiy, Denis, et al.
Veröffentlicht: (2025) -
ContextFlow++: Generalist-Specialist Flow-based Generative Models with Mixed-Variable Context Encoding
von: Gudovskiy, Denis, et al.
Veröffentlicht: (2024) -
DFM: Interpolant-free Dual Flow Matching
von: Gudovskiy, Denis, et al.
Veröffentlicht: (2024) -
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
von: Zhang, Yuan, et al.
Veröffentlicht: (2024) -
Split-Ensemble: Efficient OOD-aware Ensemble via Task and Model Splitting
von: Chen, Anthony, et al.
Veröffentlicht: (2023)