SGC-VQGAN: Towards Complex Scene Representation via Semantic Guided Clustering Codebook
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ding, Chenjing, Wang, Chiyu, Liu, Boshi, Guo, Xi, Tang, Weixuan, Wu, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
Physical Informed Driving World Model
von: Yang, Zhuoran, et al.
Veröffentlicht: (2024)
von: Yang, Zhuoran, et al.
Veröffentlicht: (2024)
Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
InfinityDrive: Breaking Time Limits in Driving World Models
von: Guo, Xi, et al.
Veröffentlicht: (2024)
von: Guo, Xi, et al.
Veröffentlicht: (2024)
InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation
von: Yang, Zhuoran, et al.
Veröffentlicht: (2026)
von: Yang, Zhuoran, et al.
Veröffentlicht: (2026)
MyGo: Consistent and Controllable Multi-View Driving Video Generation with Camera Control
von: Yao, Yining, et al.
Veröffentlicht: (2024)
von: Yao, Yining, et al.
Veröffentlicht: (2024)
PhysReaction: Physically Plausible Real-Time Humanoid Reaction Synthesis via Forward Dynamics Guided 4D Imitation
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
Toward Scene Graph and Layout Guided Complex 3D Scene Generation
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
von: Huang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
Vision-based 3D Semantic Scene Completion via Capture Dynamic Representations
von: Wang, Meng, et al.
Veröffentlicht: (2025)
von: Wang, Meng, et al.
Veröffentlicht: (2025)
CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration
von: Tang, Boshi, et al.
Veröffentlicht: (2025)
von: Tang, Boshi, et al.
Veröffentlicht: (2025)
Towards Imperceptible JPEG Image Hiding: Multi-range Representations-driven Adversarial Stego Generation
von: Yang, Junxue, et al.
Veröffentlicht: (2025)
von: Yang, Junxue, et al.
Veröffentlicht: (2025)
CoBooM: Codebook Guided Bootstrapping for Medical Image Representation Learning
von: Singh, Azad, et al.
Veröffentlicht: (2024)
von: Singh, Azad, et al.
Veröffentlicht: (2024)
Gap Completion in Point Cloud Scene occluded by Vehicles using SGC-Net
von: Feng, Yu, et al.
Veröffentlicht: (2024)
von: Feng, Yu, et al.
Veröffentlicht: (2024)
Cross-Modal Scene Semantic Alignment for Image Complexity Assessment
von: Luo, Yuqing, et al.
Veröffentlicht: (2025)
von: Luo, Yuqing, et al.
Veröffentlicht: (2025)
Layout2Scene: 3D Semantic Layout Guided Scene Generation via Geometry and Appearance Diffusion Priors
von: Chen, Minglin, et al.
Veröffentlicht: (2025)
von: Chen, Minglin, et al.
Veröffentlicht: (2025)
UrbanGS: Semantic-Guided Gaussian Splatting for Urban Scene Reconstruction
von: Li, Ziwen, et al.
Veröffentlicht: (2024)
von: Li, Ziwen, et al.
Veröffentlicht: (2024)
Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes
von: Guo, Diandian, et al.
Veröffentlicht: (2024)
von: Guo, Diandian, et al.
Veröffentlicht: (2024)
Stable Score Distillation for High-Quality 3D Generation
von: Tang, Boshi, et al.
Veröffentlicht: (2023)
von: Tang, Boshi, et al.
Veröffentlicht: (2023)
LG-VQ: Language-Guided Codebook Learning
von: Liang, Guotao, et al.
Veröffentlicht: (2024)
von: Liang, Guotao, et al.
Veröffentlicht: (2024)
PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor Scenes
von: Xu, Xinhua, et al.
Veröffentlicht: (2025)
von: Xu, Xinhua, et al.
Veröffentlicht: (2025)
SemanticMIM: Marring Masked Image Modeling with Semantics Compression for General Visual Representation
von: Yuan, Yike, et al.
Veröffentlicht: (2024)
von: Yuan, Yike, et al.
Veröffentlicht: (2024)
LACV-Net: Semantic Segmentation of Large-Scale Point Cloud Scene via Local Adaptive and Comprehensive VLAD
von: Zeng, Ziyin, et al.
Veröffentlicht: (2022)
von: Zeng, Ziyin, et al.
Veröffentlicht: (2022)
Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Text
von: Liang, Guotao, et al.
Veröffentlicht: (2025)
von: Liang, Guotao, et al.
Veröffentlicht: (2025)
UniVoxel: Fast Inverse Rendering by Unified Voxelization of Scene Representation
von: Wu, Shuang, et al.
Veröffentlicht: (2024)
von: Wu, Shuang, et al.
Veröffentlicht: (2024)
IG-Diff: Complex Night Scene Restoration with Illumination-Guided Diffusion Model
von: Chen, Yifan, et al.
Veröffentlicht: (2026)
von: Chen, Yifan, et al.
Veröffentlicht: (2026)
Learning Temporal 3D Semantic Scene Completion via Optical Flow Guidance
von: Wang, Meng, et al.
Veröffentlicht: (2025)
von: Wang, Meng, et al.
Veröffentlicht: (2025)
Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
von: Yang, Lei, et al.
Veröffentlicht: (2026)
von: Yang, Lei, et al.
Veröffentlicht: (2026)
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
von: Chen, Zisheng, et al.
Veröffentlicht: (2025)
von: Chen, Zisheng, et al.
Veröffentlicht: (2025)
InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving
von: Song, Ruiqi, et al.
Veröffentlicht: (2025)
von: Song, Ruiqi, et al.
Veröffentlicht: (2025)
CoDefend: Cross-Modal Collaborative Defense via Diffusion Purification and Prompt Optimization
von: Zhu, Fengling, et al.
Veröffentlicht: (2025)
von: Zhu, Fengling, et al.
Veröffentlicht: (2025)
MCR-VQGAN: A Scalable and Cost-Effective Tau PET Synthesis Approach for Alzheimer's Disease Imaging
von: Kim, Jin Young, et al.
Veröffentlicht: (2025)
von: Kim, Jin Young, et al.
Veröffentlicht: (2025)
SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detection
von: Lin, Xin, et al.
Veröffentlicht: (2025)
von: Lin, Xin, et al.
Veröffentlicht: (2025)
Hybrid Mesh-Gaussian Representation for Efficient Indoor Scene Reconstruction
von: Huang, Binxiao, et al.
Veröffentlicht: (2025)
von: Huang, Binxiao, et al.
Veröffentlicht: (2025)
Target Semantics Clustering via Text Representations for Robust Universal Domain Adaptation
von: He, Weinan, et al.
Veröffentlicht: (2025)
von: He, Weinan, et al.
Veröffentlicht: (2025)
MOC-RVQ: Multilevel Codebook-Assisted Digital Generative Semantic Communication
von: Zhou, Yingbin, et al.
Veröffentlicht: (2024)
von: Zhou, Yingbin, et al.
Veröffentlicht: (2024)
Panoptic Scene Graph Generation with Semantics-Prototype Learning
von: Li, Li, et al.
Veröffentlicht: (2023)
von: Li, Li, et al.
Veröffentlicht: (2023)
CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving
von: Zhang, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhang, Tianrui, et al.
Veröffentlicht: (2025)
GLARE: Low Light Image Enhancement via Generative Latent Feature based Codebook Retrieval
von: Zhou, Han, et al.
Veröffentlicht: (2024)
von: Zhou, Han, et al.
Veröffentlicht: (2024)
Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
von: Tang, Longxiang, et al.
Veröffentlicht: (2025)
von: Tang, Longxiang, et al.
Veröffentlicht: (2025)
Enhancing Underwater Images via Adaptive Semantic-aware Codebook Learning
von: Lin, Bosen, et al.
Veröffentlicht: (2026)
von: Lin, Bosen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation
von: Wu, Wei, et al.
Veröffentlicht: (2024) -
Physical Informed Driving World Model
von: Yang, Zhuoran, et al.
Veröffentlicht: (2024) -
Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
von: Zhu, Lei, et al.
Veröffentlicht: (2024) -
InfinityDrive: Breaking Time Limits in Driving World Models
von: Guo, Xi, et al.
Veröffentlicht: (2024) -
InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation
von: Yang, Zhuoran, et al.
Veröffentlicht: (2026)