BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
Fuente:
arXiv
Salvato in:
| Autori principali: | Hao, Shaozhe, Liu, Xuantong, Qi, Xianbiao, Zhao, Shihao, Zi, Bojia, Xiao, Rong, Han, Kai, Wong, Kwan-Yee K. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
di: Zhao, Shihao, et al.
Pubblicazione: (2024)
di: Zhao, Shihao, et al.
Pubblicazione: (2024)
ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction
di: Hao, Shaozhe, et al.
Pubblicazione: (2024)
di: Hao, Shaozhe, et al.
Pubblicazione: (2024)
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
di: Zi, Bojia, et al.
Pubblicazione: (2025)
di: Zi, Bojia, et al.
Pubblicazione: (2025)
Elucidating the design space of language models for image generation
di: Liu, Xuantong, et al.
Pubblicazione: (2024)
di: Liu, Xuantong, et al.
Pubblicazione: (2024)
Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
di: Zhao, Shihao, et al.
Pubblicazione: (2025)
di: Zhao, Shihao, et al.
Pubblicazione: (2025)
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
di: Zi, Bojia, et al.
Pubblicazione: (2025)
di: Zi, Bojia, et al.
Pubblicazione: (2025)
CiPR: An Efficient Framework with Cross-instance Positive Relations for Generalized Category Discovery
di: Hao, Shaozhe, et al.
Pubblicazione: (2023)
di: Hao, Shaozhe, et al.
Pubblicazione: (2023)
Refaçade: Editing Object with Given Reference Texture
di: Huang, Youze, et al.
Pubblicazione: (2025)
di: Huang, Youze, et al.
Pubblicazione: (2025)
Ctrl&Shift: High-Quality Geometry-Aware Object Manipulation in Visual Generation
di: Ruan, Penghui, et al.
Pubblicazione: (2026)
di: Ruan, Penghui, et al.
Pubblicazione: (2026)
CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
di: Zi, Bojia, et al.
Pubblicazione: (2024)
di: Zi, Bojia, et al.
Pubblicazione: (2024)
ArtiFade: Learning to Generate High-quality Subject from Blemished Images
di: Yang, Shuya, et al.
Pubblicazione: (2024)
di: Yang, Shuya, et al.
Pubblicazione: (2024)
Taming Transformer Without Using Learning Rate Warmup
di: Qi, Xianbiao, et al.
Pubblicazione: (2025)
di: Qi, Xianbiao, et al.
Pubblicazione: (2025)
CusConcept: Customized Visual Concept Decomposition with Diffusion Models
di: Xu, Zhi, et al.
Pubblicazione: (2024)
di: Xu, Zhi, et al.
Pubblicazione: (2024)
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos
di: Lin, Jiahao, et al.
Pubblicazione: (2025)
di: Lin, Jiahao, et al.
Pubblicazione: (2025)
SimpleGPT: Improving GPT via A Simple Normalization Strategy
di: Chen, Marco, et al.
Pubblicazione: (2026)
di: Chen, Marco, et al.
Pubblicazione: (2026)
Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
di: He, Wei, et al.
Pubblicazione: (2026)
di: He, Wei, et al.
Pubblicazione: (2026)
LooC: Effective Low-Dimensional Codebook for Compositional Vector Quantization
di: Li, Jie, et al.
Pubblicazione: (2026)
di: Li, Jie, et al.
Pubblicazione: (2026)
VipDiff: Towards Coherent and Diverse Video Inpainting via Training-free Denoising Diffusion Models
di: Xie, Chaohao, et al.
Pubblicazione: (2025)
di: Xie, Chaohao, et al.
Pubblicazione: (2025)
Adversarial Prompt Distillation for Vision-Language Models
di: Luo, Lin, et al.
Pubblicazione: (2024)
di: Luo, Lin, et al.
Pubblicazione: (2024)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
di: Yang, Dongquan, et al.
Pubblicazione: (2025)
di: Yang, Dongquan, et al.
Pubblicazione: (2025)
Delving into Muon and Beyond: Deep Analysis and Extensions
di: Qi, Xianbiao, et al.
Pubblicazione: (2026)
di: Qi, Xianbiao, et al.
Pubblicazione: (2026)
Ultra-low Power AMOLED Displays for Smart Wearable Applications: Theory and Practice
di: Lyu, Bojia
Pubblicazione: (2025)
di: Lyu, Bojia
Pubblicazione: (2025)
A Survey on 3D Human Avatar Modeling -- From Reconstruction to Generation
di: Wang, Ruihe, et al.
Pubblicazione: (2024)
di: Wang, Ruihe, et al.
Pubblicazione: (2024)
Control-oriented Clustering of Visual Latent Representation
di: Qi, Han, et al.
Pubblicazione: (2024)
di: Qi, Han, et al.
Pubblicazione: (2024)
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
di: Wei, Guoting, et al.
Pubblicazione: (2026)
di: Wei, Guoting, et al.
Pubblicazione: (2026)
InsMapper: Exploring Inner-instance Information for Vectorized HD Mapping
di: Xu, Zhenhua, et al.
Pubblicazione: (2023)
di: Xu, Zhenhua, et al.
Pubblicazione: (2023)
CLAP: Learning Transferable Binary Code Representations with Natural Language Supervision
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
Neural Normalized Cut: A Differential and Generalizable Approach for Spectral Clustering
di: He, Wei, et al.
Pubblicazione: (2025)
di: He, Wei, et al.
Pubblicazione: (2025)
Exploring a Principled Framework for Deep Subspace Clustering
di: Meng, Xianghan, et al.
Pubblicazione: (2025)
di: Meng, Xianghan, et al.
Pubblicazione: (2025)
AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation
di: Cao, Yukang, et al.
Pubblicazione: (2024)
di: Cao, Yukang, et al.
Pubblicazione: (2024)
Binary LCD Codes and Their Graph Representations
di: Ishizuka, Keita
Pubblicazione: (2024)
di: Ishizuka, Keita
Pubblicazione: (2024)
TrajSelector: Harnessing Latent Representations for Efficient and Effective Best-of-N in Large Reasoning Model
di: Yu, Bin, et al.
Pubblicazione: (2025)
di: Yu, Bin, et al.
Pubblicazione: (2025)
GR-Athena++: Binary Neutron Star Merger Simulations with Neutrino Transport
di: Daszuta, Boris, et al.
Pubblicazione: (2026)
di: Daszuta, Boris, et al.
Pubblicazione: (2026)
$\texttt{GR-Athena++}$ Simulations of Spinning Binary Black Hole Mergers
di: Shukla, Estuti, et al.
Pubblicazione: (2025)
di: Shukla, Estuti, et al.
Pubblicazione: (2025)
Construction and Fast Decoding of Binary Linear Sum-Rank-Metric Codes
di: Chen, Hao, et al.
Pubblicazione: (2023)
di: Chen, Hao, et al.
Pubblicazione: (2023)
GR-3 Technical Report
di: Cheang, Chilam, et al.
Pubblicazione: (2025)
di: Cheang, Chilam, et al.
Pubblicazione: (2025)
P‐9.17: Demura Algorithm Based on Generative Adversarial Network
di: Xi Li, et al.
Pubblicazione: (2024)
di: Xi Li, et al.
Pubblicazione: (2024)
DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D Generation
di: Huang, Yukun, et al.
Pubblicazione: (2023)
di: Huang, Yukun, et al.
Pubblicazione: (2023)
DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
di: Lv, Zhengyao, et al.
Pubblicazione: (2026)
di: Lv, Zhengyao, et al.
Pubblicazione: (2026)
PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis
di: Lv, Zhengyao, et al.
Pubblicazione: (2024)
di: Lv, Zhengyao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
di: Zhao, Shihao, et al.
Pubblicazione: (2024) -
ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction
di: Hao, Shaozhe, et al.
Pubblicazione: (2024) -
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
di: Zi, Bojia, et al.
Pubblicazione: (2025) -
Elucidating the design space of language models for image generation
di: Liu, Xuantong, et al.
Pubblicazione: (2024) -
Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
di: Zhao, Shihao, et al.
Pubblicazione: (2025)