BIFRÖST: 3D-Aware Image compositing with Language Instructions
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Lingxiao, Gong, Kaixiong, Li, Weihong, Dai, Xili, Chen, Tao, Yuan, Xiaojun, Yue, Xiangyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image Generation
di: Li, Lingxiao, et al.
Pubblicazione: (2024)
di: Li, Lingxiao, et al.
Pubblicazione: (2024)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
BAFNet: Bilateral Attention Fusion Network for Lightweight Semantic Segmentation of Urban Remote Sensing Images
di: Wang, Wentao, et al.
Pubblicazione: (2024)
di: Wang, Wentao, et al.
Pubblicazione: (2024)
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
di: Li, Lingxiao, et al.
Pubblicazione: (2025)
di: Li, Lingxiao, et al.
Pubblicazione: (2025)
Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model
di: Huang, Yaxuan, et al.
Pubblicazione: (2025)
di: Huang, Yaxuan, et al.
Pubblicazione: (2025)
FairGen: Enhancing Fairness in Text-to-Image Diffusion Models via Self-Discovering Latent Directions
di: Jiang, Yilei, et al.
Pubblicazione: (2024)
di: Jiang, Yilei, et al.
Pubblicazione: (2024)
OneLLM: One Framework to Align All Modalities with Language
di: Han, Jiaming, et al.
Pubblicazione: (2023)
di: Han, Jiaming, et al.
Pubblicazione: (2023)
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
di: Chu, Tianzhe, et al.
Pubblicazione: (2023)
di: Chu, Tianzhe, et al.
Pubblicazione: (2023)
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
di: Wang, Xiaoye, et al.
Pubblicazione: (2025)
di: Wang, Xiaoye, et al.
Pubblicazione: (2025)
Self-Tuning Self-Supervised Image Anomaly Detection
di: Yoo, Jaemin, et al.
Pubblicazione: (2023)
di: Yoo, Jaemin, et al.
Pubblicazione: (2023)
AmorLIP: Efficient Language-Image Pretraining via Amortization
di: Sun, Haotian, et al.
Pubblicazione: (2025)
di: Sun, Haotian, et al.
Pubblicazione: (2025)
PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware Pretraining
di: Li, Kecen, et al.
Pubblicazione: (2023)
di: Li, Kecen, et al.
Pubblicazione: (2023)
Taming Transformer Without Using Learning Rate Warmup
di: Qi, Xianbiao, et al.
Pubblicazione: (2025)
di: Qi, Xianbiao, et al.
Pubblicazione: (2025)
Native-Resolution Image Synthesis
di: Wang, Zidong, et al.
Pubblicazione: (2025)
di: Wang, Zidong, et al.
Pubblicazione: (2025)
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models
di: Li, Xu, et al.
Pubblicazione: (2024)
di: Li, Xu, et al.
Pubblicazione: (2024)
RadarOcc: Robust 3D Occupancy Prediction with 4D Imaging Radar
di: Ding, Fangqiang, et al.
Pubblicazione: (2024)
di: Ding, Fangqiang, et al.
Pubblicazione: (2024)
Video-R1: Reinforcing Video Reasoning in MLLMs
di: Feng, Kaituo, et al.
Pubblicazione: (2025)
di: Feng, Kaituo, et al.
Pubblicazione: (2025)
The Solution for Language-Enhanced Image New Category Discovery
di: Xu, Haonan, et al.
Pubblicazione: (2024)
di: Xu, Haonan, et al.
Pubblicazione: (2024)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
di: Yu, Hanxun, et al.
Pubblicazione: (2025)
di: Yu, Hanxun, et al.
Pubblicazione: (2025)
A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image Classification
di: Xiong, Xiangyu, et al.
Pubblicazione: (2023)
di: Xiong, Xiangyu, et al.
Pubblicazione: (2023)
EMR-Merging: Tuning-Free High-Performance Model Merging
di: Huang, Chenyu, et al.
Pubblicazione: (2024)
di: Huang, Chenyu, et al.
Pubblicazione: (2024)
Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing
di: He, Jingxuan, et al.
Pubblicazione: (2026)
di: He, Jingxuan, et al.
Pubblicazione: (2026)
Multi-level Cross-modal Alignment for Image Clustering
di: Qiu, Liping, et al.
Pubblicazione: (2024)
di: Qiu, Liping, et al.
Pubblicazione: (2024)
OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies
di: Chen, Runnan, et al.
Pubblicazione: (2024)
di: Chen, Runnan, et al.
Pubblicazione: (2024)
Cost-Aware Routing for Efficient Text-To-Image Generation
di: Li, Qinchan, et al.
Pubblicazione: (2025)
di: Li, Qinchan, et al.
Pubblicazione: (2025)
ConsistEdit: Highly Consistent and Precise Training-free Visual Editing
di: Yin, Zixin, et al.
Pubblicazione: (2025)
di: Yin, Zixin, et al.
Pubblicazione: (2025)
Refining 3D Medical Segmentation with Verbal Instruction
di: Xie, Kangxian, et al.
Pubblicazione: (2026)
di: Xie, Kangxian, et al.
Pubblicazione: (2026)
Language-Image Models with 3D Understanding
di: Cho, Jang Hyun, et al.
Pubblicazione: (2024)
di: Cho, Jang Hyun, et al.
Pubblicazione: (2024)
Uncertainty-Aware Vision-Language Segmentation for Medical Imaging
di: Das, Aryan, et al.
Pubblicazione: (2026)
di: Das, Aryan, et al.
Pubblicazione: (2026)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
di: Safaei, Bardia, et al.
Pubblicazione: (2025)
di: Safaei, Bardia, et al.
Pubblicazione: (2025)
CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model
di: Wang, Zhengyi, et al.
Pubblicazione: (2024)
di: Wang, Zhengyi, et al.
Pubblicazione: (2024)
pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction
di: Charatan, David, et al.
Pubblicazione: (2023)
di: Charatan, David, et al.
Pubblicazione: (2023)
SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
RAPiD-Seg: Range-Aware Pointwise Distance Distribution Networks for 3D LiDAR Segmentation
di: Li, Li, et al.
Pubblicazione: (2024)
di: Li, Li, et al.
Pubblicazione: (2024)
3D Congealing: 3D-Aware Image Alignment in the Wild
di: Zhang, Yunzhi, et al.
Pubblicazione: (2024)
di: Zhang, Yunzhi, et al.
Pubblicazione: (2024)
SmartFRZ: An Efficient Training Framework using Attention-Based Layer Freezing
di: Li, Sheng, et al.
Pubblicazione: (2024)
di: Li, Sheng, et al.
Pubblicazione: (2024)
Isotropic3D: Image-to-3D Generation Based on a Single CLIP Embedding
di: Liu, Pengkun, et al.
Pubblicazione: (2024)
di: Liu, Pengkun, et al.
Pubblicazione: (2024)
ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware Prompts
di: Petrov, Dmitry, et al.
Pubblicazione: (2024)
di: Petrov, Dmitry, et al.
Pubblicazione: (2024)
Concept-Aware Batch Sampling Improves Language-Image Pretraining
di: Ghosh, Adhiraj, et al.
Pubblicazione: (2025)
di: Ghosh, Adhiraj, et al.
Pubblicazione: (2025)
Empowering Large Language Models with 3D Situation Awareness
di: Yuan, Zhihao, et al.
Pubblicazione: (2025)
di: Yuan, Zhihao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image Generation
di: Li, Lingxiao, et al.
Pubblicazione: (2024) -
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024) -
BAFNet: Bilateral Attention Fusion Network for Lightweight Semantic Segmentation of Urban Remote Sensing Images
di: Wang, Wentao, et al.
Pubblicazione: (2024) -
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
di: Li, Lingxiao, et al.
Pubblicazione: (2025) -
Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model
di: Huang, Yaxuan, et al.
Pubblicazione: (2025)