Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
Fuente:
arXiv
Salvato in:
| Autori principali: | Fang, Ye, Sun, Zeyi, Wu, Tong, Wang, Jiaqi, Liu, Ziwei, Wetzstein, Gordon, Lin, Dahua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
di: Fang, Ye, et al.
Pubblicazione: (2025)
di: Fang, Ye, et al.
Pubblicazione: (2025)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
di: Lin, Junyan, et al.
Pubblicazione: (2026)
di: Lin, Junyan, et al.
Pubblicazione: (2026)
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
di: Yang, Shuai, et al.
Pubblicazione: (2024)
di: Yang, Shuai, et al.
Pubblicazione: (2024)
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
di: Zhang, Mengchen, et al.
Pubblicazione: (2025)
di: Zhang, Mengchen, et al.
Pubblicazione: (2025)
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
di: Sun, Zeyi, et al.
Pubblicazione: (2025)
di: Sun, Zeyi, et al.
Pubblicazione: (2025)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
Omni6D: Large-Vocabulary 3D Object Dataset for Category-Level 6D Object Pose Estimation
di: Zhang, Mengchen, et al.
Pubblicazione: (2024)
di: Zhang, Mengchen, et al.
Pubblicazione: (2024)
Video World Models with Long-term Spatial Memory
di: Wu, Tong, et al.
Pubblicazione: (2025)
di: Wu, Tong, et al.
Pubblicazione: (2025)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
di: Bin, Yi, et al.
Pubblicazione: (2024)
di: Bin, Yi, et al.
Pubblicazione: (2024)
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
GPT4Point: A Unified Framework for Point-Language Understanding and Generation
di: Qi, Zhangyang, et al.
Pubblicazione: (2023)
di: Qi, Zhangyang, et al.
Pubblicazione: (2023)
Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models
di: Li, Jinsong, et al.
Pubblicazione: (2025)
di: Li, Jinsong, et al.
Pubblicazione: (2025)
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
di: Liao, Zeyi, et al.
Pubblicazione: (2025)
di: Liao, Zeyi, et al.
Pubblicazione: (2025)
MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
di: Lei, Fei, et al.
Pubblicazione: (2025)
di: Lei, Fei, et al.
Pubblicazione: (2025)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
di: Ji, Ziwei, et al.
Pubblicazione: (2024)
di: Ji, Ziwei, et al.
Pubblicazione: (2024)
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
di: Gu, Yuzhe, et al.
Pubblicazione: (2024)
di: Gu, Yuzhe, et al.
Pubblicazione: (2024)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
di: Hua, Jiacheng, et al.
Pubblicazione: (2026)
di: Hua, Jiacheng, et al.
Pubblicazione: (2026)
AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
di: Liao, Zeyi, et al.
Pubblicazione: (2024)
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
Crossroads of Continents: Automated Artifact Extraction for Cultural Adaptation with Large Multimodal Models
di: Mukherjee, Anjishnu, et al.
Pubblicazione: (2024)
di: Mukherjee, Anjishnu, et al.
Pubblicazione: (2024)
Scaling Behavior for Large Language Models regarding Numeral Systems: An Example using Pythia
di: Zhou, Zhejian, et al.
Pubblicazione: (2024)
di: Zhou, Zhejian, et al.
Pubblicazione: (2024)
Tree-of-Table: Unleashing the Power of LLMs for Enhanced Large-Scale Table Understanding
di: Ji, Deyi, et al.
Pubblicazione: (2024)
di: Ji, Deyi, et al.
Pubblicazione: (2024)
Utilize the Flow before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning
di: Zhu, Runchuan, et al.
Pubblicazione: (2024)
di: Zhu, Runchuan, et al.
Pubblicazione: (2024)
Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
Unleashing Large Language Models' Proficiency in Zero-shot Essay Scoring
di: Lee, Sanwoo, et al.
Pubblicazione: (2024)
di: Lee, Sanwoo, et al.
Pubblicazione: (2024)
RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
di: Wu, Pengzuo, et al.
Pubblicazione: (2025)
di: Wu, Pengzuo, et al.
Pubblicazione: (2025)
Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models
di: Zhang, Che, et al.
Pubblicazione: (2024)
di: Zhang, Che, et al.
Pubblicazione: (2024)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
di: Qiao, Yuxuan, et al.
Pubblicazione: (2024)
di: Qiao, Yuxuan, et al.
Pubblicazione: (2024)
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
di: Wang, Zexuan, et al.
Pubblicazione: (2026)
di: Wang, Zexuan, et al.
Pubblicazione: (2026)
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
di: Huang, Qidong, et al.
Pubblicazione: (2024)
di: Huang, Qidong, et al.
Pubblicazione: (2024)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
di: Xing, Long, et al.
Pubblicazione: (2024)
di: Xing, Long, et al.
Pubblicazione: (2024)
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
di: Wang, Zirui, et al.
Pubblicazione: (2024)
di: Wang, Zirui, et al.
Pubblicazione: (2024)
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
di: Du, Dayou, et al.
Pubblicazione: (2024)
di: Du, Dayou, et al.
Pubblicazione: (2024)
Compress to Impress: Unleashing the Potential of Compressive Memory in Real-World Long-Term Conversations
di: Chen, Nuo, et al.
Pubblicazione: (2024)
di: Chen, Nuo, et al.
Pubblicazione: (2024)
Advantageous Parameter Expansion Training Makes Better Large Language Models
di: Gu, Naibin, et al.
Pubblicazione: (2025)
di: Gu, Naibin, et al.
Pubblicazione: (2025)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
di: Lin, Jingyang, et al.
Pubblicazione: (2025)
di: Lin, Jingyang, et al.
Pubblicazione: (2025)
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
di: Sun, Qiushi, et al.
Pubblicazione: (2025)
di: Sun, Qiushi, et al.
Pubblicazione: (2025)
Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models
di: Jiang, Songtao, et al.
Pubblicazione: (2024)
di: Jiang, Songtao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
di: Fang, Ye, et al.
Pubblicazione: (2025) -
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
di: Zhang, Yuhan, et al.
Pubblicazione: (2025) -
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
di: Lin, Junyan, et al.
Pubblicazione: (2026) -
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
di: Yang, Shuai, et al.
Pubblicazione: (2024) -
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
di: Zhang, Mengchen, et al.
Pubblicazione: (2025)