BLOCK: An Open-Source Bi-Stage MLLM Character-to-Skin Pipeline for Minecraft
Fuente:
arXiv
Saved in:
| Main Author: | Guo, Hengquan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
by: Guo, Junliang, et al.
Published: (2025)
by: Guo, Junliang, et al.
Published: (2025)
TrustSkin: A Fairness Pipeline for Trustworthy Facial Affect Analysis Across Skin Tone
by: Cabanas, Ana M., et al.
Published: (2025)
by: Cabanas, Ana M., et al.
Published: (2025)
SkinFlow: Efficient Information Transmission for Open Dermatological Diagnosis via Dynamic Visual Encoding and Staged RL
by: Liu, Lijun, et al.
Published: (2026)
by: Liu, Lijun, et al.
Published: (2026)
Tri-Reader: An Open-Access, Multi-Stage AI Pipeline for First-Pass Lung Nodule Annotation in Screening CT
by: Tushar, Fakrul Islam, et al.
Published: (2026)
by: Tushar, Fakrul Islam, et al.
Published: (2026)
Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
by: Jain, Chayan, et al.
Published: (2025)
by: Jain, Chayan, et al.
Published: (2025)
LIME: Less Is More for MLLM Evaluation
by: Zhu, King, et al.
Published: (2024)
by: Zhu, King, et al.
Published: (2024)
BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation
by: Wang, Miaowei, et al.
Published: (2026)
by: Wang, Miaowei, et al.
Published: (2026)
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
by: Xiong, Chuyan, et al.
Published: (2024)
by: Xiong, Chuyan, et al.
Published: (2024)
MLLM-based Textual Explanations for Face Comparison
by: Sony, Redwan, et al.
Published: (2026)
by: Sony, Redwan, et al.
Published: (2026)
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
by: Zhang, Zhenxing, et al.
Published: (2025)
by: Zhang, Zhenxing, et al.
Published: (2025)
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
Open-Sora Plan: Open-Source Large Video Generation Model
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
by: Chen, Yuan, et al.
Published: (2025)
by: Chen, Yuan, et al.
Published: (2025)
Robust MLLM Unlearning via Visual Knowledge Distillation
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
RampNet: A Two-Stage Pipeline for Bootstrapping Curb Ramp Detection in Streetscape Images from Open Government Metadata
by: O'Meara, John S., et al.
Published: (2025)
by: O'Meara, John S., et al.
Published: (2025)
MLLM-CL: Continual Learning for Multimodal Large Language Models
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
Character-Adapter: Prompt-Guided Region Control for High-Fidelity Character Customization
by: Ma, Yuhang, et al.
Published: (2024)
by: Ma, Yuhang, et al.
Published: (2024)
SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters
by: Shah, Arya, et al.
Published: (2026)
by: Shah, Arya, et al.
Published: (2026)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
by: Chi, Donghwan, et al.
Published: (2025)
by: Chi, Donghwan, et al.
Published: (2025)
FreeVA: Offline MLLM as Training-Free Video Assistant
by: Wu, Wenhao
Published: (2024)
by: Wu, Wenhao
Published: (2024)
Unified and Dynamic Graph for Temporal Character Grouping in Long Videos
by: Shu, Xiujun, et al.
Published: (2023)
by: Shu, Xiujun, et al.
Published: (2023)
AnatomicalNets: A Multi-Structure Segmentation and Contour-Based Distance Estimation Pipeline for Clinically Grounded Lung Cancer T-Staging
by: Chowdhury, Saniah Kayenat, et al.
Published: (2025)
by: Chowdhury, Saniah Kayenat, et al.
Published: (2025)
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
by: Luo, Zhuoyan, et al.
Published: (2024)
by: Luo, Zhuoyan, et al.
Published: (2024)
Zero-shot High-fidelity and Pose-controllable Character Animation
by: Zhu, Bingwen, et al.
Published: (2024)
by: Zhu, Bingwen, et al.
Published: (2024)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026)
by: Tzachor, Issar, et al.
Published: (2026)
Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
by: Li, Zuoou, et al.
Published: (2025)
by: Li, Zuoou, et al.
Published: (2025)
Leveraging MLLM Embeddings and Attribute Smoothing for Compositional Zero-Shot Learning
by: Yan, Xudong, et al.
Published: (2024)
by: Yan, Xudong, et al.
Published: (2024)
HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
by: Du, Yifan, et al.
Published: (2025)
by: Du, Yifan, et al.
Published: (2025)
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
by: Yue, Zhengrong, et al.
Published: (2025)
by: Yue, Zhengrong, et al.
Published: (2025)
An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation
by: Li, Bingyu, et al.
Published: (2026)
by: Li, Bingyu, et al.
Published: (2026)
StageDesigner: Artistic Stage Generation for Scenography via Theater Scripts
by: Gan, Zhaoxing, et al.
Published: (2025)
by: Gan, Zhaoxing, et al.
Published: (2025)
Animate Any Character in Any World
by: Wang, Yitong, et al.
Published: (2025)
by: Wang, Yitong, et al.
Published: (2025)
No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
by: Dai, Zunkai, et al.
Published: (2026)
by: Dai, Zunkai, et al.
Published: (2026)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
by: Yu, An, et al.
Published: (2025)
by: Yu, An, et al.
Published: (2025)
Exploring the Impact of Skin Color on Skin Lesion Segmentation
by: Paxton, Kuniko, et al.
Published: (2026)
by: Paxton, Kuniko, et al.
Published: (2026)
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness
by: Sun, Yueru, et al.
Published: (2026)
by: Sun, Yueru, et al.
Published: (2026)
Similar Items
-
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
by: Guo, Junliang, et al.
Published: (2025) -
TrustSkin: A Fairness Pipeline for Trustworthy Facial Affect Analysis Across Skin Tone
by: Cabanas, Ana M., et al.
Published: (2025) -
SkinFlow: Efficient Information Transmission for Open Dermatological Diagnosis via Dynamic Visual Encoding and Staged RL
by: Liu, Lijun, et al.
Published: (2026) -
Tri-Reader: An Open-Access, Multi-Stage AI Pipeline for First-Pass Lung Nodule Annotation in Screening CT
by: Tushar, Fakrul Islam, et al.
Published: (2026) -
Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
by: Jain, Chayan, et al.
Published: (2025)