Model LEGO: Creating Models Like Disassembling and Assembling Building Blocks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hu, Jiacong, Gao, Jing, Ye, Jingwen, Gao, Yang, Wang, Xingen, Feng, Zunlei, Song, Mingli |
|---|---|
| Format: | Preprint |
| Publié: |
2022
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Token-Level Inference-Time Alignment for Vision-Language Models
par: Chen, Kejia, et autres
Publié: (2025)
par: Chen, Kejia, et autres
Publié: (2025)
Improving Adversarial Robustness via Feature Pattern Consistency Constraint
par: Hu, Jiacong, et autres
Publié: (2024)
par: Hu, Jiacong, et autres
Publié: (2024)
Efficient and Comprehensive Feature Extraction in Large Vision-Language Model for Pathology Analysis
par: Zhang, Shengxuming, et autres
Publié: (2024)
par: Zhang, Shengxuming, et autres
Publié: (2024)
SHAPE : Self-Improved Visual Preference Alignment by Iteratively Generating Holistic Winner
par: Chen, Kejia, et autres
Publié: (2025)
par: Chen, Kejia, et autres
Publié: (2025)
Angle Robustness Unmanned Aerial Vehicle Navigation in GNSS-Denied Scenarios
par: Wang, Yuxin, et autres
Publié: (2024)
par: Wang, Yuxin, et autres
Publié: (2024)
Deep Feature Response Discriminative Calibration
par: Xu, Wenxiang, et autres
Publié: (2024)
par: Xu, Wenxiang, et autres
Publié: (2024)
On the Concept Trustworthiness in Concept Bottleneck Models
par: Huang, Qihan, et autres
Publié: (2024)
par: Huang, Qihan, et autres
Publié: (2024)
RS3DBench: A Comprehensive Benchmark for 3D Spatial Perception in Remote Sensing
par: Wang, Jiayu, et autres
Publié: (2025)
par: Wang, Jiayu, et autres
Publié: (2025)
Dataset Ownership Verification in Contrastive Pre-trained Models
par: Xie, Yuechen, et autres
Publié: (2025)
par: Xie, Yuechen, et autres
Publié: (2025)
Knowledge Amalgamation for Object Detection with Transformers
par: Zhang, Haofei, et autres
Publié: (2022)
par: Zhang, Haofei, et autres
Publié: (2022)
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants
par: Huang, Haochen, et autres
Publié: (2025)
par: Huang, Haochen, et autres
Publié: (2025)
LG-CAV: Train Any Concept Activation Vector with Language Guidance
par: Huang, Qihan, et autres
Publié: (2024)
par: Huang, Qihan, et autres
Publié: (2024)
$D^3$-RSMDE: 40$\times$ Faster and High-Fidelity Remote Sensing Monocular Depth Estimation
par: Wang, Ruizhi, et autres
Publié: (2026)
par: Wang, Ruizhi, et autres
Publié: (2026)
Diffusion Model Quantization: A Review
par: Zeng, Qian, et autres
Publié: (2025)
par: Zeng, Qian, et autres
Publié: (2025)
A Large-scale Universal Evaluation Benchmark For Face Forgery Detection
par: Bei, Yijun, et autres
Publié: (2024)
par: Bei, Yijun, et autres
Publié: (2024)
SEW: Self-calibration Enhanced Whole Slide Pathology Image Analysis
par: Luo, Haoming, et autres
Publié: (2024)
par: Luo, Haoming, et autres
Publié: (2024)
Adaptive Forensic Feature Refinement via Intrinsic Importance Perception
par: Yang, Jiazhen, et autres
Publié: (2026)
par: Yang, Jiazhen, et autres
Publié: (2026)
Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?
par: Xie, Yuechen, et autres
Publié: (2025)
par: Xie, Yuechen, et autres
Publié: (2025)
The Effectiveness of a Simplified Model Structure for Crowd Counting
par: Chen, Lei, et autres
Publié: (2024)
par: Chen, Lei, et autres
Publié: (2024)
Timestep-Aware Block Masking for Efficient Diffusion Model Inference
par: He, Haodong, et autres
Publié: (2026)
par: He, Haodong, et autres
Publié: (2026)
Sampling-Aware Quantization for Diffusion Models
par: Zeng, Qian, et autres
Publié: (2025)
par: Zeng, Qian, et autres
Publié: (2025)
LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction
par: Qian, Kangan, et autres
Publié: (2025)
par: Qian, Kangan, et autres
Publié: (2025)
Training-Free Pretrained Model Merging
par: Xu, Zhengqi, et autres
Publié: (2024)
par: Xu, Zhengqi, et autres
Publié: (2024)
Learning Stackable and Skippable LEGO Bricks for Efficient, Reconfigurable, and Variable-Resolution Diffusion Modeling
par: Zheng, Huangjie, et autres
Publié: (2023)
par: Zheng, Huangjie, et autres
Publié: (2023)
Seer: Language Instructed Video Prediction with Latent Diffusion Models
par: Gu, Xianfan, et autres
Publié: (2023)
par: Gu, Xianfan, et autres
Publié: (2023)
ProtoPFormer: Concentrating on Prototypical Parts in Vision Transformers for Interpretable Image Recognition
par: Xue, Mengqi, et autres
Publié: (2022)
par: Xue, Mengqi, et autres
Publié: (2022)
Unit: Building Unit Detection Dataset
par: Zhai, Haozhou, et autres
Publié: (2025)
par: Zhai, Haozhou, et autres
Publié: (2025)
MPM: A Unified 2D-3D Human Pose Representation via Masked Pose Modeling
par: Zhang, Zhenyu, et autres
Publié: (2023)
par: Zhang, Zhenyu, et autres
Publié: (2023)
Leveraging Segment Anything Model in Identifying Buildings within Refugee Camps (SAM4Refugee) from Satellite Imagery for Humanitarian Operations
par: Gao, Yunya
Publié: (2024)
par: Gao, Yunya
Publié: (2024)
Can Large Multimodal Models Inspect Buildings? A Hierarchical Benchmark for Structural Pathology Reasoning
par: Zhong, Hui, et autres
Publié: (2026)
par: Zhong, Hui, et autres
Publié: (2026)
DMM: Building a Versatile Image Generation Model via Distillation-Based Model Merging
par: Song, Tianhui, et autres
Publié: (2025)
par: Song, Tianhui, et autres
Publié: (2025)
Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
par: Bourceanu, Radu-Andrei, et autres
Publié: (2025)
par: Bourceanu, Radu-Andrei, et autres
Publié: (2025)
Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
par: Wei, Hongyang, et autres
Publié: (2025)
par: Wei, Hongyang, et autres
Publié: (2025)
Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model
par: Zhou, Hang, et autres
Publié: (2024)
par: Zhou, Hang, et autres
Publié: (2024)
Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models
par: Xiong, Lexiang, et autres
Publié: (2026)
par: Xiong, Lexiang, et autres
Publié: (2026)
Med-LEGO: Editing and Adapting toward Generalist Medical Image Diagnosis
par: Zhu, Yitao, et autres
Publié: (2025)
par: Zhu, Yitao, et autres
Publié: (2025)
Building Vision Models upon Heat Conduction
par: Wang, Zhaozhi, et autres
Publié: (2024)
par: Wang, Zhaozhi, et autres
Publié: (2024)
DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing
par: Feng, Kailai, et autres
Publié: (2026)
par: Feng, Kailai, et autres
Publié: (2026)
A-SDM: Accelerating Stable Diffusion through Model Assembly and Feature Inheritance Strategies
par: Zhu, Jinchao, et autres
Publié: (2024)
par: Zhu, Jinchao, et autres
Publié: (2024)
Recent Advances in Embedding Methods for Multi-Object Tracking: A Survey
par: Wang, Gaoang, et autres
Publié: (2022)
par: Wang, Gaoang, et autres
Publié: (2022)
Documents similaires
-
Token-Level Inference-Time Alignment for Vision-Language Models
par: Chen, Kejia, et autres
Publié: (2025) -
Improving Adversarial Robustness via Feature Pattern Consistency Constraint
par: Hu, Jiacong, et autres
Publié: (2024) -
Efficient and Comprehensive Feature Extraction in Large Vision-Language Model for Pathology Analysis
par: Zhang, Shengxuming, et autres
Publié: (2024) -
SHAPE : Self-Improved Visual Preference Alignment by Iteratively Generating Holistic Winner
par: Chen, Kejia, et autres
Publié: (2025) -
Angle Robustness Unmanned Aerial Vehicle Navigation in GNSS-Denied Scenarios
par: Wang, Yuxin, et autres
Publié: (2024)