Unified Personalized Understanding, Generating and Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Yu, Lin, Tianwei, Zhu, Ruike, Yuan, Yuqian, Zheng, Haoyu, Liang, Liang, Zhang, Wenqiao, Shao, Feifei, Li, Haoyuan, He, Wanggui, Jiang, Hao, Zhuang, Yueting |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IAD-Unify: A Region-Grounded Unified Model for Industrial Anomaly Segmentation, Understanding, and Generation
by: Zheng, Haoyu, et al.
Published: (2026)
by: Zheng, Haoyu, et al.
Published: (2026)
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
by: Lin, Tianwei, et al.
Published: (2025)
by: Lin, Tianwei, et al.
Published: (2025)
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
by: Lin, Tianwei, et al.
Published: (2024)
by: Lin, Tianwei, et al.
Published: (2024)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Boosting Private Domain Understanding of Efficient MLLMs: A Tuning-free, Adaptive, Universal Prompt Optimization Framework
by: Liu, Jiang, et al.
Published: (2024)
by: Liu, Jiang, et al.
Published: (2024)
Fast Thinking for Large Language Models
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
by: Zhang, Wenqiao, et al.
Published: (2024)
by: Zhang, Wenqiao, et al.
Published: (2024)
MAU-GPT: Enhancing Multi-type Industrial Anomaly Understanding via Anomaly-aware and Generalist Experts Adaptation
by: Wang, Zhuonan, et al.
Published: (2026)
by: Wang, Zhuonan, et al.
Published: (2026)
HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding
by: Xie, Yihan, et al.
Published: (2025)
by: Xie, Yihan, et al.
Published: (2025)
MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation
by: Zheng, Haoyu, et al.
Published: (2026)
by: Zheng, Haoyu, et al.
Published: (2026)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
by: Li, Sijing, et al.
Published: (2025)
by: Li, Sijing, et al.
Published: (2025)
Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs
by: Dai, Yang, et al.
Published: (2025)
by: Dai, Yang, et al.
Published: (2025)
PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
by: Zheng, Haoyu, et al.
Published: (2026)
by: Zheng, Haoyu, et al.
Published: (2026)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
by: Gao, Mingjian, et al.
Published: (2026)
by: Gao, Mingjian, et al.
Published: (2026)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
by: Cao, Jie, et al.
Published: (2025)
by: Cao, Jie, et al.
Published: (2025)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
by: Wang, Xierui, et al.
Published: (2024)
by: Wang, Xierui, et al.
Published: (2024)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
by: Zheng, Haoyu, et al.
Published: (2025)
by: Zheng, Haoyu, et al.
Published: (2025)
DuetRAG: Collaborative Retrieval-Augmented Generation
by: Jiao, Dian, et al.
Published: (2024)
by: Jiao, Dian, et al.
Published: (2024)
AnyMS: Bottom-up Attention Decoupling for Layout-guided and Training-free Multi-subject Customization
by: Yu, Binhe, et al.
Published: (2025)
by: Yu, Binhe, et al.
Published: (2025)
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
by: Huang, Qihan, et al.
Published: (2024)
by: Huang, Qihan, et al.
Published: (2024)
Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
by: Huang, Hongzhe, et al.
Published: (2024)
by: Huang, Hongzhe, et al.
Published: (2024)
Enhancing Post-Training Quantization via Future Activation Awareness
by: Lv, Zheqi, et al.
Published: (2026)
by: Lv, Zheqi, et al.
Published: (2026)
Bridging Local Details and Global Context in Text-Attributed Graphs
by: Wang, Yaoke, et al.
Published: (2024)
by: Wang, Yaoke, et al.
Published: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
by: Fan, Zhaoyu, et al.
Published: (2025)
by: Fan, Zhaoyu, et al.
Published: (2025)
OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis
by: Lin, Tianwei, et al.
Published: (2026)
by: Lin, Tianwei, et al.
Published: (2026)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
by: Dai, Yang, et al.
Published: (2026)
by: Dai, Yang, et al.
Published: (2026)
Rendering Multi-Human and Multi-Object with 3D Gaussian Splatting
by: Wang, Weiquan, et al.
Published: (2026)
by: Wang, Weiquan, et al.
Published: (2026)
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
by: Chen, Siqi, et al.
Published: (2025)
by: Chen, Siqi, et al.
Published: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
IDEAL: Leveraging Infinite and Dynamic Characterizations of Large Language Models for Query-focused Summarization
by: Cao, Jie, et al.
Published: (2024)
by: Cao, Jie, et al.
Published: (2024)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
by: Xiao, Wenyi, et al.
Published: (2024)
by: Xiao, Wenyi, et al.
Published: (2024)
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
by: Yuan, Yuqian, et al.
Published: (2025)
by: Yuan, Yuqian, et al.
Published: (2025)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
by: Gao, Minghe, et al.
Published: (2023)
by: Gao, Minghe, et al.
Published: (2023)
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text
by: Yin, Aoxiong, et al.
Published: (2024)
by: Yin, Aoxiong, et al.
Published: (2024)
Similar Items
-
IAD-Unify: A Region-Grounded Unified Model for Industrial Anomaly Segmentation, Understanding, and Generation
by: Zheng, Haoyu, et al.
Published: (2026) -
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
by: Lin, Tianwei, et al.
Published: (2025) -
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
by: Yuan, Yuqian, et al.
Published: (2026) -
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
by: Lin, Tianwei, et al.
Published: (2024) -
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
by: Wang, Wei, et al.
Published: (2026)