MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Haozhe, Cai, Zefan, Si, Shuzheng, Chen, Liang, Gu, Jiuxiang, Xiao, Wen, Zhang, Minjia, Hu, Junjie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
di: Cai, Zefan, et al.
Pubblicazione: (2025)
di: Cai, Zefan, et al.
Pubblicazione: (2025)
MMGR: Multi-Modal Generative Reasoning
di: Cai, Zefan, et al.
Pubblicazione: (2025)
di: Cai, Zefan, et al.
Pubblicazione: (2025)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
di: Zhao, Haozhe, et al.
Pubblicazione: (2023)
di: Zhao, Haozhe, et al.
Pubblicazione: (2023)
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
di: Zhao, Haozhe, et al.
Pubblicazione: (2026)
di: Zhao, Haozhe, et al.
Pubblicazione: (2026)
Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
BabyVision: Visual Reasoning Beyond Language
di: Chen, Liang, et al.
Pubblicazione: (2026)
di: Chen, Liang, et al.
Pubblicazione: (2026)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
di: Zhou, Shijie, et al.
Pubblicazione: (2025)
SANTA: Separate Strategies for Inaccurate and Incomplete Annotation Noise in Distantly-Supervised Named Entity Recognition
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
di: Wang, Zihang, et al.
Pubblicazione: (2026)
di: Wang, Zihang, et al.
Pubblicazione: (2026)
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
di: An, Kaikai, et al.
Pubblicazione: (2024)
di: An, Kaikai, et al.
Pubblicazione: (2024)
UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
di: Tang, Hongxuan, et al.
Pubblicazione: (2025)
di: Tang, Hongxuan, et al.
Pubblicazione: (2025)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
di: Chen, Liang, et al.
Pubblicazione: (2025)
di: Chen, Liang, et al.
Pubblicazione: (2025)
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
DiSA: Diffusion Step Annealing in Autoregressive Image Generation
di: Zhao, Qinyu, et al.
Pubblicazione: (2025)
di: Zhao, Qinyu, et al.
Pubblicazione: (2025)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
di: Li, Bingxuan, et al.
Pubblicazione: (2025)
di: Li, Bingxuan, et al.
Pubblicazione: (2025)
Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning
di: Koleilat, Taha, et al.
Pubblicazione: (2026)
di: Koleilat, Taha, et al.
Pubblicazione: (2026)
Autoregressive Models in Vision: A Survey
di: Xiong, Jing, et al.
Pubblicazione: (2024)
di: Xiong, Jing, et al.
Pubblicazione: (2024)
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
di: Liu, Wenjie, et al.
Pubblicazione: (2026)
di: Liu, Wenjie, et al.
Pubblicazione: (2026)
From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models
di: Wang, Qidong, et al.
Pubblicazione: (2026)
di: Wang, Qidong, et al.
Pubblicazione: (2026)
Delta Attention Residuals
di: Luo, Cheng, et al.
Pubblicazione: (2026)
di: Luo, Cheng, et al.
Pubblicazione: (2026)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
di: Huang, Jinsheng, et al.
Pubblicazione: (2024)
di: Huang, Jinsheng, et al.
Pubblicazione: (2024)
MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension
di: Liu, Ting, et al.
Pubblicazione: (2024)
di: Liu, Ting, et al.
Pubblicazione: (2024)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
di: Wei, Xinyu, et al.
Pubblicazione: (2025)
di: Wei, Xinyu, et al.
Pubblicazione: (2025)
FaithLens: Detecting and Explaining Faithfulness Hallucination
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Pseudo-Prompt Generating in Pre-trained Vision-Language Models for Multi-Label Medical Image Classification
di: Ye, Yaoqin, et al.
Pubblicazione: (2024)
di: Ye, Yaoqin, et al.
Pubblicazione: (2024)
Towards Efficient Vision-Language Tuning: More Information Density, More Generalizability
di: Hao, Tianxiang, et al.
Pubblicazione: (2023)
di: Hao, Tianxiang, et al.
Pubblicazione: (2023)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
di: Cai, Zefan, et al.
Pubblicazione: (2025)
di: Cai, Zefan, et al.
Pubblicazione: (2025)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
di: Chern, Ethan, et al.
Pubblicazione: (2024)
di: Chern, Ethan, et al.
Pubblicazione: (2024)
ImageFolder: Autoregressive Image Generation with Folded Tokens
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
di: Cai, Zefan, et al.
Pubblicazione: (2025) -
MMGR: Multi-Modal Generative Reasoning
di: Cai, Zefan, et al.
Pubblicazione: (2025) -
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
di: Zhao, Haozhe, et al.
Pubblicazione: (2023) -
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
di: Chen, Liang, et al.
Pubblicazione: (2024) -
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)