MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Wanggui, Fu, Siming, Liu, Mushui, Wang, Xierui, Xiao, Wenyi, Shu, Fangxun, Wang, Yi, Zhang, Lei, Yu, Zhelun, Li, Haoyuan, Huang, Ziwei, Gan, LeiLei, Jiang, Hao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
di: Huang, Ziwei, et al.
Pubblicazione: (2024)
di: Huang, Ziwei, et al.
Pubblicazione: (2024)
MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
di: Wang, Xierui, et al.
Pubblicazione: (2024)
di: Wang, Xierui, et al.
Pubblicazione: (2024)
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
di: Zhang, Guanghao, et al.
Pubblicazione: (2025)
di: Zhang, Guanghao, et al.
Pubblicazione: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
di: Di, Shangzhe, et al.
Pubblicazione: (2025)
di: Di, Shangzhe, et al.
Pubblicazione: (2025)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
Backdoor Attacks on Dense Retrieval via Public and Unintentional Triggers
di: Long, Quanyu, et al.
Pubblicazione: (2024)
di: Long, Quanyu, et al.
Pubblicazione: (2024)
MARS: Enabling Autoregressive Models Multi-Token Generation
di: Jin, Ziqi, et al.
Pubblicazione: (2026)
di: Jin, Ziqi, et al.
Pubblicazione: (2026)
Filter & Align: Leveraging Human Knowledge to Curate Image-Text Data
di: Zhang, Lei, et al.
Pubblicazione: (2023)
di: Zhang, Lei, et al.
Pubblicazione: (2023)
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
di: Liu, Jinlong, et al.
Pubblicazione: (2026)
di: Liu, Jinlong, et al.
Pubblicazione: (2026)
The Fibrillatory Wave Amplitude of ECG Decreases Over Time in Patients With Persistent AF: A Retrospective Cohort Study
di: Jing Li, et al.
Pubblicazione: (2026)
di: Jing Li, et al.
Pubblicazione: (2026)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
di: Zhang, Wenqiao, et al.
Pubblicazione: (2024)
di: Zhang, Wenqiao, et al.
Pubblicazione: (2024)
Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning
di: Ma, LeiLei, et al.
Pubblicazione: (2025)
di: Ma, LeiLei, et al.
Pubblicazione: (2025)
MARS: Mesh AutoRegressive Model for 3D Shape Detailization
di: Gao, Jingnan, et al.
Pubblicazione: (2025)
di: Gao, Jingnan, et al.
Pubblicazione: (2025)
Boosting Private Domain Understanding of Efficient MLLMs: A Tuning-free, Adaptive, Universal Prompt Optimization Framework
di: Liu, Jiang, et al.
Pubblicazione: (2024)
di: Liu, Jiang, et al.
Pubblicazione: (2024)
Analysis of changes and correlation in condyle‐fossa relationship after maxillary skeletal expansion
di: Yezi Qi, et al.
Pubblicazione: (2024)
di: Yezi Qi, et al.
Pubblicazione: (2024)
Mixture-of-Experts with Gradient Conflict-Driven Subspace Topology Pruning for Emergent Modularity
di: Gan, Yuxing, et al.
Pubblicazione: (2025)
di: Gan, Yuxing, et al.
Pubblicazione: (2025)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
di: She, D., et al.
Pubblicazione: (2025)
di: She, D., et al.
Pubblicazione: (2025)
Fine-grained Text to Image Synthesis
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
Task Adaptive Feature Distribution Based Network for Few-shot Fine-grained Target Classification
di: Li, Ping, et al.
Pubblicazione: (2024)
di: Li, Ping, et al.
Pubblicazione: (2024)
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
di: She, Dong, et al.
Pubblicazione: (2025)
di: She, Dong, et al.
Pubblicazione: (2025)
Bridge to Non-Barrier Communication: Gloss-Prompted Fine-grained Cued Speech Gesture Generation with Diffusion Model
di: Lei, Wentao, et al.
Pubblicazione: (2024)
di: Lei, Wentao, et al.
Pubblicazione: (2024)
EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models
di: Wang, Kun, et al.
Pubblicazione: (2025)
di: Wang, Kun, et al.
Pubblicazione: (2025)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
di: Huang, Wei, et al.
Pubblicazione: (2026)
di: Huang, Wei, et al.
Pubblicazione: (2026)
A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
di: Xiao, Wenyi, et al.
Pubblicazione: (2024)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
di: Yao, Jixun, et al.
Pubblicazione: (2025)
di: Yao, Jixun, et al.
Pubblicazione: (2025)
Multimodal Fine-grained Reasoning for Post Quality Evaluation
di: Guo, Xiaoxu, et al.
Pubblicazione: (2025)
di: Guo, Xiaoxu, et al.
Pubblicazione: (2025)
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
di: Sun, Haiyang, et al.
Pubblicazione: (2025)
di: Sun, Haiyang, et al.
Pubblicazione: (2025)
Fully Fine-tuned CLIP Models are Efficient Few-Shot Learners
di: Liu, Mushui, et al.
Pubblicazione: (2024)
di: Liu, Mushui, et al.
Pubblicazione: (2024)
STAR: Scale-wise Text-conditioned AutoRegressive image generation
di: Ma, Xiaoxiao, et al.
Pubblicazione: (2024)
di: Ma, Xiaoxiao, et al.
Pubblicazione: (2024)
Auto-Regressive Surface Cutting
di: Li, Yang, et al.
Pubblicazione: (2025)
di: Li, Yang, et al.
Pubblicazione: (2025)
MARS: Unleashing the Power of Speculative Decoding via Margin-Aware Verification
di: Song, Jingwei, et al.
Pubblicazione: (2026)
di: Song, Jingwei, et al.
Pubblicazione: (2026)
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
di: Lin, Tianwei, et al.
Pubblicazione: (2024)
di: Lin, Tianwei, et al.
Pubblicazione: (2024)
Reinforcement Fine-Tuning for Materials Design
di: Cao, Zhendong, et al.
Pubblicazione: (2025)
di: Cao, Zhendong, et al.
Pubblicazione: (2025)
Causal Interpretation of Regressions With Ranks
di: Lei, Lihua
Pubblicazione: (2024)
di: Lei, Lihua
Pubblicazione: (2024)
AU-Blendshape for Fine-grained Stylized 3D Facial Expression Manipulation
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
Beyond Model Base Retrieval: Weaving Knowledge to Master Fine-grained Neural Network Design
di: Wang, Jialiang, et al.
Pubblicazione: (2025)
di: Wang, Jialiang, et al.
Pubblicazione: (2025)
Fine-grained Testing for Autonomous Driving Software: a Study on Autoware with LLM-driven Unit Testing
di: Wang, Wenhan, et al.
Pubblicazione: (2025)
di: Wang, Wenhan, et al.
Pubblicazione: (2025)
Finite‐Time H∞ Comprehensive Bumpless Transfer Control for Switched Systems
di: Haoyuan Zhang, et al.
Pubblicazione: (2026)
di: Haoyuan Zhang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
di: Xiao, Wenyi, et al.
Pubblicazione: (2024) -
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
di: Huang, Ziwei, et al.
Pubblicazione: (2024) -
MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
di: Wang, Xierui, et al.
Pubblicazione: (2024) -
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
di: Zhang, Guanghao, et al.
Pubblicazione: (2025) -
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
di: Di, Shangzhe, et al.
Pubblicazione: (2025)