Towards Better Text-to-Image Generation Alignment via Attention Modulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Yihang, Cao, Xiao, Li, Kaixin, Chen, Zitan, Wang, Haonan, Meng, Lei, Huang, Zhiyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
von: Ma, Zhiyong, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyong, et al.
Veröffentlicht: (2025)
Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
von: Wu, Zichen, et al.
Veröffentlicht: (2025)
von: Wu, Zichen, et al.
Veröffentlicht: (2025)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
von: Lu, Jinghui, et al.
Veröffentlicht: (2024)
von: Lu, Jinghui, et al.
Veröffentlicht: (2024)
On Copyright Risks of Text-to-Image Diffusion Models
von: Zhang, Yang, et al.
Veröffentlicht: (2023)
von: Zhang, Yang, et al.
Veröffentlicht: (2023)
Shapley Value-based Contrastive Alignment for Multimodal Information Extraction
von: Luo, Wen, et al.
Veröffentlicht: (2024)
von: Luo, Wen, et al.
Veröffentlicht: (2024)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
PTA: Enhancing Multimodal Sentiment Analysis through Pipelined Prediction and Translation-based Alignment
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
Knowledge-Guided Dynamic Modality Attention Fusion Framework for Multimodal Sentiment Analysis
von: Feng, Xinyu, et al.
Veröffentlicht: (2024)
von: Feng, Xinyu, et al.
Veröffentlicht: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
"Is This It?": Towards Ecologically Valid Benchmarks for Situated Collaboration
von: Bohus, Dan, et al.
Veröffentlicht: (2024)
von: Bohus, Dan, et al.
Veröffentlicht: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
Towards Robust Multimodal Sentiment Analysis with Incomplete Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Generation for Electrocardiogram-Language Models
von: Song, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Song, Xiaoyu, et al.
Veröffentlicht: (2025)
Audio ControlNet for Fine-Grained Audio Generation and Editing
von: Zhu, Haina, et al.
Veröffentlicht: (2026)
von: Zhu, Haina, et al.
Veröffentlicht: (2026)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
A Survey on Image-text Multimodal Models
von: Guo, Ruifeng, et al.
Veröffentlicht: (2023)
von: Guo, Ruifeng, et al.
Veröffentlicht: (2023)
A Survey of Generative Categories and Techniques in Multimodal Generative Models
von: Han, Longzhen, et al.
Veröffentlicht: (2025)
von: Han, Longzhen, et al.
Veröffentlicht: (2025)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
EEG2TEXT-CN: An Exploratory Study of Open-Vocabulary Chinese Text-EEG Alignment via Large Language Model and Contrastive Learning on ChineseEEG
von: Lu, Jacky Tai-Yu, et al.
Veröffentlicht: (2025)
von: Lu, Jacky Tai-Yu, et al.
Veröffentlicht: (2025)
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
Memory-Centric Embodied Question Answering
von: Zhai, Mingliang, et al.
Veröffentlicht: (2025)
von: Zhai, Mingliang, et al.
Veröffentlicht: (2025)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
von: Han, Wei, et al.
Veröffentlicht: (2023)
von: Han, Wei, et al.
Veröffentlicht: (2023)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond
von: Wei, Tianxin, et al.
Veröffentlicht: (2024)
von: Wei, Tianxin, et al.
Veröffentlicht: (2024)
History-Guided Iterative Visual Reasoning with Self-Correction
von: Yang, Xinglong, et al.
Veröffentlicht: (2026)
von: Yang, Xinglong, et al.
Veröffentlicht: (2026)
New Job, New Gender? Measuring the Social Bias in Image Generation Models
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
von: Ku, Max, et al.
Veröffentlicht: (2023)
von: Ku, Max, et al.
Veröffentlicht: (2023)
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives
von: Chen, Yanxi, et al.
Veröffentlicht: (2025)
von: Chen, Yanxi, et al.
Veröffentlicht: (2025)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal Reasoning
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
von: S, Sridhar, et al.
Veröffentlicht: (2025)
von: S, Sridhar, et al.
Veröffentlicht: (2025)
Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective
von: Zhu, Xiangru, et al.
Veröffentlicht: (2024)
von: Zhu, Xiangru, et al.
Veröffentlicht: (2024)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
von: Ma, Zhiyong, et al.
Veröffentlicht: (2025) -
Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
von: Wu, Zichen, et al.
Veröffentlicht: (2025) -
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
von: Lu, Jinghui, et al.
Veröffentlicht: (2024) -
On Copyright Risks of Text-to-Image Diffusion Models
von: Zhang, Yang, et al.
Veröffentlicht: (2023) -
Shapley Value-based Contrastive Alignment for Multimodal Information Extraction
von: Luo, Wen, et al.
Veröffentlicht: (2024)