Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Salman, Shaeke, Shams, Md Montasir Bin, Liu, Xiuwen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
di: Salman, Shaeke, et al.
Pubblicazione: (2024)
Malicious Path Manipulations via Exploitation of Representation Vulnerabilities of Vision-Language Navigation Systems
di: Islam, Chashi Mahiul, et al.
Pubblicazione: (2024)
di: Islam, Chashi Mahiul, et al.
Pubblicazione: (2024)
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
di: Shams, Montasir, et al.
Pubblicazione: (2025)
di: Shams, Montasir, et al.
Pubblicazione: (2025)
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
di: Zhan, Jun, et al.
Pubblicazione: (2024)
di: Zhan, Jun, et al.
Pubblicazione: (2024)
Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
di: Tang, Yiwen, et al.
Pubblicazione: (2024)
di: Tang, Yiwen, et al.
Pubblicazione: (2024)
Towards Understanding Camera Motions in Any Video
di: Lin, Zhiqiu, et al.
Pubblicazione: (2025)
di: Lin, Zhiqiu, et al.
Pubblicazione: (2025)
Self-Evaluation Unlocks Any-Step Text-to-Image Generation
di: Yu, Xin, et al.
Pubblicazione: (2025)
di: Yu, Xin, et al.
Pubblicazione: (2025)
OpenClaw-RL: Train Any Agent Simply by Talking
di: Wang, Yinjie, et al.
Pubblicazione: (2026)
di: Wang, Yinjie, et al.
Pubblicazione: (2026)
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
di: Wang, Xinghao, et al.
Pubblicazione: (2024)
di: Wang, Xinghao, et al.
Pubblicazione: (2024)
AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any Scenario
di: Li, Yuhan, et al.
Pubblicazione: (2024)
di: Li, Yuhan, et al.
Pubblicazione: (2024)
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
di: Bachmann, Roman, et al.
Pubblicazione: (2024)
di: Bachmann, Roman, et al.
Pubblicazione: (2024)
AnyTrans: Translate AnyText in the Image with Large Scale Models
di: Qian, Zhipeng, et al.
Pubblicazione: (2024)
di: Qian, Zhipeng, et al.
Pubblicazione: (2024)
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
di: Ismithdeen, Mohamed Insaf, et al.
Pubblicazione: (2025)
di: Ismithdeen, Mohamed Insaf, et al.
Pubblicazione: (2025)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
di: Li, Bingyu, et al.
Pubblicazione: (2024)
di: Li, Bingyu, et al.
Pubblicazione: (2024)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image Generation
di: Chadebec, Clément, et al.
Pubblicazione: (2024)
di: Chadebec, Clément, et al.
Pubblicazione: (2024)
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
di: Horawalavithana, Sameera, et al.
Pubblicazione: (2023)
di: Horawalavithana, Sameera, et al.
Pubblicazione: (2023)
Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs
di: Luo, Jinqi, et al.
Pubblicazione: (2026)
di: Luo, Jinqi, et al.
Pubblicazione: (2026)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
di: Zhu, Kangyu, et al.
Pubblicazione: (2024)
di: Zhu, Kangyu, et al.
Pubblicazione: (2024)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
Towards Visual Text Grounding of Multimodal Large Language Model
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
A Multimodal, Multitask System for Generating E Commerce Text Listings from Images
di: Singh, Nayan Kumar
Pubblicazione: (2025)
di: Singh, Nayan Kumar
Pubblicazione: (2025)
DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any Architecture
di: Xiang, Qianlong, et al.
Pubblicazione: (2024)
di: Xiang, Qianlong, et al.
Pubblicazione: (2024)
Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation
di: Molino, Daniele, et al.
Pubblicazione: (2025)
di: Molino, Daniele, et al.
Pubblicazione: (2025)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
di: Wang, Haochen, et al.
Pubblicazione: (2025)
di: Wang, Haochen, et al.
Pubblicazione: (2025)
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
di: Wu, Shengguang, et al.
Pubblicazione: (2025)
di: Wu, Shengguang, et al.
Pubblicazione: (2025)
Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
di: Yang, Kai, et al.
Pubblicazione: (2023)
di: Yang, Kai, et al.
Pubblicazione: (2023)
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant
di: Penamakuri, Abhirama Subramanyam, et al.
Pubblicazione: (2024)
di: Penamakuri, Abhirama Subramanyam, et al.
Pubblicazione: (2024)
TAO-Amodal: A Benchmark for Tracking Any Object Amodally
di: Hsieh, Cheng-Yen, et al.
Pubblicazione: (2023)
di: Hsieh, Cheng-Yen, et al.
Pubblicazione: (2023)
Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
di: Rojas, Kevin, et al.
Pubblicazione: (2025)
di: Rojas, Kevin, et al.
Pubblicazione: (2025)
AnySR: Realizing Image Super-Resolution as Any-Scale, Any-Resource
di: Zhan, Wengyi, et al.
Pubblicazione: (2024)
di: Zhan, Wengyi, et al.
Pubblicazione: (2024)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
di: Lin, Han, et al.
Pubblicazione: (2024)
di: Lin, Han, et al.
Pubblicazione: (2024)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
di: Yuan, Huizhuo, et al.
Pubblicazione: (2024)
di: Yuan, Huizhuo, et al.
Pubblicazione: (2024)
A Generalist Model for Diverse Text-Guided Medical Image Synthesis
di: Cho, Joseph, et al.
Pubblicazione: (2024)
di: Cho, Joseph, et al.
Pubblicazione: (2024)
CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models
di: Xue, Yuyang, et al.
Pubblicazione: (2025)
di: Xue, Yuyang, et al.
Pubblicazione: (2025)
Match & Choose: Model Selection Framework for Fine-tuning Text-to-Image Diffusion Models
di: Lewandowski, Basile, et al.
Pubblicazione: (2025)
di: Lewandowski, Basile, et al.
Pubblicazione: (2025)
CapsFusion: Rethinking Image-Text Data at Scale
di: Yu, Qiying, et al.
Pubblicazione: (2023)
di: Yu, Qiying, et al.
Pubblicazione: (2023)
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
di: Salman, Shaeke, et al.
Pubblicazione: (2024) -
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
di: Salman, Shaeke, et al.
Pubblicazione: (2024) -
Malicious Path Manipulations via Exploitation of Representation Vulnerabilities of Vision-Language Navigation Systems
di: Islam, Chashi Mahiul, et al.
Pubblicazione: (2024) -
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
di: Shams, Montasir, et al.
Pubblicazione: (2025) -
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
di: Zhan, Jun, et al.
Pubblicazione: (2024)