Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chao, Li, Tianhong, Nuthalapati, Sai Vidyaranya, Chen, Hong-You, Shukla, Satya Narayan, Cheng, Jianpeng, Yang, Yonghuan, Xiao, Jun, Fan, Xiangjun, Singh, Aashu, Katabi, Dina, Mishra, Shlok Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Return of Unconditional Generation: A Self-supervised Representation Generation Method
by: Li, Tianhong, et al.
Published: (2023)
by: Li, Tianhong, et al.
Published: (2023)
An Attribute-Based Measure of Video Complexity
by: Sarkar, Aditya, et al.
Published: (2026)
by: Sarkar, Aditya, et al.
Published: (2026)
Think Then Embed: Generative Context Improves Multimodal Embedding
by: Cui, Xuanming, et al.
Published: (2025)
by: Cui, Xuanming, et al.
Published: (2025)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
GEM: Empowering LLM for both Embedding Generation and Language Understanding
by: Zhang, Caojin, et al.
Published: (2025)
by: Zhang, Caojin, et al.
Published: (2025)
Reparo: Loss-Resilient Generative Codec for Video Conferencing
by: Li, Tianhong, et al.
Published: (2023)
by: Li, Tianhong, et al.
Published: (2023)
Xray-Visual Models: Scaling Vision models on Industry Scale Data
by: Mishra, Shlok, et al.
Published: (2026)
by: Mishra, Shlok, et al.
Published: (2026)
Reason to Contrast: A Cascaded Multimodal Retrieval Framework
by: Cui, Xuanming, et al.
Published: (2025)
by: Cui, Xuanming, et al.
Published: (2025)
Transfer between Modalities with MetaQueries
by: Pan, Xichen, et al.
Published: (2025)
by: Pan, Xichen, et al.
Published: (2025)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
Socratic Students: Teaching Language Models to Learn by Asking Questions
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
Heteroscedastic Temporal Variational Autoencoder For Irregular Time Series
by: Shukla, Satya Narayan, et al.
Published: (2021)
by: Shukla, Satya Narayan, et al.
Published: (2021)
Challenges in Building a Global Supply Chain in the Apparel Industry.
by: Vidyaranya B. Gargeya
Published: (2001)
by: Vidyaranya B. Gargeya
Published: (2001)
Physiology as Language: Translating Respiration to Sleep EEG
by: Zha, Kaiwen, et al.
Published: (2026)
by: Zha, Kaiwen, et al.
Published: (2026)
Global Content Localization Strategies: Translation Management Integration for Marketing Technology Ecosystems
by: Sudhakar Nuthalapati
Published: (2026)
by: Sudhakar Nuthalapati
Published: (2026)
Cultural and Historical Identity in Amitav Ghosh’s River of Smoke: A Postcolonial Perspective
by: Satya Narayan
Published: (2021)
by: Satya Narayan
Published: (2021)
Depicting Culture and Identity in Amitav Ghosh’s The Shadow Lines
by: Satya Narayan
Published: (2017)
by: Satya Narayan
Published: (2017)
Cultural and Historical Identity in Amitav Ghosh’s River of Smoke: A Postcolonial Perspective
by: Satya Narayan
Published: (2021)
by: Satya Narayan
Published: (2021)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
by: Zha, Kaiwen, et al.
Published: (2025)
by: Zha, Kaiwen, et al.
Published: (2025)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
Generative Multi-Objective Bayesian Optimization with Scalable Batch Evaluations for Sample-Efficient De Novo Molecular Design
by: Muthyala, Madhav R., et al.
Published: (2025)
by: Muthyala, Madhav R., et al.
Published: (2025)
Back to Basics: Let Denoising Generative Models Denoise
by: Li, Tianhong, et al.
Published: (2025)
by: Li, Tianhong, et al.
Published: (2025)
Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular Diversity
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
Language-Guided Image Tokenization for Generation
by: Zha, Kaiwen, et al.
Published: (2024)
by: Zha, Kaiwen, et al.
Published: (2024)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
Nutritional Impacts of Direct Selling to Supermarkets: The Case of Farm Households in India
by: Chandra S. Nuthalapati, et al.
Published: (2026)
by: Chandra S. Nuthalapati, et al.
Published: (2026)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
by: Han, Jiaming, et al.
Published: (2025)
by: Han, Jiaming, et al.
Published: (2025)
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
by: Zhao, Yue, et al.
Published: (2025)
by: Zhao, Yue, et al.
Published: (2025)
LumiX: Structured and Coherent Text-to-Intrinsic Generation
by: Han, Xu, et al.
Published: (2025)
by: Han, Xu, et al.
Published: (2025)
Impact of Bipolar Disorder Medications on Fracture Risk: Insights from a Retrospective Cohort Study
by: Mahita C. Nuthalapati, et al.
Published: (2025)
by: Mahita C. Nuthalapati, et al.
Published: (2025)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
Thought2Text: Text Generation from EEG Signal using Large Language Models (LLMs)
by: Mishra, Abhijit, et al.
Published: (2024)
by: Mishra, Abhijit, et al.
Published: (2024)
CompCap: Improving Multimodal Large Language Models with Composite Captions
by: Chen, Xiaohui, et al.
Published: (2024)
by: Chen, Xiaohui, et al.
Published: (2024)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
by: Wang, Peiyu, et al.
Published: (2025)
by: Wang, Peiyu, et al.
Published: (2025)
Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Compositional Visual Planning via Inference-Time Diffusion Scaling
by: Zhang, Yixin, et al.
Published: (2026)
by: Zhang, Yixin, et al.
Published: (2026)
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
by: Wu, Yecheng, et al.
Published: (2024)
by: Wu, Yecheng, et al.
Published: (2024)
Similar Items
-
Return of Unconditional Generation: A Self-supervised Representation Generation Method
by: Li, Tianhong, et al.
Published: (2023) -
An Attribute-Based Measure of Video Complexity
by: Sarkar, Aditya, et al.
Published: (2026) -
Think Then Embed: Generative Context Improves Multimodal Embedding
by: Cui, Xuanming, et al.
Published: (2025) -
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025) -
GEM: Empowering LLM for both Embedding Generation and Language Understanding
by: Zhang, Caojin, et al.
Published: (2025)