Think Then Embed: Generative Context Improves Multimodal Embedding
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Xuanming, Cheng, Jianpeng, Chen, Hong-you, Shukla, Satya Narayan, Awasthi, Abhijeet, Pan, Xichen, Ahuja, Chaitanya, Mishra, Shlok Kumar, Yang, Yonghuan, Xiao, Jun, Guo, Qi, Lim, Ser-Nam, Singh, Aashu, Fan, Xiangjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens
by: Cheng, Jianpeng, et al.
Published: (2026)
by: Cheng, Jianpeng, et al.
Published: (2026)
Reason to Contrast: A Cascaded Multimodal Retrieval Framework
by: Cui, Xuanming, et al.
Published: (2025)
by: Cui, Xuanming, et al.
Published: (2025)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
Xray-Visual Models: Scaling Vision models on Industry Scale Data
by: Mishra, Shlok, et al.
Published: (2026)
by: Mishra, Shlok, et al.
Published: (2026)
Transfer between Modalities with MetaQueries
by: Pan, Xichen, et al.
Published: (2025)
by: Pan, Xichen, et al.
Published: (2025)
AirSketch: Generative Motion to Sketch
by: Lim, Hui Xian Grace, et al.
Published: (2024)
by: Lim, Hui Xian Grace, et al.
Published: (2024)
Socratic Students: Teaching Language Models to Learn by Asking Questions
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
by: Cui, Xuanming, et al.
Published: (2025)
by: Cui, Xuanming, et al.
Published: (2025)
CompCap: Improving Multimodal Large Language Models with Composite Captions
by: Chen, Xiaohui, et al.
Published: (2024)
by: Chen, Xiaohui, et al.
Published: (2024)
A Simple and Effective Reinforcement Learning Method for Text-to-Image Diffusion Fine-tuning
by: Gupta, Shashank, et al.
Published: (2025)
by: Gupta, Shashank, et al.
Published: (2025)
Heteroscedastic Temporal Variational Autoencoder For Irregular Time Series
by: Shukla, Satya Narayan, et al.
Published: (2021)
by: Shukla, Satya Narayan, et al.
Published: (2021)
Towards Chunk-Wise Generation for Long Videos
by: Zhang, Siyang, et al.
Published: (2024)
by: Zhang, Siyang, et al.
Published: (2024)
Cultural and Historical Identity in Amitav Ghosh’s River of Smoke: A Postcolonial Perspective
by: Satya Narayan
Published: (2021)
by: Satya Narayan
Published: (2021)
Depicting Culture and Identity in Amitav Ghosh’s The Shadow Lines
by: Satya Narayan
Published: (2017)
by: Satya Narayan
Published: (2017)
Cultural and Historical Identity in Amitav Ghosh’s River of Smoke: A Postcolonial Perspective
by: Satya Narayan
Published: (2021)
by: Satya Narayan
Published: (2021)
Pooling Attention: Evaluating Pretrained Transformer Embeddings for Deception Classification
by: Mamtani, Sumit, et al.
Published: (2025)
by: Mamtani, Sumit, et al.
Published: (2025)
An Attribute-Based Measure of Video Complexity
by: Sarkar, Aditya, et al.
Published: (2026)
by: Sarkar, Aditya, et al.
Published: (2026)
TractoEmbed: Modular Multi-level Embedding framework for white matter tract segmentation
by: Goel, Anoushkrit, et al.
Published: (2024)
by: Goel, Anoushkrit, et al.
Published: (2024)
VideoMerge: Towards Training-free Long Video Generation
by: Zhang, Siyang, et al.
Published: (2025)
by: Zhang, Siyang, et al.
Published: (2025)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Is Cosine-Similarity of Embeddings Really About Similarity?
by: Steck, Harald, et al.
Published: (2024)
by: Steck, Harald, et al.
Published: (2024)
Realised Volatility Forecasting: Machine Learning via Financial Word Embedding
by: Rahimikia, Eghbal, et al.
Published: (2021)
by: Rahimikia, Eghbal, et al.
Published: (2021)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
by: Chen, Harold Haodong, et al.
Published: (2024)
by: Chen, Harold Haodong, et al.
Published: (2024)
Robust Learning of Diverse Code Edits
by: Aggarwal, Tushar, et al.
Published: (2025)
by: Aggarwal, Tushar, et al.
Published: (2025)
Effective Embedding of Integer Linear Inequalities for Variational Quantum Algorithms
by: Hess, Maximilian, et al.
Published: (2024)
by: Hess, Maximilian, et al.
Published: (2024)
Self-Adaptive Reconstruction with Contrastive Learning for Unsupervised Sentence Embeddings
by: Liu, Junlong, et al.
Published: (2024)
by: Liu, Junlong, et al.
Published: (2024)
On Global Embedding of Assisted Fibre Inflation
by: Leontaris, George K., et al.
Published: (2026)
by: Leontaris, George K., et al.
Published: (2026)
DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
by: Long, Tao, et al.
Published: (2025)
by: Long, Tao, et al.
Published: (2025)
Embedding Empirical Distributions for Computing Optimal Transport Maps
by: Jiang, Mingchen, et al.
Published: (2025)
by: Jiang, Mingchen, et al.
Published: (2025)
TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval
by: Chen, Jialin, et al.
Published: (2025)
by: Chen, Jialin, et al.
Published: (2025)
LASER: A Neuro-Symbolic Framework for Learning Spatial-Temporal Scene Graphs with Weak Supervision
by: Huang, Jiani, et al.
Published: (2023)
by: Huang, Jiani, et al.
Published: (2023)
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
by: Tu, Yuanpeng, et al.
Published: (2025)
by: Tu, Yuanpeng, et al.
Published: (2025)
Towards Unified 3D Object Detection via Algorithm and Data Unification
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Composing Object Relations and Attributes for Image-Text Matching
by: Pham, Khoi, et al.
Published: (2024)
by: Pham, Khoi, et al.
Published: (2024)
Fast Encoding and Decoding for Implicit Video Representation
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
by: Qian, Zhaofang, et al.
Published: (2024)
by: Qian, Zhaofang, et al.
Published: (2024)
Delta Activations: A Representation for Finetuned Large Language Models
by: Xu, Zhiqiu, et al.
Published: (2025)
by: Xu, Zhiqiu, et al.
Published: (2025)
PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
by: Bang, Hayeon, et al.
Published: (2025)
by: Bang, Hayeon, et al.
Published: (2025)
Similar Items
-
TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens
by: Cheng, Jianpeng, et al.
Published: (2026) -
Reason to Contrast: A Cascaded Multimodal Retrieval Framework
by: Cui, Xuanming, et al.
Published: (2025) -
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026) -
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025) -
Xray-Visual Models: Scaling Vision models on Industry Scale Data
by: Mishra, Shlok, et al.
Published: (2026)