jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Fuente:
arXiv
Saved in:
| Main Authors: | Koukounas, Andreas, Mastrapas, Georgios, Eslami, Sedigheh, Wang, Bo, Akram, Mohammad Kalim, Günther, Michael, Mohr, Isabelle, Sturua, Saba, Wang, Nan, Xiao, Han |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
by: Sturua, Saba, et al.
Published: (2024)
by: Sturua, Saba, et al.
Published: (2024)
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
by: Günther, Michael, et al.
Published: (2025)
by: Günther, Michael, et al.
Published: (2025)
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
by: Jha, Rohan, et al.
Published: (2024)
by: Jha, Rohan, et al.
Published: (2024)
jina-vlm: Small Multilingual Vision Language Model
by: Koukounas, Andreas, et al.
Published: (2025)
by: Koukounas, Andreas, et al.
Published: (2025)
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
by: Günther, Michael, et al.
Published: (2023)
by: Günther, Michael, et al.
Published: (2023)
Jina CLIP: Your CLIP Model Is Also Your Text Retriever
by: Koukounas, Andreas, et al.
Published: (2024)
by: Koukounas, Andreas, et al.
Published: (2024)
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
by: Mohr, Isabelle, et al.
Published: (2024)
by: Mohr, Isabelle, et al.
Published: (2024)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
by: Asanuma, Haruka, et al.
Published: (2025)
by: Asanuma, Haruka, et al.
Published: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
by: Dai, Song, et al.
Published: (2025)
by: Dai, Song, et al.
Published: (2025)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
by: Li, Huibin, et al.
Published: (2025)
by: Li, Huibin, et al.
Published: (2025)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
by: Boumber, Dainis, et al.
Published: (2024)
by: Boumber, Dainis, et al.
Published: (2024)
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
by: Bian, Zhipeng, et al.
Published: (2026)
by: Bian, Zhipeng, et al.
Published: (2026)
PC-SNN: Predictive Coding-based Local Hebbian Plasticity Learning in Spiking Neural Networks
by: Wang, Haidong, et al.
Published: (2022)
by: Wang, Haidong, et al.
Published: (2022)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Learning the meanings of function words from grounded language using a visual question answering model
by: Portelance, Eva, et al.
Published: (2023)
by: Portelance, Eva, et al.
Published: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
by: Yang, Shan
Published: (2026)
by: Yang, Shan
Published: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
by: Bucher, Martin JJ., et al.
Published: (2025)
by: Bucher, Martin JJ., et al.
Published: (2025)
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
by: Tang, Guowei
Published: (2026)
by: Tang, Guowei
Published: (2026)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
by: Le, Van-Truong
Published: (2026)
by: Le, Van-Truong
Published: (2026)
Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers
by: Zhao, Yiming
Published: (2026)
by: Zhao, Yiming
Published: (2026)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
Text-to-Events: Synthetic Event Camera Streams from Conditional Text Input
by: Ott, Joachim, et al.
Published: (2024)
by: Ott, Joachim, et al.
Published: (2024)
LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
by: Wang, Fusang, et al.
Published: (2026)
by: Wang, Fusang, et al.
Published: (2026)
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
by: Dong, Mingkang, et al.
Published: (2026)
by: Dong, Mingkang, et al.
Published: (2026)
SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
by: Liu, Sicheng, et al.
Published: (2024)
by: Liu, Sicheng, et al.
Published: (2024)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
by: Cao, Jingtao, et al.
Published: (2024)
by: Cao, Jingtao, et al.
Published: (2024)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
vS-Graphs: Tightly Coupling Visual SLAM and 3D Scene Graphs Exploiting Hierarchical Scene Understanding
by: Tourani, Ali, et al.
Published: (2025)
by: Tourani, Ali, et al.
Published: (2025)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
by: Qian, Wenxu, et al.
Published: (2025)
by: Qian, Wenxu, et al.
Published: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
by: Adžemović, Momir
Published: (2025)
by: Adžemović, Momir
Published: (2025)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
by: Han, Yudong, et al.
Published: (2026)
by: Han, Yudong, et al.
Published: (2026)
Semantic Leakage from Image Embeddings
by: Chen, Yiyi, et al.
Published: (2026)
by: Chen, Yiyi, et al.
Published: (2026)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
by: Ferenczi, Bryce, et al.
Published: (2023)
by: Ferenczi, Bryce, et al.
Published: (2023)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
by: Perera, Amal S., et al.
Published: (2025)
by: Perera, Amal S., et al.
Published: (2025)
From Latent to Engine Manifolds: Analyzing ImageBind's Multimodal Embedding Space
by: Hamara, Andrew, et al.
Published: (2024)
by: Hamara, Andrew, et al.
Published: (2024)
Deep Learning for automated multi-scale functional field boundaries extraction using multi-date Sentinel-2 and PlanetScope imagery: Case Study of Netherlands and Pakistan
by: Zahid, Saba, et al.
Published: (2024)
by: Zahid, Saba, et al.
Published: (2024)
Similar Items
-
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
by: Sturua, Saba, et al.
Published: (2024) -
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
by: Günther, Michael, et al.
Published: (2025) -
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
by: Jha, Rohan, et al.
Published: (2024) -
jina-vlm: Small Multilingual Vision Language Model
by: Koukounas, Andreas, et al.
Published: (2025) -
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
by: Günther, Michael, et al.
Published: (2023)