Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
Fuente:
arXiv
Saved in:
| Main Authors: | Günther, Michael, Mohr, Isabelle, Williams, Daniel James, Wang, Bo, Xiao, Han |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
by: Günther, Michael, et al.
Published: (2025)
by: Günther, Michael, et al.
Published: (2025)
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
by: Sturua, Saba, et al.
Published: (2024)
by: Sturua, Saba, et al.
Published: (2024)
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
by: Jha, Rohan, et al.
Published: (2024)
by: Jha, Rohan, et al.
Published: (2024)
Efficient Code Embeddings from Code Generation Models
by: Kryvosheieva, Daria, et al.
Published: (2025)
by: Kryvosheieva, Daria, et al.
Published: (2025)
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
by: Mohr, Isabelle, et al.
Published: (2024)
by: Mohr, Isabelle, et al.
Published: (2024)
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
by: Koukounas, Andreas, et al.
Published: (2024)
by: Koukounas, Andreas, et al.
Published: (2024)
jina-reranker-v3: Last but Not Late Interaction for Listwise Document Reranking
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
A Study into Investigating Temporal Robustness of LLMs
by: Wallat, Jonas, et al.
Published: (2025)
by: Wallat, Jonas, et al.
Published: (2025)
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
by: Günther, Michael, et al.
Published: (2023)
by: Günther, Michael, et al.
Published: (2023)
Extracting Sentence Embeddings from Pretrained Transformer Models
by: Stankevičius, Lukas, et al.
Published: (2024)
by: Stankevičius, Lukas, et al.
Published: (2024)
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
Jina CLIP: Your CLIP Model Is Also Your Text Retriever
by: Koukounas, Andreas, et al.
Published: (2024)
by: Koukounas, Andreas, et al.
Published: (2024)
Knowledge Distillation of Domain-adapted LLMs for Question-Answering in Telecom
by: Sen, Rishika, et al.
Published: (2025)
by: Sen, Rishika, et al.
Published: (2025)
Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO
by: Butt, Umer, et al.
Published: (2024)
by: Butt, Umer, et al.
Published: (2024)
Automatic Cardiac Risk Management Classification using large-context Electronic Patients Health Records
by: Vitale, Jacopo, et al.
Published: (2026)
by: Vitale, Jacopo, et al.
Published: (2026)
ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature Summary
by: Li, Yutong, et al.
Published: (2024)
by: Li, Yutong, et al.
Published: (2024)
On Self-improving Token Embeddings
by: Kubek, Mario M., et al.
Published: (2025)
by: Kubek, Mario M., et al.
Published: (2025)
Evaluation of Table Representations to Answer Questions from Tables in Documents : A Case Study using 3GPP Specifications
by: Roychowdhury, Sujoy, et al.
Published: (2024)
by: Roychowdhury, Sujoy, et al.
Published: (2024)
Predicting the Geolocation of Tweets Using transformer models on Customized Data
by: Lutsai, Kateryna, et al.
Published: (2023)
by: Lutsai, Kateryna, et al.
Published: (2023)
Controllable Evidence Selection in Retrieval-Augmented Question Answering via Deterministic Utility Gating
by: Unda, Victor P.
Published: (2026)
by: Unda, Victor P.
Published: (2026)
Exploring new Approaches for Information Retrieval through Natural Language Processing
by: Raj, Manak, et al.
Published: (2025)
by: Raj, Manak, et al.
Published: (2025)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
by: Demir, M. Mikail, et al.
Published: (2025)
by: Demir, M. Mikail, et al.
Published: (2025)
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
by: Vileikytė, Brigita, et al.
Published: (2024)
by: Vileikytė, Brigita, et al.
Published: (2024)
AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs
by: Perera, Manoj Madushanka, et al.
Published: (2026)
by: Perera, Manoj Madushanka, et al.
Published: (2026)
Cognis: Context-Aware Memory for Conversational AI Agents
by: Daftari, Parshva, et al.
Published: (2026)
by: Daftari, Parshva, et al.
Published: (2026)
skLEP: A Slovak General Language Understanding Benchmark
by: Šuppa, Marek, et al.
Published: (2025)
by: Šuppa, Marek, et al.
Published: (2025)
Relating Word Embedding Gender Biases to Gender Gaps: A Cross-Cultural Analysis
by: Friedman, Scott, et al.
Published: (2026)
by: Friedman, Scott, et al.
Published: (2026)
A Case Study of Balanced Query Recommendation on Wikipedia
by: Mishra, Harshit, et al.
Published: (2025)
by: Mishra, Harshit, et al.
Published: (2025)
MeVer at CheckThat! 2026: Cluster-Aware Hard-Negative Mining for Multilingual Scientific-Source Retrieval
by: Bakagianni, Juli, et al.
Published: (2026)
by: Bakagianni, Juli, et al.
Published: (2026)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
by: Kim, Jaemin, et al.
Published: (2024)
by: Kim, Jaemin, et al.
Published: (2024)
Evaluation of RAG Metrics for Question Answering in the Telecom Domain
by: Roychowdhury, Sujoy, et al.
Published: (2024)
by: Roychowdhury, Sujoy, et al.
Published: (2024)
A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation
by: Chen, Ziyang, et al.
Published: (2025)
by: Chen, Ziyang, et al.
Published: (2025)
Combating data scarcity in recommendation services: Integrating cognitive types of VARK and neural network technologies (LLM)
by: Zmanovskii, Nikita
Published: (2026)
by: Zmanovskii, Nikita
Published: (2026)
Disambiguation of Emotion Annotations by Contextualizing Events in Plausible Narratives
by: Schäfer, Johannes, et al.
Published: (2025)
by: Schäfer, Johannes, et al.
Published: (2025)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
by: Sun, Jingyi, et al.
Published: (2024)
by: Sun, Jingyi, et al.
Published: (2024)
From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering
by: Santos, José Guilherme Marques dos, et al.
Published: (2026)
by: Santos, José Guilherme Marques dos, et al.
Published: (2026)
XPath Agent: An Efficient XPath Programming Agent Based on LLM for Web Crawler
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
CAG: Chunked Augmented Generation for Google Chrome's Built-in Gemini Nano
by: Surulimuthu, Vivek Vellaiyappan, et al.
Published: (2024)
by: Surulimuthu, Vivek Vellaiyappan, et al.
Published: (2024)
Understanding and Improving Information Preservation in Prompt Compression for LLMs
by: Łajewska, Weronika, et al.
Published: (2025)
by: Łajewska, Weronika, et al.
Published: (2025)
Evaluating Named Entity Recognition Models for Russian Cultural News Texts: From BERT to LLM
by: Levchenko, Maria
Published: (2025)
by: Levchenko, Maria
Published: (2025)
Similar Items
-
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
by: Günther, Michael, et al.
Published: (2025) -
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
by: Sturua, Saba, et al.
Published: (2024) -
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
by: Jha, Rohan, et al.
Published: (2024) -
Efficient Code Embeddings from Code Generation Models
by: Kryvosheieva, Daria, et al.
Published: (2025) -
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
by: Mohr, Isabelle, et al.
Published: (2024)