HeceTokenizer: A Syllable-Based Tokenization Approach for Turkish Retrieval
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Gulgonul, Senol |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking the Role of Token Retrieval in Multi-Vector Retrieval
von: Lee, Jinhyuk, et al.
Veröffentlicht: (2023)
von: Lee, Jinhyuk, et al.
Veröffentlicht: (2023)
FGR-ColBERT: Identifying Fine-Grained Relevance Tokens During Retrieval
von: Jarolím, Antonín, et al.
Veröffentlicht: (2026)
von: Jarolím, Antonín, et al.
Veröffentlicht: (2026)
Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
von: An, Ruize, et al.
Veröffentlicht: (2025)
von: An, Ruize, et al.
Veröffentlicht: (2025)
A Theory for Token-Level Harmonization in Retrieval-Augmented Generation
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
Towards Lossless Token Pruning in Late-Interaction Retrieval Models
von: Zong, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zong, Yuxuan, et al.
Veröffentlicht: (2025)
Semi-Parametric Retrieval via Binary Bag-of-Tokens Index
von: Zhou, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhou, Jiawei, et al.
Veröffentlicht: (2024)
AdaGATE: Adaptive Gap-Aware Token-Efficient Evidence Assembly for Multi-Hop Retrieval-Augmented Generation
von: Guo, Yilin, et al.
Veröffentlicht: (2026)
von: Guo, Yilin, et al.
Veröffentlicht: (2026)
From Tokens to Concepts: Leveraging SAE for SPLADE
von: Zong, Yuxuan, et al.
Veröffentlicht: (2026)
von: Zong, Yuxuan, et al.
Veröffentlicht: (2026)
Scaling Retrieval-Based Language Models with a Trillion-Token Datastore
von: Shao, Rulin, et al.
Veröffentlicht: (2024)
von: Shao, Rulin, et al.
Veröffentlicht: (2024)
xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
TokenRec: Learning to Tokenize ID for LLM-based Generative Recommendation
von: Qu, Haohao, et al.
Veröffentlicht: (2024)
von: Qu, Haohao, et al.
Veröffentlicht: (2024)
Token and Span Classification for Entity Recognition in French Historical Encyclopedias
von: Moncla, Ludovic, et al.
Veröffentlicht: (2025)
von: Moncla, Ludovic, et al.
Veröffentlicht: (2025)
Structured Distillation for Personalized Agent Memory: 11x Token Reduction with Retrieval Preservation
von: Lewis, Sydney
Veröffentlicht: (2026)
von: Lewis, Sydney
Veröffentlicht: (2026)
Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling
von: Clavié, Benjamin, et al.
Veröffentlicht: (2024)
von: Clavié, Benjamin, et al.
Veröffentlicht: (2024)
FIT-RAG: Black-Box RAG with Factual Information and Token Reduction
von: Mao, Yuren, et al.
Veröffentlicht: (2024)
von: Mao, Yuren, et al.
Veröffentlicht: (2024)
Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
von: Albarede, Lucas, et al.
Veröffentlicht: (2025)
von: Albarede, Lucas, et al.
Veröffentlicht: (2025)
RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish
von: Köse, Süha Kağan, et al.
Veröffentlicht: (2026)
von: Köse, Süha Kağan, et al.
Veröffentlicht: (2026)
An Early FIRST Reproduction and Improvements to Single-Token Decoding for Fast Listwise Reranking
von: Chen, Zijian, et al.
Veröffentlicht: (2024)
von: Chen, Zijian, et al.
Veröffentlicht: (2024)
Token-wise Influential Training Data Retrieval for Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2024)
von: Lin, Huawei, et al.
Veröffentlicht: (2024)
A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering
von: Sidiropoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Sidiropoulos, Georgios, et al.
Veröffentlicht: (2024)
Pctx: Tokenizing Personalized Context for Generative Recommendation
von: Zhong, Qiyong, et al.
Veröffentlicht: (2025)
von: Zhong, Qiyong, et al.
Veröffentlicht: (2025)
TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
von: Ezerceli, Özay, et al.
Veröffentlicht: (2025)
von: Ezerceli, Özay, et al.
Veröffentlicht: (2025)
TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
von: Ezerceli, Özay, et al.
Veröffentlicht: (2025)
von: Ezerceli, Özay, et al.
Veröffentlicht: (2025)
NEAR$^2$: A Nested Embedding Approach to Efficient Product Retrieval and Ranking
von: Qian, Shenbin, et al.
Veröffentlicht: (2025)
von: Qian, Shenbin, et al.
Veröffentlicht: (2025)
Rethinking Schema Linking: A Context-Aware Bidirectional Retrieval Approach for Text-to-SQL
von: Nahid, Md Mahadi Hasan, et al.
Veröffentlicht: (2025)
von: Nahid, Md Mahadi Hasan, et al.
Veröffentlicht: (2025)
RAPID: Retrieval-Augmented Parallel Inference Drafting for Text-Based Video Event Retrieval
von: Nguyen, Long, et al.
Veröffentlicht: (2025)
von: Nguyen, Long, et al.
Veröffentlicht: (2025)
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
von: Tavakoli, Mohammad, et al.
Veröffentlicht: (2025)
von: Tavakoli, Mohammad, et al.
Veröffentlicht: (2025)
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
von: Xing, Tiancheng, et al.
Veröffentlicht: (2025)
von: Xing, Tiancheng, et al.
Veröffentlicht: (2025)
Synergistic Approach for Simultaneous Optimization of Monolingual, Cross-lingual, and Multilingual Information Retrieval
von: Elmahdy, Adel, et al.
Veröffentlicht: (2024)
von: Elmahdy, Adel, et al.
Veröffentlicht: (2024)
Are Information Retrieval Approaches Good at Harmonising Longitudinal Survey Questions in Social Science?
von: Li, Wing Yan, et al.
Veröffentlicht: (2025)
von: Li, Wing Yan, et al.
Veröffentlicht: (2025)
PersonalAI: A Systematic Comparison of Knowledge Graph Storage and Retrieval Approaches for Personalized LLM agents
von: Menschikov, Mikhail, et al.
Veröffentlicht: (2025)
von: Menschikov, Mikhail, et al.
Veröffentlicht: (2025)
A MapReduce Approach to Effectively Utilize Long Context Information in Retrieval Augmented Language Models
von: Zhang, Gongbo, et al.
Veröffentlicht: (2024)
von: Zhang, Gongbo, et al.
Veröffentlicht: (2024)
Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
GRACE: Generative Recommendation via Journey-Aware Sparse Attention on Chain-of-Thought Tokenization
von: Ma, Luyi, et al.
Veröffentlicht: (2025)
von: Ma, Luyi, et al.
Veröffentlicht: (2025)
On the Robustness of LLM-Based Dense Retrievers: A Systematic Analysis of Generalizability and Stability
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
von: Li, Yongkang, et al.
Veröffentlicht: (2026)
DREditor: An Time-efficient Approach for Building a Domain-specific Dense Retrieval Model
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
Graph-Based Retriever Captures the Long Tail of Biomedical Knowledge
von: Delile, Julien, et al.
Veröffentlicht: (2024)
von: Delile, Julien, et al.
Veröffentlicht: (2024)
Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
von: Chitale, Pranjal A., et al.
Veröffentlicht: (2025)
von: Chitale, Pranjal A., et al.
Veröffentlicht: (2025)
Evaluating the Retrieval Component in LLM-Based Question Answering Systems
von: Alinejad, Ashkan, et al.
Veröffentlicht: (2024)
von: Alinejad, Ashkan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Rethinking the Role of Token Retrieval in Multi-Vector Retrieval
von: Lee, Jinhyuk, et al.
Veröffentlicht: (2023) -
FGR-ColBERT: Identifying Fine-Grained Relevance Tokens During Retrieval
von: Jarolím, Antonín, et al.
Veröffentlicht: (2026) -
Text2Token: Unsupervised Text Representation Learning with Token Target Prediction
von: An, Ruize, et al.
Veröffentlicht: (2025) -
A Theory for Token-Level Harmonization in Retrieval-Augmented Generation
von: Xu, Shicheng, et al.
Veröffentlicht: (2024) -
A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)