The Pre-Training Study of Expanded-SPLADE Models on Web Document Titles
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Hiun, Lee, Tae Kwan, Won, Taeryun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficiency and Effectiveness of SPLADE Models on Billion-Scale Web Document Title
di: Won, Taeryun, et al.
Pubblicazione: (2025)
di: Won, Taeryun, et al.
Pubblicazione: (2025)
The Role of Vocabularies in Learning Sparse Representations for Ranking
di: Kim, Hiun, et al.
Pubblicazione: (2025)
di: Kim, Hiun, et al.
Pubblicazione: (2025)
SPLADE-v3: New baselines for SPLADE
di: Lassance, Carlos, et al.
Pubblicazione: (2024)
di: Lassance, Carlos, et al.
Pubblicazione: (2024)
From Tokens to Concepts: Leveraging SAE for SPLADE
di: Zong, Yuxuan, et al.
Pubblicazione: (2026)
di: Zong, Yuxuan, et al.
Pubblicazione: (2026)
An Alternative to FLOPS Regularization to Effectively Productionize SPLADE-Doc
di: Porco, Aldo, et al.
Pubblicazione: (2025)
di: Porco, Aldo, et al.
Pubblicazione: (2025)
Mistral-SPLADE: LLMs for better Learned Sparse Retrieval
di: Doshi, Meet, et al.
Pubblicazione: (2024)
di: Doshi, Meet, et al.
Pubblicazione: (2024)
MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language Queries
di: Kim, Jonghwi, et al.
Pubblicazione: (2025)
di: Kim, Jonghwi, et al.
Pubblicazione: (2025)
CoSPLADE: Contextualizing SPLADE for Conversational Information Retrieval
di: Hai, Nam Le, et al.
Pubblicazione: (2023)
di: Hai, Nam Le, et al.
Pubblicazione: (2023)
Two-Step SPLADE: Simple, Efficient and Effective Approximation of SPLADE
di: Lassance, Carlos, et al.
Pubblicazione: (2024)
di: Lassance, Carlos, et al.
Pubblicazione: (2024)
Large Language Model Empowered Recommendation Meets All-domain Continual Pre-Training
di: Ma, Haokai, et al.
Pubblicazione: (2025)
di: Ma, Haokai, et al.
Pubblicazione: (2025)
From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines
di: Lee, Sunkyung, et al.
Pubblicazione: (2026)
di: Lee, Sunkyung, et al.
Pubblicazione: (2026)
Large language models are good medical coders, if provided with tools
di: Kwan, Keith
Pubblicazione: (2024)
di: Kwan, Keith
Pubblicazione: (2024)
HEISIR: Hierarchical Expansion of Inverted Semantic Indexing for Training-free Retrieval of Conversational Data using LLMs
di: Kim, Sangyeop, et al.
Pubblicazione: (2025)
di: Kim, Sangyeop, et al.
Pubblicazione: (2025)
An Open-Source Web-Based Tool for Evaluating Open-Source Large Language Models Leveraging Information Retrieval from Custom Documents
di: I, Godfrey
Pubblicazione: (2025)
di: I, Godfrey
Pubblicazione: (2025)
CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking
di: Vonášek, Josef, et al.
Pubblicazione: (2024)
di: Vonášek, Josef, et al.
Pubblicazione: (2024)
CorpusBrain++: A Continual Generative Pre-Training Framework for Knowledge-Intensive Language Tasks
di: Guo, Jiafeng, et al.
Pubblicazione: (2024)
di: Guo, Jiafeng, et al.
Pubblicazione: (2024)
Ontology-Based Knowledge Graph Framework for Industrial Standard Documents via Hierarchical and Propositional Structuring
di: Park, Jiin, et al.
Pubblicazione: (2025)
di: Park, Jiin, et al.
Pubblicazione: (2025)
BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback
di: Kim, Hyunseo, et al.
Pubblicazione: (2025)
di: Kim, Hyunseo, et al.
Pubblicazione: (2025)
HybridRAG: A Practical LLM-based ChatBot Framework based on Pre-Generated Q&A over Raw Unstructured Documents
di: Kim, Sungmoon, et al.
Pubblicazione: (2025)
di: Kim, Sungmoon, et al.
Pubblicazione: (2025)
Contextualization with SPLADE for High Recall Retrieval
di: Yang, Eugene
Pubblicazione: (2024)
di: Yang, Eugene
Pubblicazione: (2024)
Combining Language and Graph Models for Semi-structured Information Extraction on the Web
di: Hong, Zhi, et al.
Pubblicazione: (2024)
di: Hong, Zhi, et al.
Pubblicazione: (2024)
Understanding Fairness-Accuracy Trade-offs in Machine Learning Models: Does Promoting Fairness Undermine Performance?
di: Liu, Junhua, et al.
Pubblicazione: (2024)
di: Liu, Junhua, et al.
Pubblicazione: (2024)
Study on LLMs for Promptagator-Style Dense Retriever Training
di: Gwon, Daniel, et al.
Pubblicazione: (2025)
di: Gwon, Daniel, et al.
Pubblicazione: (2025)
Optimized Quran Passage Retrieval Using an Expanded QA Dataset and Fine-Tuned Language Models
di: Basem, Mohamed, et al.
Pubblicazione: (2024)
di: Basem, Mohamed, et al.
Pubblicazione: (2024)
Best Practices for Distilling Large Language Models into BERT for Web Search Ranking
di: Ye, Dezhi, et al.
Pubblicazione: (2024)
di: Ye, Dezhi, et al.
Pubblicazione: (2024)
Pre-training vs. Fine-tuning: A Reproducibility Study on Dense Retrieval Knowledge Acquisition
di: Yao, Zheng, et al.
Pubblicazione: (2025)
di: Yao, Zheng, et al.
Pubblicazione: (2025)
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
di: Chaffin, Antoine, et al.
Pubblicazione: (2026)
di: Chaffin, Antoine, et al.
Pubblicazione: (2026)
Language Modeling Using Tensor Trains
di: Su, Zhan, et al.
Pubblicazione: (2024)
di: Su, Zhan, et al.
Pubblicazione: (2024)
RAG-Enhanced Large Language Models for Dynamic Content Expiration Prediction in Web Search
di: Chen, Tingyu, et al.
Pubblicazione: (2026)
di: Chen, Tingyu, et al.
Pubblicazione: (2026)
Enhancing Language Models for Financial Relation Extraction with Named Entities and Part-of-Speech
di: Li, Menglin, et al.
Pubblicazione: (2024)
di: Li, Menglin, et al.
Pubblicazione: (2024)
BASES: Large-scale Web Search User Simulation with Large Language Model based Agents
di: Ren, Ruiyang, et al.
Pubblicazione: (2024)
di: Ren, Ruiyang, et al.
Pubblicazione: (2024)
When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIR
di: Ko, Dayoon, et al.
Pubblicazione: (2025)
di: Ko, Dayoon, et al.
Pubblicazione: (2025)
Leveraging Generative Models for Real-Time Query-Driven Text Summarization in Large-Scale Web Search
di: Xiong, Zeyu, et al.
Pubblicazione: (2025)
di: Xiong, Zeyu, et al.
Pubblicazione: (2025)
Re-Ranking Step by Step: Investigating Pre-Filtering for Re-Ranking with Large Language Models
di: Nouriinanloo, Baharan, et al.
Pubblicazione: (2024)
di: Nouriinanloo, Baharan, et al.
Pubblicazione: (2024)
LARA: Linguistic-Adaptive Retrieval-Augmentation for Multi-Turn Intent Classification
di: Liu, Junhua, et al.
Pubblicazione: (2024)
di: Liu, Junhua, et al.
Pubblicazione: (2024)
DiffuRank: Effective Document Reranking with Diffusion Language Models
di: Liu, Qi, et al.
Pubblicazione: (2026)
di: Liu, Qi, et al.
Pubblicazione: (2026)
Web Page Classification using LLMs for Crawling Support
di: Sasazawa, Yuichi, et al.
Pubblicazione: (2025)
di: Sasazawa, Yuichi, et al.
Pubblicazione: (2025)
Multilingual Attribute Extraction from News Web Pages
di: Bedrin, Pavel, et al.
Pubblicazione: (2025)
di: Bedrin, Pavel, et al.
Pubblicazione: (2025)
Summarization-Based Document IDs for Generative Retrieval with Language Models
di: Li, Haoxin, et al.
Pubblicazione: (2023)
di: Li, Haoxin, et al.
Pubblicazione: (2023)
Unifying Adversarial Robustness and Training Across Text Scoring Models
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2026)
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Efficiency and Effectiveness of SPLADE Models on Billion-Scale Web Document Title
di: Won, Taeryun, et al.
Pubblicazione: (2025) -
The Role of Vocabularies in Learning Sparse Representations for Ranking
di: Kim, Hiun, et al.
Pubblicazione: (2025) -
SPLADE-v3: New baselines for SPLADE
di: Lassance, Carlos, et al.
Pubblicazione: (2024) -
From Tokens to Concepts: Leveraging SAE for SPLADE
di: Zong, Yuxuan, et al.
Pubblicazione: (2026) -
An Alternative to FLOPS Regularization to Effectively Productionize SPLADE-Doc
di: Porco, Aldo, et al.
Pubblicazione: (2025)