Training With "Paraphrasing the Original Text" Teaches LLM to Better Retrieve in Long-context Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Yijiong, Huang, Yongfeng, Qi, Zhixiao, Zhou, Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
by: Yu, Yijiong, et al.
Published: (2024)
by: Yu, Yijiong, et al.
Published: (2024)
Do LLMs Really Think Step-by-step In Implicit Reasoning?
by: Yu, Yijiong
Published: (2024)
by: Yu, Yijiong
Published: (2024)
An Effective Framework to Help Large Language Models Handle Numeric-involved Long-context Tasks
by: Yu, Yijiong
Published: (2024)
by: Yu, Yijiong
Published: (2024)
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
by: Li, Yafu, et al.
Published: (2024)
by: Li, Yafu, et al.
Published: (2024)
AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts
by: Murakami, Soichiro, et al.
Published: (2025)
by: Murakami, Soichiro, et al.
Published: (2025)
Membox: Weaving Topic Continuity into Long-Range Memory for LLM Agents
by: Tao, Dehao, et al.
Published: (2026)
by: Tao, Dehao, et al.
Published: (2026)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
KLong: Training LLM Agent for Extremely Long-horizon Tasks
by: Liu, Yue, et al.
Published: (2026)
by: Liu, Yue, et al.
Published: (2026)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
by: Jiang, Ziyan, et al.
Published: (2024)
by: Jiang, Ziyan, et al.
Published: (2024)
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
by: Huo, Jiahao, et al.
Published: (2026)
by: Huo, Jiahao, et al.
Published: (2026)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Rethinking Text-based Protein Understanding: Retrieval or LLM?
by: Wu, Juntong, et al.
Published: (2025)
by: Wu, Juntong, et al.
Published: (2025)
SEMA-RAG: A Self-Evolving Multi-Agent Retrieval-Augmented Generation Framework for Medical Reasoning
by: Huang, Yongfeng, et al.
Published: (2026)
by: Huang, Yongfeng, et al.
Published: (2026)
Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
by: Li, Shiyu, et al.
Published: (2025)
by: Li, Shiyu, et al.
Published: (2025)
LLM-Confidence Reranker: A Training-Free Approach for Enhancing Retrieval-Augmented Generation Systems
by: Song, Zhipeng, et al.
Published: (2026)
by: Song, Zhipeng, et al.
Published: (2026)
Task-Aligned Tool Recommendation for Large Language Models
by: Gao, Hang, et al.
Published: (2024)
by: Gao, Hang, et al.
Published: (2024)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026)
by: Kaneko, Masahiro
Published: (2026)
MemLong: Memory-Augmented Retrieval for Long Text Modeling
by: Liu, Weijie, et al.
Published: (2024)
by: Liu, Weijie, et al.
Published: (2024)
PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
by: Zha, Yiwei, et al.
Published: (2025)
by: Zha, Yiwei, et al.
Published: (2025)
C$^3$TG: Conflict-aware, Composite, and Collaborative Controlled Text Generation
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
by: Horvitz, Zachary, et al.
Published: (2023)
by: Horvitz, Zachary, et al.
Published: (2023)
Long-context LLMs Struggle with Long In-context Learning
by: Li, Tianle, et al.
Published: (2024)
by: Li, Tianle, et al.
Published: (2024)
Supervised Knowledge Makes Large Language Models Better In-context Learners
by: Yang, Linyi, et al.
Published: (2023)
by: Yang, Linyi, et al.
Published: (2023)
Retrieval-Augmented Generation with Hierarchical Knowledge
by: Huang, Haoyu, et al.
Published: (2025)
by: Huang, Haoyu, et al.
Published: (2025)
Fine-grained Stateful Knowledge Exploration: Effective and Efficient Graph Retrieval with Large Language Models
by: Tao, Dehao, et al.
Published: (2024)
by: Tao, Dehao, et al.
Published: (2024)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Action Controlled Paraphrasing
by: Shi, Ning, et al.
Published: (2024)
by: Shi, Ning, et al.
Published: (2024)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
by: Ning, Xuefei, et al.
Published: (2024)
by: Ning, Xuefei, et al.
Published: (2024)
Retrieval Meets Reasoning: Dynamic In-Context Editing for Long-Text Understanding
by: Fei, Weizhi, et al.
Published: (2024)
by: Fei, Weizhi, et al.
Published: (2024)
LLMs Can Teach Themselves to Better Predict the Future
by: Turtel, Benjamin, et al.
Published: (2025)
by: Turtel, Benjamin, et al.
Published: (2025)
Rationale-Augmented Retrieval with Constrained LLM Re-Ranking for Task Discovery
by: Wei, Bowen
Published: (2025)
by: Wei, Bowen
Published: (2025)
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models
by: Qiu, Yifu, et al.
Published: (2025)
by: Qiu, Yifu, et al.
Published: (2025)
Language Modeling and Understanding Through Paraphrase Generation and Detection
by: Wahle, Jan Philip
Published: (2026)
by: Wahle, Jan Philip
Published: (2026)
Analyzing Persuasive Strategies in Meme Texts: A Fusion of Language Models with Paraphrase Enrichment
by: Nayak, Kota Shamanth Ramanath, et al.
Published: (2024)
by: Nayak, Kota Shamanth Ramanath, et al.
Published: (2024)
BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers
by: Xu, Ran, et al.
Published: (2024)
by: Xu, Ran, et al.
Published: (2024)
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
by: Huang, Tenghao, et al.
Published: (2026)
by: Huang, Tenghao, et al.
Published: (2026)
LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
by: Ma, Guangyuan, et al.
Published: (2025)
by: Ma, Guangyuan, et al.
Published: (2025)
With Greater Text Comes Greater Necessity: Inference-Time Training Helps Long Text Generation
by: Wang, Y., et al.
Published: (2024)
by: Wang, Y., et al.
Published: (2024)
Can reasoning models comprehend mathematical problems in Chinese ancient texts? An empirical study based on data from Suanjing Shishu
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
Similar Items
-
Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
by: Yu, Yijiong, et al.
Published: (2024) -
Do LLMs Really Think Step-by-step In Implicit Reasoning?
by: Yu, Yijiong
Published: (2024) -
An Effective Framework to Help Large Language Models Handle Numeric-involved Long-context Tasks
by: Yu, Yijiong
Published: (2024) -
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
by: Li, Yafu, et al.
Published: (2024) -
AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts
by: Murakami, Soichiro, et al.
Published: (2025)