Lost-in-the-Middle in Long-Text Generation: Synthetic Dataset, Evaluation Framework, and Mitigation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Junhao, Zhang, Richong, Kong, Fanshuang, Miao, Ziyang, Ye, Yanhan, Zheng, Yaowei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
por: Zheng, Yaowei, et al.
Publicado: (2024)
por: Zheng, Yaowei, et al.
Publicado: (2024)
EasyRAG: Efficient Retrieval-Augmented Generation Framework for Automated Network Operations
por: Feng, Zhangchi, et al.
Publicado: (2024)
por: Feng, Zhangchi, et al.
Publicado: (2024)
LH-Mix: Local Hierarchy Correlation Guided Mixup over Hierarchical Prompt Tuning
por: Kong, Fanshuang, et al.
Publicado: (2024)
por: Kong, Fanshuang, et al.
Publicado: (2024)
Activated Parameter Locating via Causal Intervention for Model Merging
por: Kong, Fanshuang, et al.
Publicado: (2024)
por: Kong, Fanshuang, et al.
Publicado: (2024)
What Works for 'Lost-in-the-Middle' in LLMs? A Study on GM-Extract and Mitigations
por: Gupte, Mihir, et al.
Publicado: (2025)
por: Gupte, Mihir, et al.
Publicado: (2025)
Never Lost in the Middle: Mastering Long-Context Question Answering with Position-Agnostic Decompositional Training
por: He, Junqing, et al.
Publicado: (2023)
por: He, Junqing, et al.
Publicado: (2023)
What Has Been Lost with Synthetic Evaluation?
por: Gill, Alexander, et al.
Publicado: (2025)
por: Gill, Alexander, et al.
Publicado: (2025)
General Table Question Answering via Answer-Formula Joint Generation
por: Wang, Zhongyuan, et al.
Publicado: (2025)
por: Wang, Zhongyuan, et al.
Publicado: (2025)
Dynamic Task Vector Grouping for Efficient Multi-Task Prompt Tuning
por: Zhang, Pieyi, et al.
Publicado: (2025)
por: Zhang, Pieyi, et al.
Publicado: (2025)
Easy Dataset: A Unified and Extensible Framework for Synthesizing LLM Fine-Tuning Data from Unstructured Documents
por: Miao, Ziyang, et al.
Publicado: (2025)
por: Miao, Ziyang, et al.
Publicado: (2025)
Generating Synthetic Datasets for Few-shot Prompt Tuning
por: Guo, Xu, et al.
Publicado: (2024)
por: Guo, Xu, et al.
Publicado: (2024)
CoRect: Context-Aware Logit Contrast for Hidden State Rectification to Resolve Knowledge Conflicts
por: Ma, Xuhua, et al.
Publicado: (2026)
por: Ma, Xuhua, et al.
Publicado: (2026)
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
por: Huang, Hailang, et al.
Publicado: (2024)
por: Huang, Hailang, et al.
Publicado: (2024)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
por: Li, Junjie, et al.
Publicado: (2026)
por: Li, Junjie, et al.
Publicado: (2026)
Measuring Diversity in Synthetic Datasets
por: Zhu, Yuchang, et al.
Publicado: (2025)
por: Zhu, Yuchang, et al.
Publicado: (2025)
Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models
por: Zhao, Juntu, et al.
Publicado: (2024)
por: Zhao, Juntu, et al.
Publicado: (2024)
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides
por: Zheng, Hao, et al.
Publicado: (2025)
por: Zheng, Hao, et al.
Publicado: (2025)
Code-Style In-Context Learning for Knowledge-Based Question Answering
por: Nie, Zhijie, et al.
Publicado: (2023)
por: Nie, Zhijie, et al.
Publicado: (2023)
PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
por: Tan, Haochen, et al.
Publicado: (2024)
por: Tan, Haochen, et al.
Publicado: (2024)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
por: Fang, Hao, et al.
Publicado: (2025)
por: Fang, Hao, et al.
Publicado: (2025)
Parameterized Synthetic Text Generation with SimpleStories
por: Finke, Lennart, et al.
Publicado: (2025)
por: Finke, Lennart, et al.
Publicado: (2025)
When Text Embedding Meets Large Language Model: A Comprehensive Survey
por: Nie, Zhijie, et al.
Publicado: (2024)
por: Nie, Zhijie, et al.
Publicado: (2024)
inversedMixup: Data Augmentation via Inverting Mixed Embeddings
por: Kong, Fanshuang, et al.
Publicado: (2026)
por: Kong, Fanshuang, et al.
Publicado: (2026)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
por: Chen, Ziyang, et al.
Publicado: (2026)
por: Chen, Ziyang, et al.
Publicado: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
por: Hao, Jitai, et al.
Publicado: (2026)
por: Hao, Jitai, et al.
Publicado: (2026)
Bayesian WeakS-to-Strong from Text Classification to Generation
por: Cui, Ziyun, et al.
Publicado: (2024)
por: Cui, Ziyun, et al.
Publicado: (2024)
Style Transfer as Bias Mitigation: Diffusion Models for Synthetic Mental Health Text for Arabic
por: Mankarious, Saad, et al.
Publicado: (2026)
por: Mankarious, Saad, et al.
Publicado: (2026)
Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias
por: Chowdhury, Borun D
Publicado: (2026)
por: Chowdhury, Borun D
Publicado: (2026)
Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation
por: Subramanian, Seganrasan, et al.
Publicado: (2025)
por: Subramanian, Seganrasan, et al.
Publicado: (2025)
Equilibrium Dynamics and Mitigation of Gender Bias in Synthetically Generated Data
por: Kattamuri, Ashish, et al.
Publicado: (2025)
por: Kattamuri, Ashish, et al.
Publicado: (2025)
Cooperative Memory Paging with Keyword Bookmarks for Long-Horizon LLM Conversations
por: Liu, Ziyang
Publicado: (2026)
por: Liu, Ziyang
Publicado: (2026)
Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset
por: Yue, Chongjian, et al.
Publicado: (2024)
por: Yue, Chongjian, et al.
Publicado: (2024)
Synthetic Dialogue Dataset Generation using LLM Agents
por: Abdullin, Yelaman, et al.
Publicado: (2024)
por: Abdullin, Yelaman, et al.
Publicado: (2024)
ChallengeMe: An Adversarial Learning-enabled Text Summarization Framework
por: Deng, Xiaoyu, et al.
Publicado: (2025)
por: Deng, Xiaoyu, et al.
Publicado: (2025)
FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
por: Jin, Song, et al.
Publicado: (2025)
por: Jin, Song, et al.
Publicado: (2025)
ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information
por: Hui, Zheng, et al.
Publicado: (2024)
por: Hui, Zheng, et al.
Publicado: (2024)
Lost in the Source Language: How Large Language Models Evaluate the Quality of Machine Translation
por: Huang, Xu, et al.
Publicado: (2024)
por: Huang, Xu, et al.
Publicado: (2024)
Generation of Synthetic Clinical Text: A Systematic Review
por: Alshaikhdeeb, Basel, et al.
Publicado: (2025)
por: Alshaikhdeeb, Basel, et al.
Publicado: (2025)
Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation
por: Salemi, Alireza, et al.
Publicado: (2025)
por: Salemi, Alireza, et al.
Publicado: (2025)
Do LLMs Align with My Task? Evaluating Text-to-SQL via Dataset Alignment
por: Rafiei, Davood, et al.
Publicado: (2025)
por: Rafiei, Davood, et al.
Publicado: (2025)
Ejemplares similares
-
LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
por: Zheng, Yaowei, et al.
Publicado: (2024) -
EasyRAG: Efficient Retrieval-Augmented Generation Framework for Automated Network Operations
por: Feng, Zhangchi, et al.
Publicado: (2024) -
LH-Mix: Local Hierarchy Correlation Guided Mixup over Hierarchical Prompt Tuning
por: Kong, Fanshuang, et al.
Publicado: (2024) -
Activated Parameter Locating via Causal Intervention for Model Merging
por: Kong, Fanshuang, et al.
Publicado: (2024) -
What Works for 'Lost-in-the-Middle' in LLMs? A Study on GM-Extract and Mitigations
por: Gupte, Mihir, et al.
Publicado: (2025)