GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Byerly, Adam, Khashabi, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
von: Byerly, Adam, et al.
Veröffentlicht: (2024)
von: Byerly, Adam, et al.
Veröffentlicht: (2024)
Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
Reasoning on Multiple Needles In A Haystack
von: Wang, Yidong
Veröffentlicht: (2025)
von: Wang, Yidong
Veröffentlicht: (2025)
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
von: Bianchi, Owen, et al.
Veröffentlicht: (2025)
von: Bianchi, Owen, et al.
Veröffentlicht: (2025)
Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long Contexts
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
DENIAHL: In-Context Features Influence LLM Needle-In-A-Haystack Abilities
von: Dai, Hui, et al.
Veröffentlicht: (2024)
von: Dai, Hui, et al.
Veröffentlicht: (2024)
U-NIAH: Unified RAG and LLM Evaluation for Long Context Needle-In-A-Haystack
von: Gao, Yunfan, et al.
Veröffentlicht: (2025)
von: Gao, Yunfan, et al.
Veröffentlicht: (2025)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
von: Hengle, Amey, et al.
Veröffentlicht: (2024)
von: Hengle, Amey, et al.
Veröffentlicht: (2024)
Needle in the Haystack for Memory Based Large Language Models
von: Nelson, Elliot, et al.
Veröffentlicht: (2024)
von: Nelson, Elliot, et al.
Veröffentlicht: (2024)
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
von: Sharma, Aditya, et al.
Veröffentlicht: (2024)
von: Sharma, Aditya, et al.
Veröffentlicht: (2024)
Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks?
von: Roberts, Jonathan, et al.
Veröffentlicht: (2024)
von: Roberts, Jonathan, et al.
Veröffentlicht: (2024)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024)
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024)
From Haystack to Needle: Label Space Reduction for Zero-shot Classification
von: Vandemoortele, Nathan, et al.
Veröffentlicht: (2025)
von: Vandemoortele, Nathan, et al.
Veröffentlicht: (2025)
An Infinite Needle in a Finite Haystack: Finding Infinite Counter-Models in Deductive Verification
von: Elad, Neta, et al.
Veröffentlicht: (2023)
von: Elad, Neta, et al.
Veröffentlicht: (2023)
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)
von: Sileo, Damien
Veröffentlicht: (2025)
von: Sileo, Damien
Veröffentlicht: (2025)
Finding a Needle in the Adversarial Haystack: A Targeted Paraphrasing Approach For Uncovering Edge Cases with Minimal Distribution Distortion
von: Kassem, Aly M., et al.
Veröffentlicht: (2024)
von: Kassem, Aly M., et al.
Veröffentlicht: (2024)
Do pretrained Transformers Learn In-Context by Gradient Descent?
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024)
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
von: Kim, Sungwon, et al.
Veröffentlicht: (2025)
von: Kim, Sungwon, et al.
Veröffentlicht: (2025)
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
von: Li, Tianjian, et al.
Veröffentlicht: (2025)
von: Li, Tianjian, et al.
Veröffentlicht: (2025)
ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers
von: Fang, Zhouxiang, et al.
Veröffentlicht: (2025)
von: Fang, Zhouxiang, et al.
Veröffentlicht: (2025)
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
von: Laban, Philippe, et al.
Veröffentlicht: (2024)
von: Laban, Philippe, et al.
Veröffentlicht: (2024)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
von: Li, Mufei, et al.
Veröffentlicht: (2025)
von: Li, Mufei, et al.
Veröffentlicht: (2025)
Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation
von: Li, Zichong, et al.
Veröffentlicht: (2026)
von: Li, Zichong, et al.
Veröffentlicht: (2026)
Needle In A Multimodal Haystack
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
von: Vodrahalli, Kiran, et al.
Veröffentlicht: (2024)
von: Vodrahalli, Kiran, et al.
Veröffentlicht: (2024)
Evaluating the Evaluators: Are readability metrics good measures of readability?
von: Cachola, Isabel, et al.
Veröffentlicht: (2025)
von: Cachola, Isabel, et al.
Veröffentlicht: (2025)
Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
Can Memory-Augmented Language Models Generalize on Reasoning-in-a-Haystack Tasks?
von: Das, Payel, et al.
Veröffentlicht: (2025)
von: Das, Payel, et al.
Veröffentlicht: (2025)
Principled Context Engineering for RAG: Statistical Guarantees via Conformal Prediction
von: Chakraborty, Debashish, et al.
Veröffentlicht: (2025)
von: Chakraborty, Debashish, et al.
Veröffentlicht: (2025)
Jailbreaking in the Haystack
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
von: Xu, Xiaoyue, et al.
Veröffentlicht: (2024)
von: Xu, Xiaoyue, et al.
Veröffentlicht: (2024)
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2025)
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2025)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2026)
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2026)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
von: Ou, Jiefu, et al.
Veröffentlicht: (2024)
von: Ou, Jiefu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
von: Byerly, Adam, et al.
Veröffentlicht: (2024) -
Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
von: Lu, Taiming, et al.
Veröffentlicht: (2024) -
Reasoning on Multiple Needles In A Haystack
von: Wang, Yidong
Veröffentlicht: (2025) -
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
von: Bianchi, Owen, et al.
Veröffentlicht: (2025) -
Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long Contexts
von: Yu, Yifei, et al.
Veröffentlicht: (2025)