Reasoning on Multiple Needles In A Haystack
Fuente:
arXiv
Saved in:
| Main Author: | Wang, Yidong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Needle in the Haystack for Memory Based Large Language Models
by: Nelson, Elliot, et al.
Published: (2024)
by: Nelson, Elliot, et al.
Published: (2024)
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
by: Bianchi, Owen, et al.
Published: (2025)
by: Bianchi, Owen, et al.
Published: (2025)
DENIAHL: In-Context Features Influence LLM Needle-In-A-Haystack Abilities
by: Dai, Hui, et al.
Published: (2024)
by: Dai, Hui, et al.
Published: (2024)
From Haystack to Needle: Label Space Reduction for Zero-shot Classification
by: Vandemoortele, Nathan, et al.
Published: (2025)
by: Vandemoortele, Nathan, et al.
Published: (2025)
Needle In A Multimodal Haystack
by: Wang, Weiyun, et al.
Published: (2024)
by: Wang, Weiyun, et al.
Published: (2024)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
by: Kuratov, Yuri, et al.
Published: (2024)
by: Kuratov, Yuri, et al.
Published: (2024)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
by: Xiong, Zheyang, et al.
Published: (2024)
by: Xiong, Zheyang, et al.
Published: (2024)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
by: Sharma, Aditya, et al.
Published: (2024)
by: Sharma, Aditya, et al.
Published: (2024)
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
by: Wang, Hengyi, et al.
Published: (2024)
by: Wang, Hengyi, et al.
Published: (2024)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
by: Kuratov, Yuri, et al.
Published: (2024)
by: Kuratov, Yuri, et al.
Published: (2024)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
by: Aksoy, Sinan G., et al.
Published: (2026)
by: Aksoy, Sinan G., et al.
Published: (2026)
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025)
by: Shah, Rishi Rajesh, et al.
Published: (2025)
Two Causally Related Needles in a Video Haystack
by: Li, Miaoyu, et al.
Published: (2025)
by: Li, Miaoyu, et al.
Published: (2025)
GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
by: Byerly, Adam, et al.
Published: (2025)
by: Byerly, Adam, et al.
Published: (2025)
Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
by: Li, Mufei, et al.
Published: (2025)
by: Li, Mufei, et al.
Published: (2025)
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
by: Wang, Yumeng, et al.
Published: (2025)
by: Wang, Yumeng, et al.
Published: (2025)
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
by: Xu, Xiaoyue, et al.
Published: (2024)
by: Xu, Xiaoyue, et al.
Published: (2024)
Mirror: A Multiple-perspective Self-Reflection Method for Knowledge-rich Reasoning
by: Yan, Hanqi, et al.
Published: (2024)
by: Yan, Hanqi, et al.
Published: (2024)
NeedleChain: Measuring Intact Context Comprehension Capability of Large Language Models
by: Moon, Hyeonseok, et al.
Published: (2025)
by: Moon, Hyeonseok, et al.
Published: (2025)
Answering Questions by Meta-Reasoning over Multiple Chains of Thought
by: Yoran, Ori, et al.
Published: (2023)
by: Yoran, Ori, et al.
Published: (2023)
LLM Augmentations to support Analytical Reasoning over Multiple Documents
by: Yousuf, Raquib Bin, et al.
Published: (2024)
by: Yousuf, Raquib Bin, et al.
Published: (2024)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024)
by: Palta, Shramay, et al.
Published: (2024)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
by: Zhang, Yushun, et al.
Published: (2026)
by: Zhang, Yushun, et al.
Published: (2026)
Beyond the Needle's Illusion: Decoupled Evaluation of Evidence Access and Use under Semantic Interference at 326M-Token Scale
by: Lin, Tianwei, et al.
Published: (2026)
by: Lin, Tianwei, et al.
Published: (2026)
Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
by: Schumann, Raphael, et al.
Published: (2025)
by: Schumann, Raphael, et al.
Published: (2025)
Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
by: Kim, Bosung, et al.
Published: (2025)
by: Kim, Bosung, et al.
Published: (2025)
Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long Contexts
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
by: Huybrechts, Goeric, et al.
Published: (2025)
by: Huybrechts, Goeric, et al.
Published: (2025)
Aligning AI Research with the Needs of Clinical Coding Workflows: Eight Recommendations Based on US Data Analysis and Critical Review
by: Gan, Yidong, et al.
Published: (2024)
by: Gan, Yidong, et al.
Published: (2024)
Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection
by: Kim, Jun Seo, et al.
Published: (2025)
by: Kim, Jun Seo, et al.
Published: (2025)
Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
by: Wang, Yuhui, et al.
Published: (2025)
by: Wang, Yuhui, et al.
Published: (2025)
CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
TrustDataFilter:Leveraging Trusted Knowledge Base Data for More Effective Filtering of Unknown Information
by: Zhang, Jinghong, et al.
Published: (2025)
by: Zhang, Jinghong, et al.
Published: (2025)
KG-Reasoner: A Reinforced Model for End-to-End Multi-Hop Knowledge Graph Reasoning
by: Wang, Shuai, et al.
Published: (2026)
by: Wang, Shuai, et al.
Published: (2026)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
by: Jin, Keyan, et al.
Published: (2025)
by: Jin, Keyan, et al.
Published: (2025)
A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning
by: Wang, Ziqi, et al.
Published: (2025)
by: Wang, Ziqi, et al.
Published: (2025)
Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application
by: Yang, Chuanpeng, et al.
Published: (2024)
by: Yang, Chuanpeng, et al.
Published: (2024)
Similar Items
-
Needle in the Haystack for Memory Based Large Language Models
by: Nelson, Elliot, et al.
Published: (2024) -
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
by: Bianchi, Owen, et al.
Published: (2025) -
DENIAHL: In-Context Features Influence LLM Needle-In-A-Haystack Abilities
by: Dai, Hui, et al.
Published: (2024) -
From Haystack to Needle: Label Space Reduction for Zero-shot Classification
by: Vandemoortele, Nathan, et al.
Published: (2025) -
Needle In A Multimodal Haystack
by: Wang, Weiyun, et al.
Published: (2024)