Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Xiaoyue, Ye, Qinyuan, Ren, Xiang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
par: Ye, Qinyuan, et autres
Publié: (2025)
par: Ye, Qinyuan, et autres
Publié: (2025)
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
par: Guo, Yipin, et autres
Publié: (2024)
par: Guo, Yipin, et autres
Publié: (2024)
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
par: Wang, Hengyi, et autres
Publié: (2024)
par: Wang, Hengyi, et autres
Publié: (2024)
Auto-ICL: In-Context Learning without Human Supervision
par: Yang, Jinghan, et autres
Publié: (2023)
par: Yang, Jinghan, et autres
Publié: (2023)
Needle in the Haystack for Memory Based Large Language Models
par: Nelson, Elliot, et autres
Publié: (2024)
par: Nelson, Elliot, et autres
Publié: (2024)
RetICL: Sequential Retrieval of In-Context Examples with Reinforcement Learning
par: Scarlatos, Alexander, et autres
Publié: (2023)
par: Scarlatos, Alexander, et autres
Publié: (2023)
DENIAHL: In-Context Features Influence LLM Needle-In-A-Haystack Abilities
par: Dai, Hui, et autres
Publié: (2024)
par: Dai, Hui, et autres
Publié: (2024)
Revisiting In-Context Learning with Long Context Language Models
par: Baek, Jinheon, et autres
Publié: (2024)
par: Baek, Jinheon, et autres
Publié: (2024)
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
par: Huybrechts, Goeric, et autres
Publié: (2025)
par: Huybrechts, Goeric, et autres
Publié: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
par: Zeng, Zhiyuan, et autres
Publié: (2025)
par: Zeng, Zhiyuan, et autres
Publié: (2025)
CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations
par: Kothapalli, Vignesh, et autres
Publié: (2025)
par: Kothapalli, Vignesh, et autres
Publié: (2025)
Jailbreaking in the Haystack
par: Shah, Rishi Rajesh, et autres
Publié: (2025)
par: Shah, Rishi Rajesh, et autres
Publié: (2025)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
par: Mujahid, Zain Muhammad, et autres
Publié: (2025)
par: Mujahid, Zain Muhammad, et autres
Publié: (2025)
Artificial Hippocampus Networks for Efficient Long-Context Modeling
par: Fang, Yunhao, et autres
Publié: (2025)
par: Fang, Yunhao, et autres
Publié: (2025)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
par: Kuratov, Yuri, et autres
Publié: (2024)
par: Kuratov, Yuri, et autres
Publié: (2024)
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
par: Bianchi, Owen, et autres
Publié: (2025)
par: Bianchi, Owen, et autres
Publié: (2025)
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
par: Chen, Yukang, et autres
Publié: (2023)
par: Chen, Yukang, et autres
Publié: (2023)
Lifelong Safety Alignment for Language Models
par: Wang, Haoyu, et autres
Publié: (2025)
par: Wang, Haoyu, et autres
Publié: (2025)
Prompt Engineering a Prompt Engineer
par: Ye, Qinyuan, et autres
Publié: (2023)
par: Ye, Qinyuan, et autres
Publié: (2023)
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
par: Xu, Chejian, et autres
Publié: (2025)
par: Xu, Chejian, et autres
Publié: (2025)
UltraEdit: Training-, Subject-, and Memory-Free Lifelong Editing in Language Models
par: Gu, Xiaojie, et autres
Publié: (2025)
par: Gu, Xiaojie, et autres
Publié: (2025)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
par: Mishra, Aayush, et autres
Publié: (2025)
par: Mishra, Aayush, et autres
Publié: (2025)
From Haystack to Needle: Label Space Reduction for Zero-shot Classification
par: Vandemoortele, Nathan, et autres
Publié: (2025)
par: Vandemoortele, Nathan, et autres
Publié: (2025)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
par: Wang, Wentian, et autres
Publié: (2024)
par: Wang, Wentian, et autres
Publié: (2024)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
par: MiniCPM Team, et autres
Publié: (2026)
par: MiniCPM Team, et autres
Publié: (2026)
ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models
par: Thonet, Thibaut, et autres
Publié: (2024)
par: Thonet, Thibaut, et autres
Publié: (2024)
Retrieval meets Long Context Large Language Models
par: Xu, Peng, et autres
Publié: (2023)
par: Xu, Peng, et autres
Publié: (2023)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
par: Li, Zeju, et autres
Publié: (2026)
par: Li, Zeju, et autres
Publié: (2026)
The Impossibility Triangle of Long-Context Modeling
par: Zhou, Yan
Publié: (2026)
par: Zhou, Yan
Publié: (2026)
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
par: Li, Junsong, et autres
Publié: (2025)
par: Li, Junsong, et autres
Publié: (2025)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
par: Chen, Yingfa, et autres
Publié: (2025)
par: Chen, Yingfa, et autres
Publié: (2025)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
par: Zhang, Junyang, et autres
Publié: (2025)
par: Zhang, Junyang, et autres
Publié: (2025)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
par: Zhao, James Xu, et autres
Publié: (2025)
par: Zhao, James Xu, et autres
Publié: (2025)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
par: Shi, Dachuan, et autres
Publié: (2025)
par: Shi, Dachuan, et autres
Publié: (2025)
Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
par: Zheng, Yicong, et autres
Publié: (2025)
par: Zheng, Yicong, et autres
Publié: (2025)
HiCI: Hierarchical Construction-Integration for Long-Context Attention
par: Zeng, Xiangyu, et autres
Publié: (2026)
par: Zeng, Xiangyu, et autres
Publié: (2026)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
par: Kuratov, Yuri, et autres
Publié: (2024)
par: Kuratov, Yuri, et autres
Publié: (2024)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
par: Xiong, Zheyang, et autres
Publié: (2024)
par: Xiong, Zheyang, et autres
Publié: (2024)
ELDER: Enhancing Lifelong Model Editing with Mixture-of-LoRA
par: Li, Jiaang, et autres
Publié: (2024)
par: Li, Jiaang, et autres
Publié: (2024)
Large Visual-Language Models Are Also Good Classifiers: A Study of In-Context Multimodal Fake News Detection
par: Jiang, Ye, et autres
Publié: (2024)
par: Jiang, Ye, et autres
Publié: (2024)
Documents similaires
-
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
par: Ye, Qinyuan, et autres
Publié: (2025) -
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
par: Guo, Yipin, et autres
Publié: (2024) -
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
par: Wang, Hengyi, et autres
Publié: (2024) -
Auto-ICL: In-Context Learning without Human Supervision
par: Yang, Jinghan, et autres
Publié: (2023) -
Needle in the Haystack for Memory Based Large Language Models
par: Nelson, Elliot, et autres
Publié: (2024)