One Thousand and One Pairs: A "novel" challenge for long-context language models
Fuente:
arXiv
Saved in:
| Main Authors: | Karpinska, Marzena, Thai, Katherine, Lo, Kyle, Goyal, Tanya, Iyyer, Mohit |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025)
by: Kim, Yekyung, et al.
Published: (2025)
Does quantization affect models' performance on long-context tasks?
by: Mekala, Anmol, et al.
Published: (2025)
by: Mekala, Anmol, et al.
Published: (2025)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
FABLES: Evaluating faithfulness and content selection in book-length summarization
by: Kim, Yekyung, et al.
Published: (2024)
by: Kim, Yekyung, et al.
Published: (2024)
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
by: Chang, Yapei, et al.
Published: (2023)
by: Chang, Yapei, et al.
Published: (2023)
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
by: Arora, Shane, et al.
Published: (2024)
by: Arora, Shane, et al.
Published: (2024)
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature
by: Srivastava, Alisha, et al.
Published: (2025)
by: Srivastava, Alisha, et al.
Published: (2025)
AI use in American newspapers is widespread, uneven, and rarely disclosed
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
Literary Evidence Retrieval via Long-Context Language Models
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
BEARCUBS: A benchmark for computer-using web agents
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization
by: Huang, Chengyu, et al.
Published: (2025)
by: Huang, Chengyu, et al.
Published: (2025)
Argument Collapse: LLMs Flatten Long-Form Public Debate
by: Kim, Yekyung, et al.
Published: (2026)
by: Kim, Yekyung, et al.
Published: (2026)
CLIPPER: Compression enables long-context synthetic data generation
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
EditLens: Quantifying the Extent of AI Editing in Text
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
Are language models rational? The case of coherence norms and belief revision
by: Hofweber, Thomas, et al.
Published: (2024)
by: Hofweber, Thomas, et al.
Published: (2024)
Dissociating language and thought in large language models
by: Mahowald, Kyle, et al.
Published: (2023)
by: Mahowald, Kyle, et al.
Published: (2023)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
Updating Parametric Knowledge with Context Distillation Retains Post-Training Capabilities
by: Padmanabhan, Shankar, et al.
Published: (2026)
by: Padmanabhan, Shankar, et al.
Published: (2026)
PostMark: A Robust Blackbox Watermark for Large Language Models
by: Chang, Yapei, et al.
Published: (2024)
by: Chang, Yapei, et al.
Published: (2024)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
by: Li, Aochong Oliver, et al.
Published: (2025)
by: Li, Aochong Oliver, et al.
Published: (2025)
The advantages of context specific language models: the case of the Erasmian Language Model
by: Gonçalves, João, et al.
Published: (2024)
by: Gonçalves, João, et al.
Published: (2024)
Can large language models explore in-context?
by: Krishnamurthy, Akshay, et al.
Published: (2024)
by: Krishnamurthy, Akshay, et al.
Published: (2024)
A systematic framework for generating novel experimental hypotheses from language models
by: Misra, Kanishka, et al.
Published: (2024)
by: Misra, Kanishka, et al.
Published: (2024)
Just-in-time and distributed task representations in language models
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy
by: Yoshida, Davis, et al.
Published: (2023)
by: Yoshida, Davis, et al.
Published: (2023)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
by: Chang, Yapei, et al.
Published: (2025)
by: Chang, Yapei, et al.
Published: (2025)
A survey on fairness of large language models in e-commerce: progress, application, and challenge
by: Ren, Qingyang, et al.
Published: (2024)
by: Ren, Qingyang, et al.
Published: (2024)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
by: Song, Yixiao, et al.
Published: (2024)
by: Song, Yixiao, et al.
Published: (2024)
Why is constrained neural language generation particularly challenging?
by: Garbacea, Cristina, et al.
Published: (2022)
by: Garbacea, Cristina, et al.
Published: (2022)
The "LLM World of Words" English free association norms generated by large language models
by: Abramski, Katherine, et al.
Published: (2024)
by: Abramski, Katherine, et al.
Published: (2024)
One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
by: Khan, Zaid, et al.
Published: (2025)
by: Khan, Zaid, et al.
Published: (2025)
A novel language model for predicting serious adverse event results in clinical trials from their prospective registrations
by: Hu, Qixuan, et al.
Published: (2025)
by: Hu, Qixuan, et al.
Published: (2025)
Adjoint sharding for very long context training of state space models
by: Xu, Xingzi, et al.
Published: (2025)
by: Xu, Xingzi, et al.
Published: (2025)
On the generalization of language models from in-context learning and finetuning: a controlled study
by: Lampinen, Andrew K., et al.
Published: (2025)
by: Lampinen, Andrew K., et al.
Published: (2025)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
AIC CTU@FEVER 8: On-premise fact checking through long context RAG
by: Ullrich, Herbert, et al.
Published: (2025)
by: Ullrich, Herbert, et al.
Published: (2025)
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
by: Shen, Ye, et al.
Published: (2025)
by: Shen, Ye, et al.
Published: (2025)
A survey of textual cyber abuse detection using cutting-edge language models and large language models
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
by: Diaz-Garcia, Jose A., et al.
Published: (2025)
Towards robust long-context understanding of large language model via active recap learning
by: Hui, Chenyu
Published: (2026)
by: Hui, Chenyu
Published: (2026)
Fluent dreaming for language models
by: Thompson, T. Ben, et al.
Published: (2024)
by: Thompson, T. Ben, et al.
Published: (2024)
Similar Items
-
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025) -
Does quantization affect models' performance on long-context tasks?
by: Mekala, Anmol, et al.
Published: (2025) -
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025) -
FABLES: Evaluating faithfulness and content selection in book-length summarization
by: Kim, Yekyung, et al.
Published: (2024) -
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
by: Chang, Yapei, et al.
Published: (2023)