Does quantization affect models' performance on long-context tasks?
Fuente:
arXiv
Saved in:
| Main Authors: | Mekala, Anmol, Atmakuru, Anirudh, Song, Yixiao, Karpinska, Marzena, Iyyer, Mohit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
One Thousand and One Pairs: A "novel" challenge for long-context language models
by: Karpinska, Marzena, et al.
Published: (2024)
by: Karpinska, Marzena, et al.
Published: (2024)
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025)
by: Kim, Yekyung, et al.
Published: (2025)
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
by: Arora, Shane, et al.
Published: (2024)
by: Arora, Shane, et al.
Published: (2024)
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature
by: Srivastava, Alisha, et al.
Published: (2025)
by: Srivastava, Alisha, et al.
Published: (2025)
FABLES: Evaluating faithfulness and content selection in book-length summarization
by: Kim, Yekyung, et al.
Published: (2024)
by: Kim, Yekyung, et al.
Published: (2024)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
by: Song, Yixiao, et al.
Published: (2024)
by: Song, Yixiao, et al.
Published: (2024)
BEARCUBS: A benchmark for computer-using web agents
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
Argument Collapse: LLMs Flatten Long-Form Public Debate
by: Kim, Yekyung, et al.
Published: (2026)
by: Kim, Yekyung, et al.
Published: (2026)
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
by: Chang, Yapei, et al.
Published: (2023)
by: Chang, Yapei, et al.
Published: (2023)
CLIPPER: Compression enables long-context synthetic data generation
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
AI use in American newspapers is widespread, uneven, and rarely disclosed
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
Localizing and Mitigating Errors in Long-form Question Answering
by: Sachdeva, Rachneet, et al.
Published: (2024)
by: Sachdeva, Rachneet, et al.
Published: (2024)
PostMark: A Robust Blackbox Watermark for Large Language Models
by: Chang, Yapei, et al.
Published: (2024)
by: Chang, Yapei, et al.
Published: (2024)
Superhuman performance of a large language model on the reasoning tasks of a physician
by: Brodeur, Peter G., et al.
Published: (2024)
by: Brodeur, Peter G., et al.
Published: (2024)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
by: Chang, Yapei, et al.
Published: (2025)
by: Chang, Yapei, et al.
Published: (2025)
Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
by: Hegde, Pradyoth, et al.
Published: (2025)
by: Hegde, Pradyoth, et al.
Published: (2025)
When is the consistent prediction likely to be a correct prediction?
by: Nguyen, Alex, et al.
Published: (2024)
by: Nguyen, Alex, et al.
Published: (2024)
Adjoint sharding for very long context training of state space models
by: Xu, Xingzi, et al.
Published: (2025)
by: Xu, Xingzi, et al.
Published: (2025)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
by: Kodali, Prashant, et al.
Published: (2024)
by: Kodali, Prashant, et al.
Published: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
AIC CTU@FEVER 8: On-premise fact checking through long context RAG
by: Ullrich, Herbert, et al.
Published: (2025)
by: Ullrich, Herbert, et al.
Published: (2025)
Just-in-time and distributed task representations in language models
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Literary Evidence Retrieval via Long-Context Language Models
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
Code-enabled language models can outperform reasoning models on diverse tasks
by: Zhang, Cedegao E., et al.
Published: (2025)
by: Zhang, Cedegao E., et al.
Published: (2025)
Auxiliary task demands mask the capabilities of smaller language models
by: Hu, Jennifer, et al.
Published: (2024)
by: Hu, Jennifer, et al.
Published: (2024)
Integration of cognitive tasks into artificial general intelligence test for large models
by: Qu, Youzhi, et al.
Published: (2024)
by: Qu, Youzhi, et al.
Published: (2024)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
by: Ilia, Evgenia, et al.
Published: (2024)
by: Ilia, Evgenia, et al.
Published: (2024)
Does GPT-4 surpass human performance in linguistic pragmatics?
by: Bojic, Ljubisa, et al.
Published: (2023)
by: Bojic, Ljubisa, et al.
Published: (2023)
Frankentext: Stitching random text fragments into long-form narratives
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
Are language models rational? The case of coherence norms and belief revision
by: Hofweber, Thomas, et al.
Published: (2024)
by: Hofweber, Thomas, et al.
Published: (2024)
Clinical ModernBERT: An efficient and long context encoder for biomedical text
by: Lee, Simon A., et al.
Published: (2025)
by: Lee, Simon A., et al.
Published: (2025)
Large Language Model as a Universal Clinical Multi-task Decoder
by: Wu, Yujiang, et al.
Published: (2024)
by: Wu, Yujiang, et al.
Published: (2024)
Language translation, and change of accent for speech-to-speech task using diffusion model
by: Mishra, Abhishek, et al.
Published: (2025)
by: Mishra, Abhishek, et al.
Published: (2025)
Adaptively profiling models with task elicitation
by: Brown, Davis, et al.
Published: (2025)
by: Brown, Davis, et al.
Published: (2025)
Evidence from counterfactual tasks supports emergent analogical reasoning in large language models
by: Webb, Taylor, et al.
Published: (2024)
by: Webb, Taylor, et al.
Published: (2024)
Cartridges: Lightweight and general-purpose long context representations via self-study
by: Eyuboglu, Sabri, et al.
Published: (2025)
by: Eyuboglu, Sabri, et al.
Published: (2025)
PLD+: Accelerating LLM inference by leveraging Language Model Artifacts
by: Somasundaram, Shwetha, et al.
Published: (2024)
by: Somasundaram, Shwetha, et al.
Published: (2024)
Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis
by: Attaluri, Kaushal, et al.
Published: (2024)
by: Attaluri, Kaushal, et al.
Published: (2024)
The advantages of context specific language models: the case of the Erasmian Language Model
by: Gonçalves, João, et al.
Published: (2024)
by: Gonçalves, João, et al.
Published: (2024)
Similar Items
-
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025) -
One Thousand and One Pairs: A "novel" challenge for long-context language models
by: Karpinska, Marzena, et al.
Published: (2024) -
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025) -
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
by: Arora, Shane, et al.
Published: (2024) -
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature
by: Srivastava, Alisha, et al.
Published: (2025)