NoLiMa: Long-Context Evaluation Beyond Literal Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Modarressi, Ali, Deilamsalehy, Hanieh, Dernoncourt, Franck, Bui, Trung, Rossi, Ryan A., Yoon, Seunghyun, Schütze, Hinrich |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMs
von: Parmar, Mihir, et al.
Veröffentlicht: (2024)
von: Parmar, Mihir, et al.
Veröffentlicht: (2024)
CORG: Generating Answers from Complex, Interrelated Contexts
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models
von: Nguyen, Minh, et al.
Veröffentlicht: (2024)
von: Nguyen, Minh, et al.
Veröffentlicht: (2024)
Consistent Document-Level Relation Extraction via Counterfactuals
von: Modarressi, Ali, et al.
Veröffentlicht: (2024)
von: Modarressi, Ali, et al.
Veröffentlicht: (2024)
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
RET-LLM: Towards a General Read-Write Memory for Large Language Models
von: Modarressi, Ali, et al.
Veröffentlicht: (2023)
von: Modarressi, Ali, et al.
Veröffentlicht: (2023)
Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation
von: Fang, Jiangnan, et al.
Veröffentlicht: (2026)
von: Fang, Jiangnan, et al.
Veröffentlicht: (2026)
SlimLM: An Efficient Small Language Model for On-Device Document Assistance
von: Pham, Thang M., et al.
Veröffentlicht: (2024)
von: Pham, Thang M., et al.
Veröffentlicht: (2024)
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
von: Hakimi, Ahmad Dawar, et al.
Veröffentlicht: (2025)
von: Hakimi, Ahmad Dawar, et al.
Veröffentlicht: (2025)
ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge
von: Taghavi, Zeinab Sadat, et al.
Veröffentlicht: (2025)
von: Taghavi, Zeinab Sadat, et al.
Veröffentlicht: (2025)
DynaSaur: Large Language Agents Beyond Predefined Actions
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Scaling Up Video Summarization Pretraining with Large Language Models
von: Argaw, Dawit Mureja, et al.
Veröffentlicht: (2024)
von: Argaw, Dawit Mureja, et al.
Veröffentlicht: (2024)
MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
von: Modarressi, Ali, et al.
Veröffentlicht: (2024)
von: Modarressi, Ali, et al.
Veröffentlicht: (2024)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
Drift No More? Context Equilibria in Multi-Turn LLM Interactions
von: Dongre, Vardhan, et al.
Veröffentlicht: (2025)
von: Dongre, Vardhan, et al.
Veröffentlicht: (2025)
Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
von: Lee, Hyunji, et al.
Veröffentlicht: (2025)
MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2024)
von: Kargaran, Amir Hossein, et al.
Veröffentlicht: (2024)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2024)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
von: Zang, Yuan, et al.
Veröffentlicht: (2025)
Lizard: An Efficient Linearization Framework for Large Language Models
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2025)
von: Van Nguyen, Chien, et al.
Veröffentlicht: (2025)
Personalized Graph-Based Retrieval for Large Language Models
von: Au, Steven, et al.
Veröffentlicht: (2025)
von: Au, Steven, et al.
Veröffentlicht: (2025)
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2024)
von: Gallegos, Isabel O., et al.
Veröffentlicht: (2024)
Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams
von: Kim, Jiyeon, et al.
Veröffentlicht: (2026)
von: Kim, Jiyeon, et al.
Veröffentlicht: (2026)
Retrieval Augmented Generation for Domain-specific Question Answering
von: Sharma, Sanat, et al.
Veröffentlicht: (2024)
von: Sharma, Sanat, et al.
Veröffentlicht: (2024)
LongLaMP: A Benchmark for Personalized Long-form Text Generation
von: Kumar, Ishita, et al.
Veröffentlicht: (2024)
von: Kumar, Ishita, et al.
Veröffentlicht: (2024)
A Multi-LLM Debiasing Framework
von: Owens, Deonna M., et al.
Veröffentlicht: (2024)
von: Owens, Deonna M., et al.
Veröffentlicht: (2024)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
Multi-LLM Text Summarization
von: Fang, Jiangnan, et al.
Veröffentlicht: (2024)
von: Fang, Jiangnan, et al.
Veröffentlicht: (2024)
ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution
von: Goswami, Kanika, et al.
Veröffentlicht: (2025)
von: Goswami, Kanika, et al.
Veröffentlicht: (2025)
PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback
von: Goswami, Kanika, et al.
Veröffentlicht: (2025)
von: Goswami, Kanika, et al.
Veröffentlicht: (2025)
PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents
von: Goswami, Kanika, et al.
Veröffentlicht: (2025)
von: Goswami, Kanika, et al.
Veröffentlicht: (2025)
Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation
von: Kennedy, Molly, et al.
Veröffentlicht: (2026)
von: Kennedy, Molly, et al.
Veröffentlicht: (2026)
XAMPLER: Learning to Retrieve Cross-Lingual In-Context Examples
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
von: Tian, Yanzhi, et al.
Veröffentlicht: (2026)
von: Tian, Yanzhi, et al.
Veröffentlicht: (2026)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
von: Liu, Yihong, et al.
Veröffentlicht: (2026)
von: Liu, Yihong, et al.
Veröffentlicht: (2026)
MaLA-500: Massive Language Adaptation of Large Language Models
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
von: Lin, Peiqin, et al.
Veröffentlicht: (2024)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
von: Weissweiler, Leonie, et al.
Veröffentlicht: (2024)
von: Weissweiler, Leonie, et al.
Veröffentlicht: (2024)
mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning
von: Ngo, Nghia Trung, et al.
Veröffentlicht: (2025)
von: Ngo, Nghia Trung, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025) -
Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMs
von: Parmar, Mihir, et al.
Veröffentlicht: (2024) -
CORG: Generating Answers from Complex, Interrelated Contexts
von: Lee, Hyunji, et al.
Veröffentlicht: (2025) -
Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models
von: Nguyen, Minh, et al.
Veröffentlicht: (2024) -
Consistent Document-Level Relation Extraction via Counterfactuals
von: Modarressi, Ali, et al.
Veröffentlicht: (2024)