CUB: Benchmarking Context Utilisation Techniques for Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hagström, Lovisa, Kim, Youna, Yu, Haeun, Lee, Sang-goo, Johansson, Richard, Cho, Hyunsoo, Augenstein, Isabelle |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Reality Check on Context Utilisation for Retrieval-Augmented Generation
di: Hagström, Lovisa, et al.
Pubblicazione: (2024)
di: Hagström, Lovisa, et al.
Pubblicazione: (2024)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
di: Yu, Haeun, et al.
Pubblicazione: (2024)
di: Yu, Haeun, et al.
Pubblicazione: (2024)
Language Model Re-rankers are Fooled by Lexical Similarities
di: Hagström, Lovisa, et al.
Pubblicazione: (2025)
di: Hagström, Lovisa, et al.
Pubblicazione: (2025)
UniKnow: A Unified Framework for Reliable Language Model Behavior across Parametric and External Knowledge
di: Kim, Youna, et al.
Pubblicazione: (2025)
di: Kim, Youna, et al.
Pubblicazione: (2025)
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
di: Pauli, Amalie Brogaard, et al.
Pubblicazione: (2024)
di: Pauli, Amalie Brogaard, et al.
Pubblicazione: (2024)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
di: Yu, Haeun, et al.
Pubblicazione: (2025)
di: Yu, Haeun, et al.
Pubblicazione: (2025)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024)
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024)
Unveiling Imitation Learning: Exploring the Impact of Data Falsity to Large Language Model
di: Cho, Hyunsoo
Pubblicazione: (2024)
di: Cho, Hyunsoo
Pubblicazione: (2024)
Adaptive Contrastive Decoding in Retrieval-Augmented Generation for Handling Noisy Contexts
di: Kim, Youna, et al.
Pubblicazione: (2024)
di: Kim, Youna, et al.
Pubblicazione: (2024)
When to Speak, When to Abstain: Contrastive Decoding with Abstention
di: Kim, Hyuhng Joon, et al.
Pubblicazione: (2024)
di: Kim, Hyuhng Joon, et al.
Pubblicazione: (2024)
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
di: Islam, Sekh Mainul, et al.
Pubblicazione: (2025)
di: Islam, Sekh Mainul, et al.
Pubblicazione: (2025)
Inertia in Moral and Value Judgments of Large Language Models
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
Claim Verification in the Age of Large Language Models: A Survey
di: Dmonte, Alphaeus, et al.
Pubblicazione: (2024)
di: Dmonte, Alphaeus, et al.
Pubblicazione: (2024)
Fact Recall, Heuristics or Pure Guesswork? Precise Interpretations of Language Models for Fact Completion
di: Saynova, Denitsa, et al.
Pubblicazione: (2024)
di: Saynova, Denitsa, et al.
Pubblicazione: (2024)
Investigating the Influence of Prompt-Specific Shortcuts in AI Generated Text Detection
di: Park, Choonghyun, et al.
Pubblicazione: (2024)
di: Park, Choonghyun, et al.
Pubblicazione: (2024)
Cleanse: Uncertainty Estimation Approach Using Clustering-based Semantic Consistency in LLMs
di: Joo, Minsuh, et al.
Pubblicazione: (2025)
di: Joo, Minsuh, et al.
Pubblicazione: (2025)
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
di: Sun, Jingyi, et al.
Pubblicazione: (2025)
di: Sun, Jingyi, et al.
Pubblicazione: (2025)
Understanding the Interplay between LLMs' Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025
di: Augenstein, Isabelle
Pubblicazione: (2026)
di: Augenstein, Isabelle
Pubblicazione: (2026)
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
di: Ghazaryan, Gayane, et al.
Pubblicazione: (2024)
di: Ghazaryan, Gayane, et al.
Pubblicazione: (2024)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
di: Arakelyan, Erik, et al.
Pubblicazione: (2024)
di: Arakelyan, Erik, et al.
Pubblicazione: (2024)
LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models
di: Kim, Yungi, et al.
Pubblicazione: (2024)
di: Kim, Yungi, et al.
Pubblicazione: (2024)
Can Community Notes Replace Professional Fact-Checkers?
di: Borenstein, Nadav, et al.
Pubblicazione: (2025)
di: Borenstein, Nadav, et al.
Pubblicazione: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
Aligning Language Models to Explicitly Handle Ambiguity
di: Kim, Hyuhng Joon, et al.
Pubblicazione: (2024)
di: Kim, Hyuhng Joon, et al.
Pubblicazione: (2024)
Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
di: Warren, Greta, et al.
Pubblicazione: (2025)
di: Warren, Greta, et al.
Pubblicazione: (2025)
M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models
di: Kwon, Yejin, et al.
Pubblicazione: (2025)
di: Kwon, Yejin, et al.
Pubblicazione: (2025)
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns
di: Pauli, Amalie Brogaard, et al.
Pubblicazione: (2026)
di: Pauli, Amalie Brogaard, et al.
Pubblicazione: (2026)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
di: Mujahid, Zain Muhammad, et al.
Pubblicazione: (2025)
di: Mujahid, Zain Muhammad, et al.
Pubblicazione: (2025)
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
di: Islam, Sekh Mainul, et al.
Pubblicazione: (2025)
di: Islam, Sekh Mainul, et al.
Pubblicazione: (2025)
1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
di: Hakimi, Ahmad Dawar, et al.
Pubblicazione: (2026)
di: Hakimi, Ahmad Dawar, et al.
Pubblicazione: (2026)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
di: Lee, Donggyu, et al.
Pubblicazione: (2025)
di: Lee, Donggyu, et al.
Pubblicazione: (2025)
Mentor-KD: Making Small Language Models Better Multi-step Reasoners
di: Lee, Hojae, et al.
Pubblicazione: (2024)
di: Lee, Hojae, et al.
Pubblicazione: (2024)
Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
di: Choi, Junhyuk, et al.
Pubblicazione: (2026)
di: Choi, Junhyuk, et al.
Pubblicazione: (2026)
Instruction Tuning with Human Curriculum
di: Lee, Bruce W., et al.
Pubblicazione: (2023)
di: Lee, Bruce W., et al.
Pubblicazione: (2023)
Modeling Public Perceptions of Science in Media
di: Pei, Jiaxin, et al.
Pubblicazione: (2025)
di: Pei, Jiaxin, et al.
Pubblicazione: (2025)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
di: Kim, Yungi, et al.
Pubblicazione: (2024)
di: Kim, Yungi, et al.
Pubblicazione: (2024)
What Happens to a Dataset Transformed by a Projection-based Concept Removal Method?
di: Johansson, Richard
Pubblicazione: (2024)
di: Johansson, Richard
Pubblicazione: (2024)
DSG-KD: Knowledge Distillation from Domain-Specific to General Language Models
di: Cho, Sangyeon, et al.
Pubblicazione: (2024)
di: Cho, Sangyeon, et al.
Pubblicazione: (2024)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
di: Sun, Jingyi, et al.
Pubblicazione: (2026)
di: Sun, Jingyi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Reality Check on Context Utilisation for Retrieval-Augmented Generation
di: Hagström, Lovisa, et al.
Pubblicazione: (2024) -
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
di: Yu, Haeun, et al.
Pubblicazione: (2024) -
Language Model Re-rankers are Fooled by Lexical Similarities
di: Hagström, Lovisa, et al.
Pubblicazione: (2025) -
UniKnow: A Unified Framework for Reliable Language Model Behavior across Parametric and External Knowledge
di: Kim, Youna, et al.
Pubblicazione: (2025) -
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
di: Pauli, Amalie Brogaard, et al.
Pubblicazione: (2024)