Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sun, Jingyi, Atanasova, Pepa, Choudhury, Sagnik Ray, Islam, Sekh Mainul, Augenstein, Isabelle |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
par: Islam, Sekh Mainul, et autres
Publié: (2025)
par: Islam, Sekh Mainul, et autres
Publié: (2025)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
par: Sun, Jingyi, et autres
Publié: (2024)
par: Sun, Jingyi, et autres
Publié: (2024)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
par: Yu, Haeun, et autres
Publié: (2024)
par: Yu, Haeun, et autres
Publié: (2024)
Graph-Guided Textual Explanation Generation Framework
par: Yuan, Shuzhou, et autres
Publié: (2024)
par: Yuan, Shuzhou, et autres
Publié: (2024)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
par: Sun, Jingyi, et autres
Publié: (2026)
par: Sun, Jingyi, et autres
Publié: (2026)
A Reality Check on Context Utilisation for Retrieval-Augmented Generation
par: Hagström, Lovisa, et autres
Publié: (2024)
par: Hagström, Lovisa, et autres
Publié: (2024)
Self-Critique and Refinement for Faithful Natural Language Explanations
par: Wang, Yingming, et autres
Publié: (2025)
par: Wang, Yingming, et autres
Publié: (2025)
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
par: Islam, Sekh Mainul, et autres
Publié: (2025)
par: Islam, Sekh Mainul, et autres
Publié: (2025)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
par: Marjanović, Sara Vera, et autres
Publié: (2024)
par: Marjanović, Sara Vera, et autres
Publié: (2024)
Understanding the Interplay between LLMs' Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025
par: Augenstein, Isabelle
Publié: (2026)
par: Augenstein, Isabelle
Publié: (2026)
CUB: Benchmarking Context Utilisation Techniques for Language Models
par: Hagström, Lovisa, et autres
Publié: (2025)
par: Hagström, Lovisa, et autres
Publié: (2025)
S-GRADES -- Studying Generalization of Student Response Assessments in Diverse Evaluative Settings
par: Seuti, Tasfia, et autres
Publié: (2026)
par: Seuti, Tasfia, et autres
Publié: (2026)
Investigating the Impact of Model Instability on Explanations and Uncertainty
par: Marjanović, Sara Vera, et autres
Publié: (2024)
par: Marjanović, Sara Vera, et autres
Publié: (2024)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
par: Wang, Qianli, et autres
Publié: (2026)
par: Wang, Qianli, et autres
Publié: (2026)
Explaining Sources of Uncertainty in Automated Fact-Checking
par: Sun, Jingyi, et autres
Publié: (2025)
par: Sun, Jingyi, et autres
Publié: (2025)
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
par: Arora, Arnav, et autres
Publié: (2022)
par: Arora, Arnav, et autres
Publié: (2022)
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
par: Pauli, Amalie Brogaard, et autres
Publié: (2024)
par: Pauli, Amalie Brogaard, et autres
Publié: (2024)
Aggregating Soft Labels from Crowd Annotations Improves Uncertainty Estimation Under Distribution Shift
par: Wright, Dustin, et autres
Publié: (2022)
par: Wright, Dustin, et autres
Publié: (2022)
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics
par: Pauli, Amalie Brogaard, et autres
Publié: (2025)
par: Pauli, Amalie Brogaard, et autres
Publié: (2025)
Social Bias Probing: Fairness Benchmarking for Language Models
par: Manerba, Marta Marchiori, et autres
Publié: (2023)
par: Manerba, Marta Marchiori, et autres
Publié: (2023)
ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
par: Anik, Anirban Saha, et autres
Publié: (2025)
par: Anik, Anirban Saha, et autres
Publié: (2025)
Can Transformers Learn $n$-gram Language Models?
par: Svete, Anej, et autres
Publié: (2024)
par: Svete, Anej, et autres
Publié: (2024)
Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting
par: Nemkova, Apollinaire Poli, et autres
Publié: (2025)
par: Nemkova, Apollinaire Poli, et autres
Publié: (2025)
Expanding Computation Spaces of LLMs at Inference Time
par: Jang, Yoonna, et autres
Publié: (2025)
par: Jang, Yoonna, et autres
Publié: (2025)
Quantifying Gender Biases Towards Politicians on Reddit
par: Marjanovic, Sara, et autres
Publié: (2021)
par: Marjanovic, Sara, et autres
Publié: (2021)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
par: Arakelyan, Erik, et autres
Publié: (2024)
par: Arakelyan, Erik, et autres
Publié: (2024)
Claim Verification in the Age of Large Language Models: A Survey
par: Dmonte, Alphaeus, et autres
Publié: (2024)
par: Dmonte, Alphaeus, et autres
Publié: (2024)
A Detailed Factor Analysis for the Political Compass Test: Navigating Ideologies of Large Language Models
par: Kamal, Sadia, et autres
Publié: (2025)
par: Kamal, Sadia, et autres
Publié: (2025)
Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations
par: Cao, Yong, et autres
Publié: (2025)
par: Cao, Yong, et autres
Publié: (2025)
Unstructured Evidence Attribution for Long Context Query Focused Summarization
par: Wright, Dustin, et autres
Publié: (2025)
par: Wright, Dustin, et autres
Publié: (2025)
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
par: Jahan, Rafid Ishrak, et autres
Publié: (2026)
par: Jahan, Rafid Ishrak, et autres
Publié: (2026)
Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
par: Warren, Greta, et autres
Publié: (2025)
par: Warren, Greta, et autres
Publié: (2025)
Revealing Fine-Grained Values and Opinions in Large Language Models
par: Wright, Dustin, et autres
Publié: (2024)
par: Wright, Dustin, et autres
Publié: (2024)
Modeling Public Perceptions of Science in Media
par: Pei, Jiaxin, et autres
Publié: (2025)
par: Pei, Jiaxin, et autres
Publié: (2025)
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
par: Ghazaryan, Gayane, et autres
Publié: (2024)
par: Ghazaryan, Gayane, et autres
Publié: (2024)
Towards Automated Situation Awareness: A RAG-Based Framework for Peacebuilding Reports
par: Nemkova, Poli A., et autres
Publié: (2025)
par: Nemkova, Poli A., et autres
Publié: (2025)
Large Language Models as Evaluators for Recommendation Explanations
par: Zhang, Xiaoyu, et autres
Publié: (2024)
par: Zhang, Xiaoyu, et autres
Publié: (2024)
What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages
par: Borenstein, Nadav, et autres
Publié: (2024)
par: Borenstein, Nadav, et autres
Publié: (2024)
Understanding Fine-grained Distortions in Reports of Scientific Findings
par: Wührl, Amelie, et autres
Publié: (2024)
par: Wührl, Amelie, et autres
Publié: (2024)
Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions
par: Kaffee, Lucie-Aimée, et autres
Publié: (2023)
par: Kaffee, Lucie-Aimée, et autres
Publié: (2023)
Documents similaires
-
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
par: Islam, Sekh Mainul, et autres
Publié: (2025) -
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
par: Sun, Jingyi, et autres
Publié: (2024) -
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
par: Yu, Haeun, et autres
Publié: (2024) -
Graph-Guided Textual Explanation Generation Framework
par: Yuan, Shuzhou, et autres
Publié: (2024) -
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
par: Sun, Jingyi, et autres
Publié: (2026)