Selective Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Paes, Lucas Monteiro, Wei, Dennis, Calmon, Flavio P. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Algorithmic Arbitrariness in Content Moderation
by: Gomez, Juan Felipe, et al.
Published: (2024)
by: Gomez, Juan Felipe, et al.
Published: (2024)
Multi-Group Fairness Evaluation via Conditional Value-at-Risk Testing
by: Paes, Lucas Monteiro, et al.
Published: (2023)
by: Paes, Lucas Monteiro, et al.
Published: (2023)
Theoretical Limits of Language Model Alignment
by: Paes, Lucas Monteiro, et al.
Published: (2026)
by: Paes, Lucas Monteiro, et al.
Published: (2026)
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025)
by: Paes, Lucas Monteiro, et al.
Published: (2025)
AI Alignment at Your Discretion
by: Buyl, Maarten, et al.
Published: (2025)
by: Buyl, Maarten, et al.
Published: (2025)
DetoxLLM: A Framework for Detoxification with Explanations
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024)
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024)
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
by: Dumitran, Adrian-Marius, et al.
Published: (2025)
by: Dumitran, Adrian-Marius, et al.
Published: (2025)
Aleatoric and Epistemic Discrimination: Fundamental Limits of Fairness Interventions
by: Wang, Hao, et al.
Published: (2023)
by: Wang, Hao, et al.
Published: (2023)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
by: Hamman, Faisal, et al.
Published: (2025)
by: Hamman, Faisal, et al.
Published: (2025)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
Strategic Demonstration Selection for Improved Fairness in LLM In-Context Learning
by: Hu, Jingyu, et al.
Published: (2024)
by: Hu, Jingyu, et al.
Published: (2024)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
by: Tsur, Dor, et al.
Published: (2025)
by: Tsur, Dor, et al.
Published: (2025)
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
by: Wilming, Rick, et al.
Published: (2024)
by: Wilming, Rick, et al.
Published: (2024)
Properties and Challenges of LLM-Generated Explanations
by: Kunz, Jenny, et al.
Published: (2024)
by: Kunz, Jenny, et al.
Published: (2024)
The Geometric Price of Discrete Logic: Context-driven Manifold Dynamics of Number Representations
by: Zhang, Long, et al.
Published: (2026)
by: Zhang, Long, et al.
Published: (2026)
Towards Explainable Evaluation Metrics for Machine Translation
by: Leiter, Christoph, et al.
Published: (2023)
by: Leiter, Christoph, et al.
Published: (2023)
Customize Multi-modal RAI Guardrails with Precedent-based predictions
by: Yang, Cheng-Fu, et al.
Published: (2025)
by: Yang, Cheng-Fu, et al.
Published: (2025)
ICX360: In-Context eXplainability 360 Toolkit
by: Wei, Dennis, et al.
Published: (2025)
by: Wei, Dennis, et al.
Published: (2025)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
by: Chehbouni, Khaoula, et al.
Published: (2024)
by: Chehbouni, Khaoula, et al.
Published: (2024)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025)
by: Bhalla, Usha, et al.
Published: (2025)
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
by: Wei, Shou'ang, et al.
Published: (2025)
by: Wei, Shou'ang, et al.
Published: (2025)
AppellateGen: A Benchmark for Appellate Legal Judgment Generation
by: Yang, Hongkun, et al.
Published: (2026)
by: Yang, Hongkun, et al.
Published: (2026)
SUV: Scalable Large Language Model Copyright Compliance with Regularized Selective Unlearning
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
by: Taghanaki, Saeid Asgari, et al.
Published: (2025)
by: Taghanaki, Saeid Asgari, et al.
Published: (2025)
The Moral Foundations Reddit Corpus
by: Trager, Jackson, et al.
Published: (2022)
by: Trager, Jackson, et al.
Published: (2022)
Protected group bias and stereotypes in Large Language Models
by: Kotek, Hadas, et al.
Published: (2024)
by: Kotek, Hadas, et al.
Published: (2024)
Diagnosing Hate Speech Classification: Where Do Humans and Machines Disagree, and Why?
by: Yang, Xilin
Published: (2024)
by: Yang, Xilin
Published: (2024)
Exploring the Potential of the Large Language Models (LLMs) in Identifying Misleading News Headlines
by: Rony, Md Main Uddin, et al.
Published: (2024)
by: Rony, Md Main Uddin, et al.
Published: (2024)
Unintended Impacts of LLM Alignment on Global Representation
by: Ryan, Michael J., et al.
Published: (2024)
by: Ryan, Michael J., et al.
Published: (2024)
Student Answer Forecasting: Transformer-Driven Answer Choice Prediction for Language Learning
by: Gado, Elena Grazia, et al.
Published: (2024)
by: Gado, Elena Grazia, et al.
Published: (2024)
Test Case-Informed Knowledge Tracing for Open-ended Coding Tasks
by: Duan, Zhangqi, et al.
Published: (2024)
by: Duan, Zhangqi, et al.
Published: (2024)
A Time-Aware Approach to Early Detection of Anorexia: UNSL at eRisk 2024
by: Thompson, Horacio, et al.
Published: (2024)
by: Thompson, Horacio, et al.
Published: (2024)
Identifying Climate Targets in National Laws and Policies using Machine Learning
by: Juhasz, Matyas, et al.
Published: (2024)
by: Juhasz, Matyas, et al.
Published: (2024)
Interpreting Latent Student Knowledge Representations in Programming Assignments
by: Fernandez, Nigel, et al.
Published: (2024)
by: Fernandez, Nigel, et al.
Published: (2024)
From Imitation to Introspection: Probing Self-Consciousness in Language Models
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
by: Arzaghi, Mina, et al.
Published: (2024)
by: Arzaghi, Mina, et al.
Published: (2024)
Balancing the Scales: Reinforcement Learning for Fair Classification
by: Eshuijs, Leon, et al.
Published: (2024)
by: Eshuijs, Leon, et al.
Published: (2024)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
by: Fleisig, Eve, et al.
Published: (2024)
by: Fleisig, Eve, et al.
Published: (2024)
ClaimVer: Explainable Claim-Level Verification and Evidence Attribution of Text Through Knowledge Graphs
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
Convolutional Neural Networks can achieve binary bail judgement classification
by: Barman, Amit, et al.
Published: (2024)
by: Barman, Amit, et al.
Published: (2024)
Similar Items
-
Algorithmic Arbitrariness in Content Moderation
by: Gomez, Juan Felipe, et al.
Published: (2024) -
Multi-Group Fairness Evaluation via Conditional Value-at-Risk Testing
by: Paes, Lucas Monteiro, et al.
Published: (2023) -
Theoretical Limits of Language Model Alignment
by: Paes, Lucas Monteiro, et al.
Published: (2026) -
DSO: Direct Steering Optimization for Bias Mitigation
by: Paes, Lucas Monteiro, et al.
Published: (2025) -
AI Alignment at Your Discretion
by: Buyl, Maarten, et al.
Published: (2025)