Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Hamman, Faisal, Dissanayake, Pasan, Fu, Yanjun, Dutta, Sanghamitra |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
by: Dissanayake, Pasan, et al.
Published: (2025)
by: Dissanayake, Pasan, et al.
Published: (2025)
Counterfactual Explanations for Model Ensembles Using Entropic Risk Measures
by: Noorani, Erfaun, et al.
Published: (2025)
by: Noorani, Erfaun, et al.
Published: (2025)
Quantifying Prediction Consistency Under Fine-Tuning Multiplicity in Tabular LLMs
by: Hamman, Faisal, et al.
Published: (2024)
by: Hamman, Faisal, et al.
Published: (2024)
Model Reconstruction Using Counterfactual Explanations: A Perspective From Polytope Theory
by: Dissanayake, Pasan, et al.
Published: (2024)
by: Dissanayake, Pasan, et al.
Published: (2024)
Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees
by: Hamman, Faisal, et al.
Published: (2023)
by: Hamman, Faisal, et al.
Published: (2023)
Towards Formalizing Spuriousness of Biased Datasets Using Partial Information Decomposition
by: Halder, Barproda, et al.
Published: (2024)
by: Halder, Barproda, et al.
Published: (2024)
Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
by: Hamman, Faisal, et al.
Published: (2025)
by: Hamman, Faisal, et al.
Published: (2025)
T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction Tuning
by: Fu, Yanjun, et al.
Published: (2025)
by: Fu, Yanjun, et al.
Published: (2025)
Demystifying Local and Global Fairness Trade-offs in Federated Learning Using Partial Information Decomposition
by: Hamman, Faisal, et al.
Published: (2023)
by: Hamman, Faisal, et al.
Published: (2023)
A Unified View of Group Fairness Tradeoffs Using Partial Information Decomposition
by: Hamman, Faisal, et al.
Published: (2024)
by: Hamman, Faisal, et al.
Published: (2024)
Quantifying Knowledge Distillation Using Partial Information Decomposition
by: Dissanayake, Pasan, et al.
Published: (2024)
by: Dissanayake, Pasan, et al.
Published: (2024)
Learning Invariant Graph Representations Through Redundant Information
by: Halder, Barproda, et al.
Published: (2025)
by: Halder, Barproda, et al.
Published: (2025)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation
by: Egea, David, et al.
Published: (2025)
by: Egea, David, et al.
Published: (2025)
KTCF: Actionable Recourse in Knowledge Tracing via Counterfactual Explanations for Education
by: Kim, Woojin, et al.
Published: (2026)
by: Kim, Woojin, et al.
Published: (2026)
Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations
by: Ezzeddine, Fatima, et al.
Published: (2024)
by: Ezzeddine, Fatima, et al.
Published: (2024)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
by: Ghosh, Rajarshi, et al.
Published: (2025)
by: Ghosh, Rajarshi, et al.
Published: (2025)
Counterfactual Explanations for Hypergraph Neural Networks
by: Veglianti, Fabiano, et al.
Published: (2026)
by: Veglianti, Fabiano, et al.
Published: (2026)
Procedural Fairness via Group Counterfactual Explanation
by: Popoola, Gideon, et al.
Published: (2026)
by: Popoola, Gideon, et al.
Published: (2026)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
by: Wilming, Rick, et al.
Published: (2024)
by: Wilming, Rick, et al.
Published: (2024)
Demystifying the Accuracy-Interpretability Trade-Off: A Case Study of Inferring Ratings from Reviews
by: Atrey, Pranjal, et al.
Published: (2025)
by: Atrey, Pranjal, et al.
Published: (2025)
Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models
by: Feng, Duanyu, et al.
Published: (2023)
by: Feng, Duanyu, et al.
Published: (2023)
Language Representation Favored Zero-Shot Cross-Domain Cognitive Diagnosis
by: Liu, Shuo, et al.
Published: (2025)
by: Liu, Shuo, et al.
Published: (2025)
TABCF: Counterfactual Explanations for Tabular Data Using a Transformer-Based VAE
by: Panagiotou, Emmanouil, et al.
Published: (2024)
by: Panagiotou, Emmanouil, et al.
Published: (2024)
BANER: Boundary-Aware LLMs for Few-Shot Named Entity Recognition
by: Guo, Quanjiang, et al.
Published: (2024)
by: Guo, Quanjiang, et al.
Published: (2024)
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes
by: Gallegos, Isabel O., et al.
Published: (2024)
by: Gallegos, Isabel O., et al.
Published: (2024)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
by: Lin, Yujie, et al.
Published: (2026)
by: Lin, Yujie, et al.
Published: (2026)
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models
by: Era, Jalisha Jashim, et al.
Published: (2025)
by: Era, Jalisha Jashim, et al.
Published: (2025)
Wikipedia in the Era of LLMs: Evolution and Risks
by: Huang, Siming, et al.
Published: (2025)
by: Huang, Siming, et al.
Published: (2025)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Properties and Challenges of LLM-Generated Explanations
by: Kunz, Jenny, et al.
Published: (2024)
by: Kunz, Jenny, et al.
Published: (2024)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
by: Han, Pengrui, et al.
Published: (2025)
by: Han, Pengrui, et al.
Published: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025)
by: Naderi, Nariman, et al.
Published: (2025)
How Far Are We From AGI: Are LLMs All We Need?
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
Similar Items
-
TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
by: Dissanayake, Pasan, et al.
Published: (2025) -
Counterfactual Explanations for Model Ensembles Using Entropic Risk Measures
by: Noorani, Erfaun, et al.
Published: (2025) -
Quantifying Prediction Consistency Under Fine-Tuning Multiplicity in Tabular LLMs
by: Hamman, Faisal, et al.
Published: (2024) -
Model Reconstruction Using Counterfactual Explanations: A Perspective From Polytope Theory
by: Dissanayake, Pasan, et al.
Published: (2024) -
Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees
by: Hamman, Faisal, et al.
Published: (2023)