Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Lewis, Ashley, White, Michael, Liu, Jing, Koike-Akino, Toshiaki, Parsons, Kieran, Wang, Ye |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Smoothed Embeddings for Robust Language Models
por: Hase, Ryo, et al.
Publicado: (2025)
por: Hase, Ryo, et al.
Publicado: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
por: Koike-Akino, Toshiaki, et al.
Publicado: (2026)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2026)
Directional Embedding Smoothing for Robust Vision Language Models
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Efficient Differentially Private Fine-Tuning of Diffusion Models
por: Liu, Jing, et al.
Publicado: (2024)
por: Liu, Jing, et al.
Publicado: (2024)
Variational Randomized Smoothing for Sample-Wise Adversarial Robustness
por: Hase, Ryo, et al.
Publicado: (2024)
por: Hase, Ryo, et al.
Publicado: (2024)
AutoHLS: Learning to Accelerate Design Space Exploration for HLS Designs
por: Ahmed, Md Rubel, et al.
Publicado: (2024)
por: Ahmed, Md Rubel, et al.
Publicado: (2024)
Why Does Differential Privacy with Large Epsilon Defend Against Practical Membership Inference Attacks?
por: Lowy, Andrew, et al.
Publicado: (2024)
por: Lowy, Andrew, et al.
Publicado: (2024)
Exploring User-level Gradient Inversion with a Diffusion Prior
por: Li, Zhuohang, et al.
Publicado: (2024)
por: Li, Zhuohang, et al.
Publicado: (2024)
Analyzing Inference Privacy Risks Through Gradients in Machine Learning
por: Li, Zhuohang, et al.
Publicado: (2024)
por: Li, Zhuohang, et al.
Publicado: (2024)
LatentLLM: Attention-Aware Joint Tensor Compression
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
Quantum Implicit Neural Compression
por: Fujihashi, Takuya, et al.
Publicado: (2024)
por: Fujihashi, Takuya, et al.
Publicado: (2024)
Random Channel Ablation for Robust Hand Gesture Classification with Multimodal Biosignals
por: Bimbraw, Keshav, et al.
Publicado: (2024)
por: Bimbraw, Keshav, et al.
Publicado: (2024)
Quantum Diffusion Models for Few-Shot Learning
por: Wang, Ruhan, et al.
Publicado: (2024)
por: Wang, Ruhan, et al.
Publicado: (2024)
GPT Sonograpy: Hand Gesture Decoding from Forearm Ultrasound Images via VLM
por: Bimbraw, Keshav, et al.
Publicado: (2024)
por: Bimbraw, Keshav, et al.
Publicado: (2024)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
por: Khattar, Vanshaj, et al.
Publicado: (2026)
por: Khattar, Vanshaj, et al.
Publicado: (2026)
Geo-ADAPT-VQE: Quantum Information Metric-Aware Circuit Optimization for Quantum Chemistry
por: Sohail, Mohammad Aamir, et al.
Publicado: (2026)
por: Sohail, Mohammad Aamir, et al.
Publicado: (2026)
AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
por: Liu, Jing, et al.
Publicado: (2025)
por: Liu, Jing, et al.
Publicado: (2025)
Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
por: Nguyen, Hieu, et al.
Publicado: (2025)
por: Nguyen, Hieu, et al.
Publicado: (2025)
Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
por: Sun, Yiqun, et al.
Publicado: (2026)
por: Sun, Yiqun, et al.
Publicado: (2026)
Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning Distillation
por: Zhang, Yong, et al.
Publicado: (2025)
por: Zhang, Yong, et al.
Publicado: (2025)
On the Generalization vs Fidelity Paradox in Knowledge Distillation
por: Ramesh, Suhas Kamasetty, et al.
Publicado: (2025)
por: Ramesh, Suhas Kamasetty, et al.
Publicado: (2025)
Combining LLMs and Knowledge Graphs to Reduce Hallucinations in Question Answering
por: Pusch, Larissa, et al.
Publicado: (2024)
por: Pusch, Larissa, et al.
Publicado: (2024)
Counterfactual Cultural Cues Reduce Medical QA Accuracy in LLMs: Identifier vs Context Effects
por: Rezaei, Amirhossein Haji Mohammad, et al.
Publicado: (2026)
por: Rezaei, Amirhossein Haji Mohammad, et al.
Publicado: (2026)
VISTA: Verification In Sequential Turn-based Assessment
por: Lewis, Ashley, et al.
Publicado: (2025)
por: Lewis, Ashley, et al.
Publicado: (2025)
TuneComp: Joint Fine-tuning and Compression for Large Foundation Models
por: Chen, Xiangyu, et al.
Publicado: (2025)
por: Chen, Xiangyu, et al.
Publicado: (2025)
Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
por: Rashid, Md Rafi Ur, et al.
Publicado: (2024)
por: Rashid, Md Rafi Ur, et al.
Publicado: (2024)
Range Image-Based Implicit Neural Compression for LiDAR Point Clouds
por: Kuwabara, Akihiro, et al.
Publicado: (2025)
por: Kuwabara, Akihiro, et al.
Publicado: (2025)
Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey
por: Agrawal, Garima, et al.
Publicado: (2023)
por: Agrawal, Garima, et al.
Publicado: (2023)
Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching
por: Li, Mingzhe, et al.
Publicado: (2025)
por: Li, Mingzhe, et al.
Publicado: (2025)
Small Wins Big: Comparing Large Language Models and Domain Fine-Tuned Models for Sarcasm Detection in Code-Mixed Hinglish Text
por: Majumder, Bitan, et al.
Publicado: (2026)
por: Majumder, Bitan, et al.
Publicado: (2026)
Knowledge Distillation with Training Wheels
por: Liu, Guanlin, et al.
Publicado: (2025)
por: Liu, Guanlin, et al.
Publicado: (2025)
Mini-Giants: "Small" Language Models and Open Source Win-Win
por: Zhou, Zhengping, et al.
Publicado: (2023)
por: Zhou, Zhengping, et al.
Publicado: (2023)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
por: Barron, Joshua, et al.
Publicado: (2025)
por: Barron, Joshua, et al.
Publicado: (2025)
Small Agent Can Also Rock! Empowering Small Language Models as Hallucination Detector
por: Cheng, Xiaoxue, et al.
Publicado: (2024)
por: Cheng, Xiaoxue, et al.
Publicado: (2024)
Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks
por: Tong, Chaodong, et al.
Publicado: (2025)
por: Tong, Chaodong, et al.
Publicado: (2025)
Small Updates, Big Doubts: Does Parameter-Efficient Fine-tuning Enhance Hallucination Detection ?
por: Hu, Xu, et al.
Publicado: (2026)
por: Hu, Xu, et al.
Publicado: (2026)
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
por: Datta, Joyeeta, et al.
Publicado: (2025)
por: Datta, Joyeeta, et al.
Publicado: (2025)
Training Language Models to Win Debates with Self-Play Improves Judge Accuracy
por: Arnesen, Samuel, et al.
Publicado: (2024)
por: Arnesen, Samuel, et al.
Publicado: (2024)
Ejemplares similares
-
Smoothed Embeddings for Robust Language Models
por: Hase, Ryo, et al.
Publicado: (2025) -
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026) -
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025) -
TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
por: Koike-Akino, Toshiaki, et al.
Publicado: (2026) -
Directional Embedding Smoothing for Robust Vision Language Models
por: Wang, Ye, et al.
Publicado: (2026)