SHAP-Based Fraud Explanation with Small Language Models via Knowledge Distillation
Tri Thuc Tran Le and Phu-Nguyen Le
In EAI FPT International Conference on Intelligent Systems and Advanced Technologies (EAI FISAT), 2026
Accepted
Automated fraud detection systems achieve high predictive accuracy but produce decisions that are difficult to explain to customers and regulators. Large language models (LLMs) can generate high-quality, structured explanations grounded in SHAP attributions, yet their direct deployment in financial systems is impeded by data-privacy constraints, latency requirements, and the need for reproducible on-premise inference. We address this gap by applying knowledge distillation to SHAP-based fraud explanation: a GPT-5.5 teacher generates a corpus of 5,000 four-step chain-of-thought explanations over IEEE-CIS fraud transactions; two small language models, Phi-4-mini-instruct and Qwen3-1.7B, are then fine-tuned on this corpus via Low-Rank Adaptation (LoRA). We introduce Faithfulness, a domain-specific metric measuring SHAP feature citation accuracy. Both distilled students achieve near-perfect faithfulness (at least 0.978), with decision-consistency gains of +0.196 and +0.148 over their zero-shot counterparts. We further show that the keyword rule commonly used to recover a decision label from generated text is systematically biased: it understates the teacher’s own label agreement by 25.0 points. We give a context-aware extractor validated label-free against the protocol-stated decisions, which it recovers on 99.8% of teacher outputs. The distilled models require no LLM API access at inference time, enabling privacy-preserving, auditable financial explanations at production scale.