Leakage-Aware LLM Augmentation for Attrition Prediction: A DecisionCentric Evaluation

Authors

  • Liao Weiquan
  • Jiayangmei Xu
  • Ekaterina A. Panova

DOI:

https://doi.org/10.65343/aiis.v2i2.141

Keywords:

employee attrition, HR analytics, large language models, actor-critic framework, synthetic data augmentation, leakage-aware evaluation, class imbalance

Abstract

Employee attrition is a high-cost, asymmetric decision problem plagued by data imbalance. To address this, we propose a leakage-aware, dual-network augmentation framework that integrates the semantic reasoning of Large Language Models (LLMs) with the distributional rigorousness of GAN discriminators. Specifically, we employ an autoregressive transformer (GReaT) as a generative "actor" to synthesize diverse minority samples, coupled with a post-hoc "critic" derived from a conditional GAN to strictly filter low-fidelity outliers. This actor-critic inspired loop ensures that synthetic records broaden coverage without introducing noise. We evaluate this pipeline using a rigorous protocol where all generation and filtering occur strictly within training folds to prevent leakage. Experiments on HR datasets show that moderate, critic-guided oversampling yields significant recall gains (e.g., raising recall from 45% to 53%) and improved F₁ scores compared to baselines, while maintaining calibration and discrimination (stable ROC-AUC). Furthermore, the approach preserves interpretability (SHAP) and fairness (low TPR gaps), offering HR practitioners a robust, decision-centric tool to identify at-risk employees without sacrificing transparency.

Downloads

Published

2026-09-23

How to Cite

Liao Weiquan, Jiayangmei Xu, & Ekaterina A. Panova. (2026). Leakage-Aware LLM Augmentation for Attrition Prediction: A DecisionCentric Evaluation. Artificial Intelligence and Internet Studies, 2(2), pp.1–14. https://doi.org/10.65343/aiis.v2i2.141