Predicting Employee Turnover through Hr Analytics
Predicting Employee Turnover through Hr Analytics
Abstract
Bu tez, çalışan devamsızlık modellerini nedensel açıklamalar değil, karar verme araçları olarak ele almaktadır. Buradaki temel sorun, kaynakların kısıtlı olduğu ve yanlış sınıflandırmanın maliyetinin asimetrik olduğu durumlarda, tahmin edilen risk puanlarına dayalı olarak çalışanları elde tutma eylemlerinin ölçülebilmesidir. İki adet halka açık İK veri seti incelenmiştir: HR v13 (310 çalışan, devamsızlık oranı %33,2) ve IBM Devamsızlık veri seti (1470 çalışan, devamsızlık oranı %16,1). Tüm analizler, tamamen çalıştırılabilir ön işleme ile Python'da gerçekleştirilmiştir. HR v13 durumunda, saatlik ücret ve maaş tablosu orta noktalarının aylık karşılık gelen değerlere dönüştürülmesiyle konsolide edilmiş aylık gelir değişkeni oluşturulmuştur. L2 düzenlemeli lojistik regresyon ve rastgele orman olmak üzere iki sınıflandırıcı karşılaştırılmıştır. Açık erişim kuralları altında, modeller AUROC, Brier puanı ve karar tabanlı metrikler kullanılarak karşılaştırılmıştır. Güvenilirlik, güvenilirlik diyagramı ve izotonik kalibrasyon yoluyla analiz edildi. 3:1'lik bir maliyet-kâr oranı, her iki veri setinde de beklenen maliyeti geleneksel 0,50'lik eşik değerinden daha düşük bir seviyeye indiren t = 0,25'lik bir çalışma noktası sağladı. Kalibrasyon ile yorumlanabilirlik artırıldı, ancak sabit eşikte beklenen maliyette bir azalma olmadı. Segment düzeyindeki temasların analizi, görev süresi, ücret ve fazla mesai kategorilerinde büyük farklılıklar gösterdi. IBM'deki performans kayıpları, HR v13'te yapılan yalın özellik setlerinin sonucuydu. Sabit erişim politikalarında, lojistik regresyon rastgele ormana eşit veya biraz daha iyi performans gösterdi. Sonuç olarak, sonuçlar, değerlendirmenin doğruluğun kendisinden ziyade karar verme maliyeti, erişim düzeyi, kalibrasyon ve segment raporuyla ilişkili olduğu durumlarda işten ayrılma modellemesinin en yararlı olduğunu göstermektedir.
This thesis considers employee attrition models to be tools of decisions and not causal explanations. The main issue here is that the retention actions based on the predicted risk scores can be measured when resources are scarce and the cost of misclassification is asymmetric. Two publicly available datasets of HR were examined, namely, HR v13 (310 employees, attrition rate 33.2%) and the IBM Attrition dataset (1,470 employees, attrition rate 16.1%). All the analysis was performed in Python with entirely executable preprocessing. In the case of HR v13, a consolidated monthly income variable was made through the transformation of hourly pay and salary-grid midpoints into monthly corresponding counterparts. Two classifiers were compared, which included L2 regularization (ridge penalty) (L2) regularized logistic regression and random forest. Under explicit outreach rules, models were compared by using Area Under the Receiver Operating Characteristics Curve (AUROC), Brier score and decision-based metrics. Reliability was analyzed by way of reliability diagram and the isotonic calibration. A cost-to-profit ratio of 3:1 gave a working point of t = 0.25 that cut expected cost in both datasets lower than the traditional cut of 0.50. Interpretability was enhanced with calibration, but at the fixed threshold, there was no decrease in expected cost. Analysis of contacts at the segment level showed large differentiation by tenure, pay, and overtime categories. The performance losses in IBM were the result of lean feature sets that were made in HR v13. In the fixed outreach policies, logistic regression was equal or slightly better than random forest. All in all, the results do indicate that attrition modelling proves most useful where the evaluation is associated with the cost of decision making, the reach level, calibration and segment report as opposed to the accuracy itself.
This thesis considers employee attrition models to be tools of decisions and not causal explanations. The main issue here is that the retention actions based on the predicted risk scores can be measured when resources are scarce and the cost of misclassification is asymmetric. Two publicly available datasets of HR were examined, namely, HR v13 (310 employees, attrition rate 33.2%) and the IBM Attrition dataset (1,470 employees, attrition rate 16.1%). All the analysis was performed in Python with entirely executable preprocessing. In the case of HR v13, a consolidated monthly income variable was made through the transformation of hourly pay and salary-grid midpoints into monthly corresponding counterparts. Two classifiers were compared, which included L2 regularization (ridge penalty) (L2) regularized logistic regression and random forest. Under explicit outreach rules, models were compared by using Area Under the Receiver Operating Characteristics Curve (AUROC), Brier score and decision-based metrics. Reliability was analyzed by way of reliability diagram and the isotonic calibration. A cost-to-profit ratio of 3:1 gave a working point of t = 0.25 that cut expected cost in both datasets lower than the traditional cut of 0.50. Interpretability was enhanced with calibration, but at the fixed threshold, there was no decrease in expected cost. Analysis of contacts at the segment level showed large differentiation by tenure, pay, and overtime categories. The performance losses in IBM were the result of lean feature sets that were made in HR v13. In the fixed outreach policies, logistic regression was equal or slightly better than random forest. All in all, the results do indicate that attrition modelling proves most useful where the evaluation is associated with the cost of decision making, the reach level, calibration and segment report as opposed to the accuracy itself.
Description
Keywords
Yönetim Bilişim Sistemleri, Management Information Systems
Turkish CoHE Thesis Center URL
Fields of Science
Citation
WoS Q
Scopus Q
Source
Volume
Issue
Start Page
End Page
132
