A Development of Predictive Model for Customer Retention using Machine Learning Techniques: A Case Study of a Parcel Service Franchise

Main Article Content

๋Jamornkul Laokietkul
Wuttinun Boonpho

Abstract

   This research aims to 1) develop a predictive model using machine learning techniques to forecast customer retention, 2) compare the performance of various machine learning models, and 3) analyze key determinants influencing customer re-usage behavior in a parcel delivery franchise. The research methodology was conducted based on the Cross-Industry Standard Process for Data Mining (CRISP-DM) The study utilized empirical transaction data from August 2025 to January 2026. Features extraction was based on the Recency, Frequency, and Monetary (RFM) analysis framework, supplemented by the Cash on Delivery (COD) service. The Tabular Generative Adversarial Networks (CTGAN) was employed to synthesize data and mitigate class imbalance, enhancing model performance. Subsequently, four machine learning algorithms as Logistic Regression, Random Forest, XGBoost, and Support Vector Machine (SVM) were evaluated. The findings successfully addressed the research objectives as follows: 1) The predictive model was successfully developed, effectively mitigating the learning bias caused by data imbalance. 2) The comparative analysis demonstrated that the XGBoost model achieved the highest predictive performance, yielding an Area Under the ROC Curve (AUC-ROC) of 0.8462, thereby proving to be the most suitable algorithm for predicting customer return. 3) The feature importance analysis indicated that the Recency of service usage was the most influential predictor, followed by the COD service usage status. The results suggest that leveraging datasets from small businesses to construct predictive models can serve as an effective decision-support tool. Furthermore, this enables businesses to formulate proactive customer retention strategies, increase re-usage rates, and establish a competitive advantage in the logistics industry.

Article Details

Section
Research Articles

References

กาญจนา หฤหรรษพงศ์ และ ปิยมาศ จิตตระ. (2561). การวิเคราะห์แบบ RFM ในการแบ่งกลุ่มผู้ใช้งานเครื่องพิมพ์และเครื่องถ่ายเอกสารในองค์กร กรณีศึกษาสํานักวิชาสารสนเทศศาสตร์ มหาวิทยาลัยวลัยลักษณ์. วารสารวิชาการการจัดการเทคโนโลยีสารสนเทศและนวัตกรรม, 5(1), 21-29.

จามรกุล เหล่าเกียรติกุล. (2558). เหมืองข้อมูลเบื้องต้น : Introduction to Data Mining. กรุงเทพฯ : แดเน็กซ์ อินเตอร์คอร์ปอเรชั่น.

ชาญชวัฒน์ ภัคดีศรี. (2566). การพัฒนากลยุทธ์ทางการตลาดโดยใช้วิทยาศาสตร์ข้อมูลในการจัดกลุ่มลูกค้าตามพฤติกรรมการซื้อสินค้า [การค้นคว้าอิสระปริญญามหาบัณฑิต, มหาวิทยาลัยธรรมศาสตร์]. วิทยาลัยนวัตกรรม. https://digital.library.tu.ac.th/tu_dc/frontend/Info/item/dc:313603

ณรรฐคุณ วิรุฬห์ศรี, ลัทธพล โชครัตน์ประภา, ณัฐณิชา ศรีสมาน และพรทิพย์ เดชพิชัย. (2565). การวิเคราะห์แบ่งกลุ่มลูกค้าโดยใช้พฤติกรรมการซื้อเชิงลึก: กรณีศึกษา บริษัทผู้ผลิตอาหารสัตว์เลี้ยงแห่งหนึ่ง. วารสารวิทยาศาสตร์ลาดกระบัง, 31(1), 103-120.

ปัณดารีย์ สุนทรวราภาส, น้ำทิพย์ ตระกูลเมฆี, และ สูรีนา มะตาหยง. (2563). การวิเคราะห์แบบ RFM ในการแบ่งกลุ่มผู้ใช้งานตามพฤติกรรมการยืมทรัพยากรสารสนเทศ ของสำนักทรัพยากรการเรียนรู้คุณหญิงหลงอรรถกระวีสุนทร มหาวิทยาลัยสงขลานครินทร์. วารสารบรรณศาสตร์ มศว, 13(1), 30-45.

ปรารถนา ด่านก่อโพธิ์. (2567). การพยากรณ์ยอดขายโดยใช้อัลกอริทึม XGBoost และ TimesFM [วิทยานิพนธ์ปริญญามหาบัณฑิต, จุฬาลงกรณ์มหาวิทยาลัย].

https://www.cp.eng.chula.ac.th/~prabhas/thesis/prathana_complete_2024.pdf

วิรากานต์ กิตติบวรกุล, ศรายุทธ นนท์ศิริ, และ พิชิตชัย คำอินทร์. (2565). การเปรียบเทียบประสิทธิภาพเทคนิคการเรียนรู้ของเครื่องสำหรับการบำรุงรักษาเชิงคาดการณ์ของเครื่องยนต์อากาศยาน. วารสารวิชาการสมาคมสถาบันอุดมศึกษาเอกชนแห่งประเทศไทย (ฉบับวิทยาศาสตร์และเทคโนโลยี), 11(1), 16–29.

Ang, L., & Buttle, F. (2006). Customer retention management processes: A survey of Australian companies. European Journal of Marketing, 40(1/2), 83-99. https://doi.org/10.1108/03090560610637329

Gundogdu, S. (2022). Hepatitis C Disease Detection Based on PCA–SVM Model. Hittite Journal of Science and Engineering, 9(2), 111-116.

Hosmer, D. W., Lemeshow, S., & Sturdivant, R. X. (2013). Applied logistic regression (3rd ed.). John Wiley & Sons.

Kim, J., & Seok, J. (2024). ctGAN: Combined transformation of gene expression and survival data with generative adversarial network. Briefings in Bioinformatics, 25(4), bbae325. https://doi.org/10.1093/bib/bbae325

Korstanje, J. (2021). Gradient Boosting with XGBoost and LightGBM. In Advanced Forecasting with Python. (pp. 203-221) Apress. Berkeley, CA. https://doi.org/10.1007/978-1-4842-7150-6_15.

Meric, E., & Ozer, C. (2022). Symptom Based Health Status Prediction via Decision Tree, KNN, XGBoost, LDA, SVM, and Random Forest. In Proceedings of International Conference on Computing Intelligence and Data Analytics, Koceli, Turkey, 193-207.

Monisha, A. S., Suresh Kumar, N., & Sreeramulu, M. (2020). Customer Segmentation and Repurchase Intention using RFM Model and Machine Learning. International Journal of Engineering and Advanced Technology, 9(4), 2142-2146.

Moon, Jaeuk & Jung, Seungwon & Park, Sungwoo & Hwang, Eenjun. (2020). Conditional Tabular GAN-Based Two-Stage Data Generation Scheme for Short-Term Load Forecasting. IEEE Access. 8. 205327-205339. https://doi.org/10.1109/ACCESS.2020.3037063.

Phanbua, P., Arwatchananukul, S., Hristov, G., & Temdee, P. (2025). CTGAN-augmented ensemble learning models for classifying dementia and heart failure. Inventions, 10(6), 101. https://doi.org/10.3390/inventions10060101

Rizkyanto, H., & Gaol, F. L. (2023). Customer segmentation of personal credit using Recency, Frequency, Monetary (RFM) and K-means on financial industry. International Journal of Advanced Computer Science and Applications, 14(4), 312-318. https://doi.org/10.14569/IJACSA.2023.0140417

Sinchana, K. C., Anthraper, M. G., Sanjaykumar, K., Kumari, S., & D., U. (2025). Synthetic data generation using CTGAN with agentic workflows and retrieval-augmented generation. Proceedings of the 5th International Conference on AI Research (ICAIR 2025), 5(1), 472-480. https://doi.org/10.34190/icair.5.1.4280

Sunarya, P.A., Rahardja, U., Chen, S.C. , Choi, K., & Ku, C. T. (2024) Deciphering Digital Social Dynamics: A Comparative Study of Logistic Regression and Random Forest in Predicting e-Commerce Customer Behavior. Journal of Applied Data Sciences, 5(1), 100-113. https://doi.org/10.47738/jads.v5i1.155

Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (Article No. 659, pp. 7335–7345). Curran Associates Inc.