A Development of Predictive Model for Customer Retention using Machine Learning Techniques: A Case Study of a Parcel Service Franchise
Main Article Content
Abstract
This research aims to 1) develop a predictive model using machine learning techniques to forecast customer retention, 2) compare the performance of various machine learning models, and 3) analyze key determinants influencing customer re-usage behavior in a parcel delivery franchise. The research methodology was conducted based on the Cross-Industry Standard Process for Data Mining (CRISP-DM) The study utilized empirical transaction data from August 2025 to January 2026. Features extraction was based on the Recency, Frequency, and Monetary (RFM) analysis framework, supplemented by the Cash on Delivery (COD) service. The Tabular Generative Adversarial Networks (CTGAN) was employed to synthesize data and mitigate class imbalance, enhancing model performance. Subsequently, four machine learning algorithms as Logistic Regression, Random Forest, XGBoost, and Support Vector Machine (SVM) were evaluated. The findings successfully addressed the research objectives as follows: 1) The predictive model was successfully developed, effectively mitigating the learning bias caused by data imbalance. 2) The comparative analysis demonstrated that the XGBoost model achieved the highest predictive performance, yielding an Area Under the ROC Curve (AUC-ROC) of 0.8462, thereby proving to be the most suitable algorithm for predicting customer return. 3) The feature importance analysis indicated that the Recency of service usage was the most influential predictor, followed by the COD service usage status. The results suggest that leveraging datasets from small businesses to construct predictive models can serve as an effective decision-support tool. Furthermore, this enables businesses to formulate proactive customer retention strategies, increase re-usage rates, and establish a competitive advantage in the logistics industry.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
ข้อความและบทความในวารสารนวัตกรรมการบริหารและการจัดการ เป็นแนวคิดของผู้เขียน ไม่ใช่ความคิดเห็นและความรับผิดชอบของคณะผู้จัดทำ บรรณาธิการ กองบรรณาธิการ วิทยาลัยนวัตกรรมการจัดการ และมหาวิทยาลัยเทคโนโลยีราชมงคลรัตนโกสินทร์
ข้อความ ข้อมูล เนื้อหา รูปภาพ ฯลฯ ที่ได้รับการีพิมพ์ในวารสารนวัตกรรมการบริหารและการจัดการ ถือเป็นลิขสิทธิ์ของวารสารนวัตกรรมการบริหารและการจัดการ หากบุคคลใดหรือหน่วยงานใดต้องการนำทั้งหมดหรือส่วนหนึ่งส่วนใดไปเผยแพร่ต่อหรือกระทำการใดๆ จะต้องได้รับอนุญาติเป็นลายลักษณ์อักษรจากวารสารนวัตกรรมการบริหารและการจัดการก่อนเท่านั้น
References
กาญจนา หฤหรรษพงศ์ และ ปิยมาศ จิตตระ. (2561). การวิเคราะห์แบบ RFM ในการแบ่งกลุ่มผู้ใช้งานเครื่องพิมพ์และเครื่องถ่ายเอกสารในองค์กร กรณีศึกษาสํานักวิชาสารสนเทศศาสตร์ มหาวิทยาลัยวลัยลักษณ์. วารสารวิชาการการจัดการเทคโนโลยีสารสนเทศและนวัตกรรม, 5(1), 21-29.
จามรกุล เหล่าเกียรติกุล. (2558). เหมืองข้อมูลเบื้องต้น : Introduction to Data Mining. กรุงเทพฯ : แดเน็กซ์ อินเตอร์คอร์ปอเรชั่น.
ชาญชวัฒน์ ภัคดีศรี. (2566). การพัฒนากลยุทธ์ทางการตลาดโดยใช้วิทยาศาสตร์ข้อมูลในการจัดกลุ่มลูกค้าตามพฤติกรรมการซื้อสินค้า [การค้นคว้าอิสระปริญญามหาบัณฑิต, มหาวิทยาลัยธรรมศาสตร์]. วิทยาลัยนวัตกรรม. https://digital.library.tu.ac.th/tu_dc/frontend/Info/item/dc:313603
ณรรฐคุณ วิรุฬห์ศรี, ลัทธพล โชครัตน์ประภา, ณัฐณิชา ศรีสมาน และพรทิพย์ เดชพิชัย. (2565). การวิเคราะห์แบ่งกลุ่มลูกค้าโดยใช้พฤติกรรมการซื้อเชิงลึก: กรณีศึกษา บริษัทผู้ผลิตอาหารสัตว์เลี้ยงแห่งหนึ่ง. วารสารวิทยาศาสตร์ลาดกระบัง, 31(1), 103-120.
ปัณดารีย์ สุนทรวราภาส, น้ำทิพย์ ตระกูลเมฆี, และ สูรีนา มะตาหยง. (2563). การวิเคราะห์แบบ RFM ในการแบ่งกลุ่มผู้ใช้งานตามพฤติกรรมการยืมทรัพยากรสารสนเทศ ของสำนักทรัพยากรการเรียนรู้คุณหญิงหลงอรรถกระวีสุนทร มหาวิทยาลัยสงขลานครินทร์. วารสารบรรณศาสตร์ มศว, 13(1), 30-45.
ปรารถนา ด่านก่อโพธิ์. (2567). การพยากรณ์ยอดขายโดยใช้อัลกอริทึม XGBoost และ TimesFM [วิทยานิพนธ์ปริญญามหาบัณฑิต, จุฬาลงกรณ์มหาวิทยาลัย].
https://www.cp.eng.chula.ac.th/~prabhas/thesis/prathana_complete_2024.pdf
วิรากานต์ กิตติบวรกุล, ศรายุทธ นนท์ศิริ, และ พิชิตชัย คำอินทร์. (2565). การเปรียบเทียบประสิทธิภาพเทคนิคการเรียนรู้ของเครื่องสำหรับการบำรุงรักษาเชิงคาดการณ์ของเครื่องยนต์อากาศยาน. วารสารวิชาการสมาคมสถาบันอุดมศึกษาเอกชนแห่งประเทศไทย (ฉบับวิทยาศาสตร์และเทคโนโลยี), 11(1), 16–29.
Ang, L., & Buttle, F. (2006). Customer retention management processes: A survey of Australian companies. European Journal of Marketing, 40(1/2), 83-99. https://doi.org/10.1108/03090560610637329
Gundogdu, S. (2022). Hepatitis C Disease Detection Based on PCA–SVM Model. Hittite Journal of Science and Engineering, 9(2), 111-116.
Hosmer, D. W., Lemeshow, S., & Sturdivant, R. X. (2013). Applied logistic regression (3rd ed.). John Wiley & Sons.
Kim, J., & Seok, J. (2024). ctGAN: Combined transformation of gene expression and survival data with generative adversarial network. Briefings in Bioinformatics, 25(4), bbae325. https://doi.org/10.1093/bib/bbae325
Korstanje, J. (2021). Gradient Boosting with XGBoost and LightGBM. In Advanced Forecasting with Python. (pp. 203-221) Apress. Berkeley, CA. https://doi.org/10.1007/978-1-4842-7150-6_15.
Meric, E., & Ozer, C. (2022). Symptom Based Health Status Prediction via Decision Tree, KNN, XGBoost, LDA, SVM, and Random Forest. In Proceedings of International Conference on Computing Intelligence and Data Analytics, Koceli, Turkey, 193-207.
Monisha, A. S., Suresh Kumar, N., & Sreeramulu, M. (2020). Customer Segmentation and Repurchase Intention using RFM Model and Machine Learning. International Journal of Engineering and Advanced Technology, 9(4), 2142-2146.
Moon, Jaeuk & Jung, Seungwon & Park, Sungwoo & Hwang, Eenjun. (2020). Conditional Tabular GAN-Based Two-Stage Data Generation Scheme for Short-Term Load Forecasting. IEEE Access. 8. 205327-205339. https://doi.org/10.1109/ACCESS.2020.3037063.
Phanbua, P., Arwatchananukul, S., Hristov, G., & Temdee, P. (2025). CTGAN-augmented ensemble learning models for classifying dementia and heart failure. Inventions, 10(6), 101. https://doi.org/10.3390/inventions10060101
Rizkyanto, H., & Gaol, F. L. (2023). Customer segmentation of personal credit using Recency, Frequency, Monetary (RFM) and K-means on financial industry. International Journal of Advanced Computer Science and Applications, 14(4), 312-318. https://doi.org/10.14569/IJACSA.2023.0140417
Sinchana, K. C., Anthraper, M. G., Sanjaykumar, K., Kumari, S., & D., U. (2025). Synthetic data generation using CTGAN with agentic workflows and retrieval-augmented generation. Proceedings of the 5th International Conference on AI Research (ICAIR 2025), 5(1), 472-480. https://doi.org/10.34190/icair.5.1.4280
Sunarya, P.A., Rahardja, U., Chen, S.C. , Choi, K., & Ku, C. T. (2024) Deciphering Digital Social Dynamics: A Comparative Study of Logistic Regression and Random Forest in Predicting e-Commerce Customer Behavior. Journal of Applied Data Sciences, 5(1), 100-113. https://doi.org/10.47738/jads.v5i1.155
Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (Article No. 659, pp. 7335–7345). Curran Associates Inc.