Customer Churn Prediction for Consumer-Facing Delivery Apps

Mufti, Muhammad Jawad (2026) Customer Churn Prediction for Consumer-Facing Delivery Apps. Masters thesis, King Fahd University of Petroleum and Minerals.

[img] PDF (MS Thesis)
202392310_Thesis_Report.pdf - Accepted Version
Restricted to Repository staff only until 31 July 2027.
Available under License Creative Commons Attribution.

Download (3MB)

Arabic Abstract

يُعد التنبؤ بانقطاع العملاء مهمة مهمة للشركات الخدمية وغير التعاقدية، حيث قد يتوقف العملاء عن استخدام الخدمة دون إلغاء علاقتهم بها بشكل رسمي. غالباً ما تمثل الدراسات التقليدية في مجال التنبؤ بانقطاع العملاء كل عميل على شكل سجل ثابت، وتركز بصورة رئيسية على التنبؤ بتسمية الانقطاع أو درجة احتمال الانقطاع. ومع ذلك، توفر هذه الأساليب دعماً محدوداً لتقييم مخاطر الانقطاع بصورة مستمرة، وتشخيص أسباب الانقطاع، واتخاذ قرارات احتفاظ قابلة للتنفيذ. تقترح هذه الرسالة إطار عمل قائماً على النافذة المتحركة للتنبؤ بانقطاع العملاء ودعم القرار في منصة غير تعاقدية لخدمات غسيل السيارات. يحول الإطار سجلات الحجز الخام ذات الطوابع الزمنية إلى أمثلة زمنية على مستوى العميل باستخدام نوافذ متطابقة لمراقبة السلوك وتقييم الانقطاع بطول 30 يوماً، و15 يوماً، و7 أيام. وتم استخراج 55 خاصية سلوكية وتشغيلية من كل نافذة مراقبة، بحيث تغطي نشاط الحجز، وحداثة التفاعل، والسلوك المالي، واستخدام الاشتراك، ونتائج الخدمة، والخصائص المكانية، والتفضيلات الفئوية. تم تدريب وتقييم نماذج تعلم آلي، شملت الانحدار اللوجستي، والغابات العشوائية، وXGBoost، ونماذج تعلم عميق، شملت LSTM وGRU، عبر آفاق التنبؤ الثلاثة. استُخدم الاستدعاء معياراً رئيسياً لاختيار النموذج لأن تحديد أكبر عدد ممكن من العملاء المنقطعين فعلياً يُعد أكثر أهمية من تعظيم الدقة العامة في سياق الاحتفاظ بالعملاء. حقق نموذج LSTM أفضل أداء قائم على الاستدعاء عبر إعدادات النوافذ الثلاثة. ففي مجموعة بيانات 30 يوماً، حقق نموذج LSTM الأساسي دقة اختبار بلغت 90.0%، ودقة تنبؤ بلغت 90.0%، واستدعاء بلغ 96.8%، ودرجة F1 بلغت 93.3%، وقيمة ROC-AUC بلغت 0.939. وفي مجموعة بيانات 15 يوماً، حقق نموذج LSTM المضبوط دقة اختبار بلغت 89.2%، ودقة تنبؤ بلغت 92.2%، واستدعاء بلغ 96.4%، ودرجة F1 بلغت 93.2%، وقيمة ROC-AUC بلغت 0.944. أما في مجموعة بيانات 7 أيام، فقد حقق نموذج LSTM المضبوط دقة اختبار بلغت 87.2%، ودقة تنبؤ بلغت 91.9%، واستدعاء بلغ 91.8%، ودرجة F1 بلغت 91.8%، وقيمة ROC-AUC بلغت 0.932. إلى جانب التنبؤ، يدمج الإطار المقترح نظام دعم قرار يستخدم التحليل الاستكشافي للبيانات والتشخيص القائم على القواعد لتحديد الأسباب المحتملة لانقطاع العملاء المتوقع انقطاعهم. يتم تقسيم العملاء وفقاً لعدد الحجوزات الناجحة، ويتم تحديد أسباب مختلفة للانقطاع عبر مراحل العلاقة المختلفة، بما في ذلك فشل الحجز، وفشل خدمة الاشتراك، والاعتماد على الرموز الترويجية، والخمول، وفقدان الاشتراك، وعبء السعر. ثم تُربط هذه الأسباب باستراتيجيات احتفاظ مستندة إلى آراء الخبراء. ولتقييم قابلية التعميم، تم تطبيق منهجية النافذة المقترحة أيضاً على مجموعة بيانات طلبات عملاء Instacart. وفي مجموعة البيانات الخارجية هذه، حقق نموذج XGBoost أفضل أداء عام بدقة اختبار بلغت 94.12% وقيمة ROC-AUC بلغت 0.9777، بينما حققت جميع النماذج المقيمة دقة اختبار أعلى من 91%. وتُظهر النتائج أن الإطار المقترح زمني، وقابل للتفسير، وموجه نحو القرار، وقابل للتعميم على مجالات أخرى من بيانات معاملات العملاء.

English Abstract

Customer churn prediction is an important task for consumer-facing delivery applications, where customers may stop using a service without formally cancelling their relationship. Traditional churn prediction studies commonly represent each customer as a static record and focus mainly on predicting a churn label or probability score. However, such approaches provide limited support for continuous churn-risk reassessment, churn reason diagnosis, and actionable retention decision-making. This thesis proposes a rolling-window churn prediction and decision support framework for a non-contractual car wash service platform. The framework transforms raw time-stamped booking records into customer-window instances using matched behavioral observation and churn-evaluation windows of 30 days, 15 days, and 7 days. A total of 54 behavioral and operational features are engineered from each observation window, capturing booking activity, recency, monetary behavior, subscription usage, service outcomes, spatial characteristics, and categorical preferences. Machine learning models, including Logistic Regression, Random Forest, and XGBoost, and deep learning models, including LSTM and GRU, are trained and evaluated across the three churn prediction horizons. Recall is used as the primary model selection criterion because identifying actual churned customers is more important than maximizing overall accuracy in a retention context. The LSTM model achieved the best recall-based performance across all three window settings. For the 30-day dataset, the baseline LSTM achieved 90.0% test accuracy, 90.0% precision, 96.8% recall, 93.3% F1-score, and 0.939 ROC-AUC. For the 15-day dataset, the tuned LSTM achieved 89.2% test accuracy, 92.2% precision, 96.4% recall, 93.2% F1-score, and 0.944 ROC-AUC. For the 7-day dataset, the tuned LSTM achieved 87.2% test accuracy, 91.9% precision, 91.8% recall, 91.8% F1-score, and 0.932 ROC-AUC. Beyond prediction, the proposed framework integrates a decision support system that uses exploratory data analysis and rule-based diagnosis to assign possible churn reasons to predicted churned customers. Customers are segmented by successful booking count, and different churn reasons are identified across relationship stages, including booking failure, subscription service failure, promo-code dependence, inactivity, subscription loss, and price burden. These reasons are then mapped to expert-informed retention strategies. To evaluate generalizability, the proposed 30-day window methodology is also applied to the Instacart customer order dataset. On this external dataset, XGBoost achieved the best overall performance with 94.12% test accuracy and 0.9777 ROC-AUC, while all evaluated models achieved test accuracies above 91%. The results show that the proposed framework is temporal, interpretable, decision-oriented, and generalizable to other customer transaction domains.

Item Type: Thesis (Masters)
Subjects: Computer
Department: College of Computing and Mathematics > Information and Computer Science
Thesis Advisor:
Omar Hammad,
Thesis Committee Members:
Moataz Ahmed, Haitham Saleh,
Depositing User: MUHAMMAD JAWAD MUFTI
Date Deposited: 04 Aug 2026 11:01
Last Modified: 04 Aug 2026 11:01
URI: https://eprints.kfupm.edu.sa/id/eprint/144662