Marketing Campaign Prediction
Built a comprehensive machine learning system in R to predict customer response to marketing campaigns, analyzing 22,141 customer records with Decision Tree, Logistic Regression, KNN, and Cluster Analysis.
Introduction & Business Context
In order to maintain a competitive edge and sustain their progress, many companies employ marketing campaigns as a strategy to engage with their customers, thereby enhancing their sales landscape and customer retention.
"Customer acquisition costs are approximately 5-6 times higher than retaining existing customers." — Colgate & Danaher, 2000
This project exemplifies the profound impact of machine learning, offering a comprehensive perspective on accurately predicting customer responses to promotional offers based on individual customer traits.
Project Objectives
- Extract customer data from MySQL database
- Preprocess and clean the dataset
- Explore data through comprehensive EDA
- Build multiple classification models
- Segment customers through cluster analysis
Dataset Overview
Dataset retrieved from MySQL database with no missing values. Majority of participants are older individuals (avg. birth year: 1969) with income ranging from $1,730 to $666,666.
Exploratory Data Analysis
Customer Demographics
- • Average birth year: 1969
- • Mean recency: ~49 days since last purchase
- • Average web visits: 5 per month
- • Online purchases: 4 avg vs Store: 6 avg
Education Level Analysis
9,558 graduates declined the campaign offer vs 1,592 who accepted. Higher decline rate observed among graduate customers compared to PhD and Master's holders.
Response Distribution
15.32% acceptance rate indicates class imbalance in the target variable, requiring careful model handling.
Model Building & Comparison
1. Decision Tree
Built using rpart
package with visualization via rpart.plot. Evaluated with
confusion matrix, ROC curve, and AUC score.
2. Logistic Regression
Binary classification with probability-based predictions. Full ROC and AUC analysis performed for model evaluation.
3. K-Nearest Neighbor (KNN)
Instance-based learning using class package. Distance-based classification with confusion
matrix evaluation.
4. Cluster Analysis
Customer segmentation using cluster and factoextra
packages. Identified distinct customer groups based on purchasing behavior and demographics.
Model Evaluation Metrics
Confusion Matrix
ROC Curve
AUC Score
Conclusions & Recommendations
Key Conclusions
- • Customer acquisition costs 5-6x more than retention
- • Graduate customers showed higher campaign decline rates
- • Customers show different patterns for online vs in-store purchases
- • Cluster analysis reveals actionable customer segments
Recommendations
- • Targeted Campaigns: Focus on high-response segments
- • Retention Strategy: Invest in existing customers
- • Channel Optimization: Tailor by purchase channel preference
- • Education-Based Targeting: Develop segment-specific strategies
Tech Stack
Dataset Features
- ID (Customer ID)
- Year_Birth
- Education
- MaritalStatus
- Income
- Recency
- NumWebPurchases
- NumStorePurchases
- NumWebVisitsMonth
- Response (Target)
Key Statistics
- Records 22,141
- Avg Income $52,514
- Response Rate 15.32%
- Models Built 4
- Missing Values 0
References
- • Ascarza et al. (2018) — Customer retention
- • Colgate & Danaher (2000) — Acquisition costs