Heart Disease Prediction
DSN X Microsoft 2024 AI Bootcamp Qualification Hackathon — Built a machine learning model to predict heart disease risk using patient data, achieving ~80.9% accuracy with LightGBM and ensemble techniques.
DSN X Microsoft 2024 AI Bootcamp
Qualification Hackathon on Zindi
Background & Problem Statement
Heart disease is a leading cause of death worldwide. Early detection and prediction are crucial for timely intervention and improved patient outcomes. Machine learning techniques have shown promising results in predicting heart disease risk based on individual patient characteristics.
This project addresses the problem of predicting heart disease using a machine learning model trained on a dataset of patient information. The goal is to develop a model that accurately classifies individuals as having or not having heart disease.
Project Objectives
- Explore and analyze the heart disease dataset
- Preprocess data, handle missing values, and engineer features
- Develop a machine learning model for prediction
- Evaluate performance and compare different algorithms
Exploratory Data Analysis
Univariate Analysis Findings
- • Age distribution slightly right-skewed (most patients middle-aged)
- • More male patients than female in dataset
- • Most experienced atypical angina or non-anginal pain
- • Majority had fasting blood sugar below 120 mg/dl
- • Class imbalance in target variable detected
Bivariate Analysis Insights
- • Age: Older patients slightly more susceptible
- • Cholesterol: No clear trend with heart disease
- • Chest pain type: Asymptomatic = higher likelihood
- • Exercise-induced angina: Strong risk indicator
- • Chest pain + max heart rate jointly influence risk
Feature Engineering
- Label Encoding: Categorical features (sex, chest pain type, ST slope) transformed for algorithm compatibility
- Age Binning: Ages grouped into Young, Adult, Middle Age, Old for capturing non-linear relationships
- SMOTE: Synthetic Minority Over-sampling Technique to address class imbalance
- StandardScaler: Normalization applied across all numerical features
Model Building & Comparison
Models Evaluated
Hyperparameter Tuning
RandomizedSearchCV employed for efficient exploration — samples different combinations and evaluates through cross-validation. Parameters: n_estimators, max_depth, min_samples_split, min_samples_leaf.
Best Model: LightGBM
Selected for highest accuracy, computational efficiency, scalability, and effective handling of large datasets.
Key Results
Conclusions & Recommendations
Key Predictors Identified
Age, chest pain type, and exercise-induced angina emerged as significant predictors. Older patients and those with asymptomatic chest pain or angina during exercise are more likely to have heart disease.
Recommendations
- • Focus on key predictors for targeted interventions
- • Investigate variable interactions (max heart rate + chest pain type)
- • Clinical application for early screening and risk assessment
Tech Stack
Dataset Features
- Age, Sex
- Chest Pain Type (cp)
- Blood Pressure (trestbps)
- Cholesterol (chol)
- Fasting Blood Sugar (fbs)
- Max Heart Rate (thalach)
- Exercise Angina (exang)
- ST Depression (oldpeak)
- ST Slope, ca, thal