Skip to Content
Machine Learning / Python

Heart Disease Prediction

DSN X Microsoft 2024 AI Bootcamp Qualification Hackathon — Built a machine learning model to predict heart disease risk using patient data, achieving ~80.9% accuracy with LightGBM and ensemble techniques.

Heart Disease Prediction Dashboard

DSN X Microsoft 2024 AI Bootcamp

Qualification Hackathon on Zindi

Background & Problem Statement

Heart disease is a leading cause of death worldwide. Early detection and prediction are crucial for timely intervention and improved patient outcomes. Machine learning techniques have shown promising results in predicting heart disease risk based on individual patient characteristics.

This project addresses the problem of predicting heart disease using a machine learning model trained on a dataset of patient information. The goal is to develop a model that accurately classifies individuals as having or not having heart disease.

Project Objectives

  • Explore and analyze the heart disease dataset
  • Preprocess data, handle missing values, and engineer features
  • Develop a machine learning model for prediction
  • Evaluate performance and compare different algorithms

Exploratory Data Analysis

Univariate Analysis Findings

  • • Age distribution slightly right-skewed (most patients middle-aged)
  • • More male patients than female in dataset
  • • Most experienced atypical angina or non-anginal pain
  • • Majority had fasting blood sugar below 120 mg/dl
  • • Class imbalance in target variable detected

Bivariate Analysis Insights

  • • Age: Older patients slightly more susceptible
  • • Cholesterol: No clear trend with heart disease
  • • Chest pain type: Asymptomatic = higher likelihood
  • • Exercise-induced angina: Strong risk indicator
  • • Chest pain + max heart rate jointly influence risk

Feature Engineering

  • Label Encoding: Categorical features (sex, chest pain type, ST slope) transformed for algorithm compatibility
  • Age Binning: Ages grouped into Young, Adult, Middle Age, Old for capturing non-linear relationships
  • SMOTE: Synthetic Minority Over-sampling Technique to address class imbalance
  • StandardScaler: Normalization applied across all numerical features

Model Building & Comparison

Models Evaluated

Random Forest
XGBoost
LightGBM ★
Stacking Ensemble

Hyperparameter Tuning

RandomizedSearchCV employed for efficient exploration — samples different combinations and evaluates through cross-validation. Parameters: n_estimators, max_depth, min_samples_split, min_samples_leaf.

Best Model: LightGBM

Selected for highest accuracy, computational efficiency, scalability, and effective handling of large datasets.

Key Results

~80.9% Model Accuracy
10-Fold Cross Validation
10K Records Analyzed
4 Algorithms Compared

Conclusions & Recommendations

Key Predictors Identified

Age, chest pain type, and exercise-induced angina emerged as significant predictors. Older patients and those with asymptomatic chest pain or angina during exercise are more likely to have heart disease.

Recommendations

  • • Focus on key predictors for targeted interventions
  • • Investigate variable interactions (max heart rate + chest pain type)
  • • Clinical application for early screening and risk assessment

Tech Stack

Python Scikit-learn XGBoost LightGBM Pandas NumPy SMOTE Matplotlib Seaborn

Dataset Features

  • Age, Sex
  • Chest Pain Type (cp)
  • Blood Pressure (trestbps)
  • Cholesterol (chol)
  • Fasting Blood Sugar (fbs)
  • Max Heart Rate (thalach)
  • Exercise Angina (exang)
  • ST Depression (oldpeak)
  • ST Slope, ca, thal

Key Risk Factors

1 Age (older patients)
2 Chest Pain Type (asymptomatic)
3 Exercise-Induced Angina