Skip to content

01 - Introduction

Get the Video Course

Before we write any code, this section shows three things: what you will build, the problem we solve and how we measure it in dollars, and the machine learning workflow the whole project follows.

Table of Contents

Step Topic
Step-01 Welcome: What You Will Build
Step-02 Before You Start
Step-03 The Fraud Problem, in Dollars
Step-04 The ML Workflow and How Our Project Follows It

Step-01: Welcome: What You Will Build

You build one complete machine learning project in Python: a credit card fraud detector, from raw data to your own prediction API. It is plain Python scripts, not notebooks, organized the way real projects are.

What you will build: five steps from generating 10,000 card transactions, exploring the fraud patterns, training a scikit-learn pipeline, improving it with new features and tuning judged in dollars, to serving it through your own FastAPI inference API with a web page

The course has nine sections, and every section adds one piece to the same project:

The course map: 01 Introduction, 02 Python Environment Setup, 03 Synthetic Dataset Generation, 04 A Model Training and Evaluation, 04 B Model Serving with an Inference API, 05 Exploratory Data Analysis, 06 Feature Engineering, 07 Hyperparameter Tuning, 08 Model Serving with the Tuned Inference API, about 14 hours of video

We judge every change by the money it saves the bank:

The journey in dollars: baseline model 39,625 dollars, plus feature engineering 39,700 dollars, plus random search tuning 37,025 dollars, plus tuning that respects the money 46,525 to 50,550 dollars depending on the machine

Who this course is for, and what you need before you start:

Who this course is for: DevOps, SRE and cloud engineers moving into machine learning, developers and students who want one real ML project end to end, and Python learners ready for their first ML project. Before you start: Python basics, no machine learning needed, a Mac, Windows or Linux computer


Step-02: Before You Start

  • Python basics: functions, lists, dictionaries and simple classes. New to Python? Take our Python for Absolute Beginners course first.
  • Your numbers may differ. Most of this course was recorded on an Intel Mac, and the final runs (the numbers you see in the slides and READMEs) were done on an Apple Silicon Mac. So you will see differences between the videos and the written numbers, for example $46,525 on camera vs $50,550 in Section 07, and the winning class_weight can come out as 1:100 or 1:200. Your own results depend on your processor, its math library, and small numerical differences in a search that does not always fully converge. The improvement story stays the same. Section 07 explains why this happens.
  • About MLflow, KServe and "advanced sections": when a lecture mentions MLflow, KServe, DVC, Docker, Kubernetes or "advanced sections", it refers to our MLOps courses. This course covers the machine learning side end to end.
  • If pip is missing in a new conda environment, run conda install pip inside it (Section 02).

Step-03: The Fraud Problem, in Dollars

Our data has 10,000 card transactions, and only 298 of them are fraud, about 3%. That makes accuracy a trap: a "model" that says legitimate to every transaction is 97% accurate and catches zero fraud.

The fraud problem: 10,000 card transactions with only 298 fraud, about 3 percent. A model that says legit to everything is 97 percent accurate and catches zero fraud, so the right question is how much money the model saves the bank

So we ask a better question: how much money does the model save the bank? Every decision has a price:

Four outcomes and their prices: catching a fraud earns 1,475 dollars, missing a fraud costs 1,500 dollars, a false alarm on an honest customer costs 25 dollars, and letting an honest customer pass costs nothing

Outcome Price Why
Fraud, we catch it + $1,475 We stop $1,500 of fraud and spend $25 to check it
Fraud, we miss it - $1,500 The fraudster keeps the money
Honest customer, we flag it (false alarm) - $25 A quick check of a good transaction
Honest customer, we let it pass $0 The normal case

Missing a fraud ($1,500) costs 60 times more than one check ($25). That is why the model will happily check some honest customers to catch more fraud.

Adding the four prices over all test transactions gives one number, Net Benefit:

Net Benefit formula: caught times 1,475 minus false alarms times 25 minus missed times 1,500. Example for the first model: 46 caught, 289 false alarms, 14 missed, Net Benefit 39,625 dollars

Net Benefit = caught x $1,475 - false alarms x $25 - missed x $1,500

First model (Section 04):  46 x $1,475 - 289 x $25 - 14 x $1,500 = $39,625

Notice: turning one missed fraud into a caught fraud moves the score by about $3,000 (from - $1,500 to + $1,475). That is why catching fraud matters so much in this project.


Step-04: The ML Workflow and How Our Project Follows It

Every machine learning project goes through the same basic steps. It is a loop, not a straight line: we come back to earlier steps whenever we learn something new.

The machine learning workflow in seven steps: data collection, explore with EDA, feature engineering, train, evaluate, tune and serve, followed by monitoring and retraining after the model goes live

In our project, each step is a Python script, and each section builds one of them:

How our project follows the workflow: p1_01 generates the data in Section 03, p3_01 trains and evaluates the baseline in Section 04 A, p3_02 and p3_03 serve the baseline API in Section 04 B, p2 explores the data in Section 05, p4 scripts do feature engineering in Section 06, p5_01 to p5_03 tune and train the production model in Section 07, and p5_04 and p5_05 serve the tuned API in Section 08

Step Script(s) Section
Data p1_01_generate_initial_dataset.py 03
Train + evaluate (baseline) p3_01_train_model_baseline.py 04 A
Serve (baseline API) p3_02_inference_api_baseline.py + p3_03 test 04 B
Explore (EDA) p2_eda_base.py + auto-EDA tools 05
Feature engineering p4_01 / p4_02 / p4_03 06
Tune + production model p5_01 / p5_02 / p5_03 07
Serve (tuned API) p5_04_inference_api.py + p5_05 test 08

We build a working baseline first (Section 04), then come back to explore and improve it.

One idea runs through every section: train on one part of the data, test on a part the model has never seen.

The 80/20 split: a training set of 8,000 transactions with about 238 frauds that the model learns from, and a test set of 2,000 transactions with 60 frauds. Like a student's exam, you are tested on new questions. Stratify keeps the same 3 percent fraud in both parts

  • Training set (80%): 8,000 transactions. The model learns from these.
  • Test set (20%): 2,000 transactions, 60 of them fraud. Every number you will see in this course, including Net Benefit, is measured on these.
  • Stratify: we keep the same 3% fraud in both parts, so the test set is a fair copy of the real world.

Next: 02 - Setup Python

Prefer to learn by watching?

The video course builds this whole project with you on screen, step by step.

Get the Video Course


Next: 02 - Setup Python Environment