Capstone: build a loan-default classifier
Split, standardize, train logistic regression, and evaluate on a held-out test set - end to end with numpy.
- Run a complete, leak-free ML pipeline
- Train and threshold a logistic regression model
- Report test-set metrics against a baseline
Time to put the whole pipeline together on a small loan dataset - income, debt-to-income ratio and years employed - predicting whether an applicant defaults. You’ll write every step yourself:
- Split - the last quarter of rows is the test set.
- Standardize - with training statistics only.
- Train logistic regression with gradient descent.
- Evaluate on the test set: accuracy, precision, recall, against the majority-class baseline.
In real projects scikit-learn wraps each step (train_test_split, StandardScaler, LogisticRegression, classification_report) - but having written them once, you’ll know exactly what those calls do.
1import numpy as np
2data = np.arange(20).reshape(10, 2)
3cut = int(len(data) * 0.75)
4print(data[:cut].shape, data[cut:].shape)(7, 2) (3, 2)
Key takeaways
Split first, then fit every transformation and the model on training data only.
Gradient descent on log loss trains logistic regression.
Evaluate once on the test set, against a baseline, with metrics that match the costs of errors.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Step 1: split and standardize
Each line is income dti years defaulted. Use the first 75% of rows (int(n * 0.75)) for training and the rest for testing. Standardize the three features with the training mean and standard deviation. Print train: A rows, test: B rows, train means: [...] and test row 1 scaled: [...] (lists rounded to 2 decimals with np.round(...).tolist()).
- Loans
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Step 2: train and evaluate
Same input and split. Standardize as in step 1, then train logistic regression from zero weights with learning rate 0.1 for 3000 steps (gradients as in the logistic regression lesson). On the test rows, predict default when and print:
1baseline accuracy: X% (always predict the majority training label)
2model accuracy: X%
3precision: P (2 decimals, for the default class)
4recall: R (2 decimals)
5largest weight: NAME (by absolute value: income, dti or years)- Loans
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…