Feature Engineering · Notebook 1 companion articleFoundations of Features
Before you change or create any column, you need to answer a simpler question: what is a feature, and what makes one good? This article answers it using a real, messy dataset of 7,043 customers. Along the way it corrects a few habits people pick up from textbook examples.
Start hereIntroduction#
Ask most beginners what makes a machine learning project work, and they'll say the model. Pick the right algorithm, tune it, and good predictions follow. That's often wrong. Two people can train the same model on the same target and get very different results. The difference is often the columns they gave it. One person gave the model columns that describe the problem. The other gave it whatever the database happened to store.
A feature is an input column the model is allowed to look at. Feature engineering is the work of choosing, cleaning and reshaping those columns so the model can use the patterns in them. It happens before model choice and tuning. If you get it wrong, tuning won't save you. If you get it right, even a simple logistic regression can do surprisingly well.
This article goes with Notebook 1 of the series. It covers the basics everything else builds on:
- what counts as a feature and what counts as the target
- why the problem type depends on the column you predict, not the dataset
- why a new feature sometimes helps a lot and sometimes does nothing
- the feature types you'll see in almost every table
- the one train/test split rule that, if broken, makes your results meaningless
Each idea is shown with real numbers, like a model that scores 100% on its training data and then does worse than always guessing the most common answer.
The dataset: Telco Customer Churn#
All examples use the Telco Customer Churn dataset. IBM released it as sample data, and it's now a common teaching dataset because it looks like a real business problem. Each of the 7,043 rows is one customer of a phone and internet company. The question is simple: did this customer cancel in the last month?
It's a good dataset for learning feature engineering because:
- It's big enough. With 7,043 rows, test scores are fairly stable. On tiny datasets, the score depends heavily on which rows land in the test set.
- It has every common feature type. Numbers, counts, yes/no flags, unordered categories and one ordered category all appear in its 21 columns.
- Its problems are real. A money column is stored as text. Eleven rows are blank for a reason that turns out to matter. Nobody added these problems for teaching.
- The business goal is easy to understand. Keeping a customer is cheaper than finding a new one, so knowing who will leave is clearly worth money.
The notebook reads the file from datasets/telco_churn.csv. Every code cell below is copied exactly from the notebook, with its real output underneath.
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, roc_auc_score
from sklearn.model_selection import StratifiedKFold, cross_val_score, train_test_split
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.tree import DecisionTreeClassifier
plt.style.use('seaborn-v0_8-whitegrid')
pd.set_option('display.max_columns', 30)
pd.set_option('display.width', 160)
RANDOM_STATE = 42
np.random.seed(RANDOM_STATE)
churn = pd.read_csv('datasets/telco_churn.csv')
print(f'Rows : {churn.shape[0]:,}')
print(f'Columns : {churn.shape[1]}')
churn.head()
Rows : 7,043 Columns : 21
customerID gender SeniorCitizen Partner Dependents tenure PhoneService MultipleLines InternetService OnlineSecurity OnlineBackup DeviceProtection \ 0 7590-VHVEG Female 0 Yes No 1 No No phone service DSL No Yes No 1 5575-GNVDE Male 0 No No 34 Yes No DSL Yes No Yes 2 3668-QPYBK Male 0 No No 2 Yes No DSL Yes Yes No 3 7795-CFOCW Male 0 No No 45 No No phone service DSL Yes No Yes 4 9237-HQITU Female 0 No No 2 Yes No Fiber optic No No No TechSupport StreamingTV StreamingMovies Contract PaperlessBilling PaymentMethod MonthlyCharges TotalCharges Churn 0 No No No Month-to-month Yes Electronic check 29.85 29.85 No 1 No No No One year No Mailed check 56.95 1889.5 No 2 No No No Month-to-month Yes Mailed check 53.85 108.15 Yes 3 Yes No No One year No Bank transfer (automatic) 42.30 1840.75 No 4 No No No Month-to-month Yes Electronic check 70.70 151.65 Yes
The shape is 7,043 rows and 21 columns, which matches the public version of this dataset. Checking the shape against the documentation is a quick way to confirm you loaded the right file. Also note that text columns like gender will show the dtype str in .info(), not object. That's explained below.
Section 1Features and targets#
And the ID column that quietly breaks models
Imagine an analyst at a phone company trying to guess who will cancel. They look at how long each person has been a customer, what they pay each month, their contract type and how many services they use. Then they make one call: will this person leave?
The things the analyst looks at are the features. The thing they're guessing is the target. Both have other names you'll hear in courses and interviews:
- Features are also called input variables, predictors, independent variables, attributes, or simply
X. - The target is also called the label, response variable, dependent variable, outcome, or simply
y.
The key idea: features are what you know at prediction time, and the target is what you want to know. A column only counts as a feature if you'd have it at the moment you make the prediction.
This is one of the most common ways models go wrong. A cancellation date, a support ticket marked "closing account" or a final bill would all make a churn model look great in testing. In real use they're worthless, because they only exist after the customer has already left.
churn.info()
<class 'pandas.DataFrame'> RangeIndex: 7043 entries, 0 to 7042 Data columns (total 21 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 customerID 7043 non-null str 1 gender 7043 non-null str 2 SeniorCitizen 7043 non-null int64 3 Partner 7043 non-null str 4 Dependents 7043 non-null str 5 tenure 7043 non-null int64 6 PhoneService 7043 non-null str 7 MultipleLines 7043 non-null str 8 InternetService 7043 non-null str 9 OnlineSecurity 7043 non-null str 10 OnlineBackup 7043 non-null str 11 DeviceProtection 7043 non-null str 12 TechSupport 7043 non-null str 13 StreamingTV 7043 non-null str 14 StreamingMovies 7043 non-null str 15 Contract 7043 non-null str 16 PaperlessBilling 7043 non-null str 17 PaymentMethod 7043 non-null str 18 MonthlyCharges 7043 non-null float64 19 TotalCharges 7043 non-null str 20 Churn 7043 non-null str dtypes: float64(1), int64(2), str(18) memory usage: 1.1 MB
.info() is usually the first thing to check on a new table. Three things stand out here. First, every column shows 7,043 non-null values, which suggests no missing data. That's wrong, and we'll see why later. Second, only SeniorCitizen, tenure and MonthlyCharges are numbers. Everything else is text. Third, TotalCharges is money but is stored as text. That's a problem, and you can spot it right here.
.info() tells you how pandas read the file, not what the data means. Treat it as something to check, not something to trust.
# The target: did this customer leave?
y = (churn['Churn'] == 'Yes').astype(int)
# Everything else is a candidate feature
X = churn.drop(columns=['Churn'])
print(f'X shape : {X.shape}')
print(f'y shape : {y.shape}')
print()
print('Target distribution:')
print(y.value_counts().rename({0: 'Stayed', 1: 'Churned'}))
print()
print(f'Churn rate: {y.mean():.1%}')
X shape : (7043, 20) y shape : (7043,) Target distribution: Churn Stayed 5174 Churned 1869 Name: count, dtype: int64 Churn rate: 26.5%
Two details in that cell matter. First, Churn was turned from 'Yes'/'No' into 1/0 by hand. Scikit-learn can handle text labels, but setting the encoding yourself means you always know which class is "positive." That matters because precision and recall are measured against the positive class.
Second, the churn rate is 26.5%. So about 3 customers stay for every 1 who leaves. A model that says "nobody churns" would be 73.5% accurate and completely useless. That's why this series uses ROC AUC instead of accuracy for this dataset. Whenever you see an accuracy number, ask what the "predict the majority" baseline would score.
The identifier trap#
X still contains customerID. It looks like any other column, and dropping it can feel like throwing data away. But it's the most dangerous column in the table. Every customer has a different ID, so a flexible model can simply memorize which IDs churned. That works perfectly on the training data and tells you nothing about new customers.
The notebook tests this directly. It trains a decision tree on the customer ID alone, then checks it on customers it hasn't seen.
print(f'Unique customerID values : {churn["customerID"].nunique():,}')
print(f'Number of rows : {len(churn):,}')
print('Every row has its own ID :', churn['customerID'].nunique() == len(churn))
Unique customerID values : 7,043 Number of rows : 7,043 Every row has its own ID : True
# Turn the ID into a number so a tree can split on it, then train on that alone
id_only = churn['customerID'].astype('category').cat.codes.to_frame('customerID_code')
id_train, id_test, y_train_id, y_test_id = train_test_split(
id_only, y, test_size=0.2, random_state=RANDOM_STATE, stratify=y
)
id_tree = DecisionTreeClassifier(random_state=RANDOM_STATE).fit(id_train, y_train_id)
train_acc = accuracy_score(y_train_id, id_tree.predict(id_train))
test_acc = accuracy_score(y_test_id, id_tree.predict(id_test))
baseline = (y_test_id == 0).mean()
print(f'Accuracy on TRAINING data : {train_acc:.1%}')
print(f'Accuracy on TEST data : {test_acc:.1%}')
print(f'Predicting "nobody churns" : {baseline:.1%}')
Accuracy on TRAINING data : 100.0% Accuracy on TEST data : 61.3% Predicting "nobody churns" : 73.5%
Training accuracy is 100%. On new customers it drops to 61.3%. That's worse than the 73.5% you'd get by always guessing "no churn." The model didn't just fail to learn. It did worse than a free guess.
Here's why. A decision tree splits data using number thresholds. With a column that's different for every row, the tree can put almost every training customer in its own group, so training accuracy is perfect. But those splits are based on ID numbers, which say nothing about the customer. A new ID just lands in whichever group its number is near, and that has nothing to do with whether the person will leave.
You can even predict the 61.3%. Suppose a model's answers have nothing to do with the truth, but it still says "churn" about as often as churn really happens (26.5% of the time). Then it will be right about 0.735 × 0.735 + 0.265 × 0.265 ≈ 61% of the time. That's almost exactly what the tree scored. It's guessing at random, just in the right proportions.
The rule: use IDs to join tables and to trace predictions back to real customers. Don't train on them. So the notebook moves the ID out of the feature table and keeps it on the side.
# Keep the ID aside so we can trace predictions back to customers later,
# but keep it out of the feature matrix
customer_ids = churn['customerID']
X = churn.drop(columns=['Churn', 'customerID'])
print(f'Feature matrix is now {X.shape[0]:,} rows x {X.shape[1]} columns')
print()
print('Candidate features:')
print(list(X.columns))
Feature matrix is now 7,043 rows x 19 columns Candidate features: ['gender', 'SeniorCitizen', 'Partner', 'Dependents', 'tenure', 'PhoneService', 'MultipleLines', 'InternetService', 'OnlineSecurity', 'OnlineBackup', 'DeviceProtection', 'TechSupport', 'StreamingTV', 'StreamingMovies', 'Contract', 'PaperlessBilling', 'PaymentMethod', 'MonthlyCharges', 'TotalCharges']
Section 2Types of machine learning problems#
It depends on the target you pick, not on the dataset
The problem type decides how you measure success, which models you can use and which features are worth building. It's set by one choice: what you put in y.
- Regression predicts a number, like monthly revenue, delivery time or a house price. You measure error in the target's own units, using RMSE or MAE.
- Classification predicts a category, like churn or stay, fraud or not fraud, or one of five product tiers. You measure it with accuracy, precision, recall, F1 or ROC AUC.
- Clustering has no target. You give the algorithm only
X, and it groups similar rows. There's no right answer to check against, so you use measures like the silhouette score plus your own judgment about what the groups mean.
Here's the part that surprises people: a dataset isn't "a classification dataset." The problem type comes from the column you choose to predict. The same 7,043 rows below become all three types, just by changing y.
# CLASSIFICATION: predict a category (did the customer leave?)
y_classification = (churn['Churn'] == 'Yes').astype(int)
print('CLASSIFICATION target -> Churn')
print(f' dtype : {y_classification.dtype}')
print(f' unique values : {sorted(int(v) for v in y_classification.unique())}')
print(f' interpretation : a label from a fixed set')
print(f' scored with : accuracy, precision, recall, ROC AUC')
CLASSIFICATION target -> Churn dtype : int64 unique values : [0, 1] interpretation : a label from a fixed set scored with : accuracy, precision, recall, ROC AUC
# REGRESSION: same rows, different column -> predict a number
y_regression = churn['MonthlyCharges']
print('REGRESSION target -> MonthlyCharges')
print(f' dtype : {y_regression.dtype}')
print(f' unique values : {y_regression.nunique():,} distinct amounts')
print(f' range : ${y_regression.min():.2f} to ${y_regression.max():.2f}')
print(f' interpretation : any value on a continuous scale')
print(f' scored with : RMSE, MAE, R-squared')
REGRESSION target -> MonthlyCharges dtype : float64 unique values : 1,585 distinct amounts range : $18.25 to $118.75 interpretation : any value on a continuous scale scored with : RMSE, MAE, R-squared
# CLUSTERING: no target at all, just look for structure in X
print('CLUSTERING target -> none')
print(f' we pass only X ({X.shape[0]:,} x {X.shape[1]}) and ask for groups')
print(f' interpretation : discover customer segments nobody labelled')
print(f' scored with : silhouette score, plus whether the groups mean anything')
CLUSTERING target -> none we pass only X (7,043 x 19) and ask for groups interpretation : discover customer segments nobody labelled scored with : silhouette score, plus whether the groups mean anything
| Problem type | What y looks like | Scored with |
|---|---|---|
| Regression | A number on a continuous scale, here 1,585 distinct dollar amounts from $18.25 to $118.75 | RMSE, MAE, R² |
| Classification | A label from a fixed, known set, here just two: churned or stayed | Accuracy, precision, recall, ROC AUC |
| Clustering | Nothing. Only X is passed in, and the algorithm finds the groups | Silhouette score, plus human judgment |
Same rows, three different problems. This matters for feature engineering because a column can be a feature for one question and the target for another. When you predict churn, MonthlyCharges is a normal feature. When you predict MonthlyCharges, churn could be a feature, but only if you'd know the churn result before you needed the prediction. Usually you wouldn't. That's the "known at prediction time" rule again.
So the first question on any project isn't "which model should I use?" It's "what am I predicting, and what will I actually know when I need to predict it?"
Section 3Why feature engineering matters#
Tested twice: once where we know the answer, once where we don't
Data and features set the limit on how good a model can be. The algorithm only decides how close you get to that limit.
People repeat this line so often that it's easy to ignore. So instead of quoting it, this section tests it twice. First on made-up data, where we know the exact rule behind the labels. Then on the real churn data, where nobody knows the rule. The results look very different.
Made-up data: plots of land#
A property company calls a plot "large" if its area is over 9,000 square feet. The data has each plot's length and width, but not its area. So the model has to learn the rule length × width > 9000.
rng = np.random.default_rng(RANDOM_STATE)
n_plots = 3000
length = rng.uniform(20, 200, n_plots)
width = rng.uniform(20, 200, n_plots)
plot_area = length * width # the model is NOT given this
is_large = (plot_area > 9000).astype(int)
plots = pd.DataFrame({'length': length, 'width': width, 'area': plot_area, 'is_large': is_large})
print(f'Plots generated : {n_plots:,}')
print(f'Labelled large : {is_large.mean():.1%} (nicely balanced, so accuracy is readable)')
plots.head()
Plots generated : 3,000 Labelled large : 53.1% (nicely balanced, so accuracy is readable)
length width area is_large 0 159.312089 169.259805 26965.133141 1 1 98.998119 126.937278 12566.551738 1 2 174.547626 64.073854 11183.939164 1 3 145.526245 154.221500 22443.275846 1 4 36.951923 35.206561 1300.950126 0
def evaluate(features, target, label):
# Train logistic regression on `features` and report held-out performance.
f_train, f_test, t_train, t_test = train_test_split(
features, target, test_size=0.25, random_state=7, stratify=target
)
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000))
model.fit(f_train, t_train)
acc = accuracy_score(t_test, model.predict(f_test))
auc = roc_auc_score(t_test, model.predict_proba(f_test)[:, 1])
print(f'{label:<34} accuracy = {acc:.3f} ROC AUC = {auc:.3f}')
return acc
raw_acc = evaluate(plots[['length', 'width']], is_large, 'Raw features (length, width)')
eng_acc = evaluate(plots[['length', 'width', 'area']], is_large, 'With engineered feature (area)')
print()
print(f'Improvement: +{(eng_acc - raw_acc) * 100:.1f} accuracy points, same algorithm, same data.')
Raw features (length, width) accuracy = 0.913 ROC AUC = 0.986 With engineered feature (area) accuracy = 0.984 ROC AUC = 0.999 Improvement: +7.1 accuracy points, same algorithm, same data.
Both models used the same rows, the same scaling and the same logistic regression. The only difference is one extra column: length times width. No new information was added. Accuracy still went from 91.3% to 98.4%.
Why can't the model work out length times width by itself? Because of how logistic regression draws its decision line. It predicts "large" when a weighted sum of the inputs goes above zero:
w₁·length + w₂·width + b = 0
Whatever the weights are, that's a straight line on a length-vs-width chart. The real rule, length × width = 9000, is a curve. A straight line can't follow a curve, so some plots will always end up on the wrong side. The chart below shows this.
train_l, test_l, train_t, test_t = train_test_split(
plots[['length', 'width']], is_large, test_size=0.25, random_state=7, stratify=is_large
)
linear_model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000)).fit(train_l, train_t)
fig, axes = plt.subplots(1, 2, figsize=(14, 5.5))
# --- Left: raw feature space ---
ax = axes[0]
sample = plots.sample(1200, random_state=1)
ax.scatter(sample['length'], sample['width'], c=sample['is_large'],
cmap='coolwarm', s=12, alpha=0.6, edgecolors='none')
grid_l, grid_w = np.meshgrid(np.linspace(20, 200, 300), np.linspace(20, 200, 300))
grid = pd.DataFrame({'length': grid_l.ravel(), 'width': grid_w.ravel()})
# what the logistic regression actually learned
ax.contour(grid_l, grid_w,
linear_model.decision_function(grid).reshape(grid_l.shape),
levels=[0], colors='black', linewidths=2.5, linestyles='--')
# the rule that actually generates the labels
ax.contour(grid_l, grid_w, (grid_l * grid_w).reshape(grid_l.shape),
levels=[9000], colors='green', linewidths=2.5)
ax.set(xlabel='length', ylabel='width', title='Raw space: a straight line cannot trace a curve')
ax.plot([], [], 'k--', lw=2.5, label='learned boundary (straight)')
ax.plot([], [], 'g-', lw=2.5, label='true boundary: length x width = 9000')
ax.legend(loc='upper right', fontsize=9)
# --- Right: engineered feature ---
ax = axes[1]
for lab, colour, name in [(0, '#3b76af', 'small'), (1, '#c4453c', 'large')]:
ax.hist(plots.loc[plots['is_large'] == lab, 'area'], bins=60, alpha=0.65,
color=colour, label=name)
ax.axvline(9000, color='green', lw=2.5, label='threshold = 9000')
ax.set(xlabel='area = length x width', ylabel='number of plots',
title='Engineered space: one clean cut separates the classes')
ax.legend(fontsize=9)
plt.tight_layout()
plt.show()
In the left chart, the green curve is the true boundary and the dashed black line is the best straight line the model could find. The line can't bend, so it gets the plots in the corners wrong. Those errors are the 7.1 points of accuracy the first model lost. In the right chart, once area is its own column, one straight cut separates the two groups. Any linear model can do that.
So here's what this kind of feature engineering does: it doesn't add new information. It reshapes information you already have so the model's assumptions fit. (Other kinds of feature engineering do bring in new information, such as joining another table or using columns the model wasn't given. You'll see one of those later in this article.) A linear model can only draw straight boundaries. A good feature changes the inputs so that a straight boundary becomes the right answer.
A warning: not every combined feature helps#
It's easy to take the wrong lesson from this example and think ratios and products always help linear models. They don't. Take a similar rule: a house is "premium" if area / rooms > 400. It looks like the same kind of problem, so you might expect a ratio feature to help. A little algebra shows it won't:
area / rooms > 400 ⇔ area − 400·rooms > 0
The right side is just a weighted sum of area and rooms compared with zero. (Multiplying both sides by rooms is only allowed because rooms is always positive. A negative number would flip the inequality.) That's already a straight line, so logistic regression can learn it without any help.
# The same experiment, but with a ratio-threshold rule instead of a product rule
rng2 = np.random.default_rng(0)
n_houses = 3000
house_area = rng2.uniform(500, 3000, n_houses)
rooms = rng2.integers(2, 8, n_houses).astype(float)
is_premium = ((house_area / rooms) > 400).astype(int)
houses = pd.DataFrame({
'area': house_area,
'rooms': rooms,
'area_per_room': house_area / rooms,
})
print(f'Premium rate: {is_premium.mean():.1%}')
print()
a = evaluate(houses[['area', 'rooms']], is_premium, 'Raw features (area, rooms)')
b = evaluate(houses[['area', 'rooms', 'area_per_room']], is_premium, 'With ratio (area_per_room)')
print()
print(f'Difference: {(b - a) * 100:+.1f} accuracy points.')
Premium rate: 48.9% Raw features (area, rooms) accuracy = 0.992 ROC AUC = 1.000 With ratio (area_per_room) accuracy = 0.993 ROC AUC = 1.000 Difference: +0.1 accuracy points.
The gain is 0.1 points, which is basically nothing. The model could already draw that line, so giving it the ratio changed almost nothing. Putting both experiments together gives a rule of thumb you can reuse:
| Engineered feature | Does it help a linear model? | Why |
|---|---|---|
Multiplying two features (length × width) | Yes, when the true rule depends on the product | Creates a curved boundary a straight line can't draw |
A ratio compared with a fixed number (area / rooms > k) | Usually not | Rearranges into a weighted sum, which is already a straight line |
| Taking the log of a skewed column | Often, yes | Turns a curved, shrinking relationship into something closer to a straight line |
Adding a squared term (x²) | Often, yes | Lets a linear model fit a curve |
Before building a feature, ask: "Can my model already learn this?" The answer depends on the model. Tree-based models like gradient boosting can build curved boundaries out of many small splits, so they need combined features much less than linear models do. Feature engineering isn't a fixed checklist. It's about filling the gap between your data and what your chosen model can learn.
The same test on real data#
With made-up data we knew the rule, so we could build the perfect feature. Real data doesn't give you that. Don't guess at features. First look for a curved relationship. Plotting churn rate against tenure shows whether that relationship is straight or curved.
churn_num = churn.copy()
churn_num['TotalCharges'] = pd.to_numeric(churn_num['TotalCharges'], errors='coerce')
churn_num = churn_num[churn_num['tenure'] > 0].copy() # explained in section 4
y_num = (churn_num['Churn'] == 'Yes').astype(int)
buckets = pd.cut(churn_num['tenure'], [0, 3, 6, 12, 24, 48, 72])
by_bucket = churn_num.assign(churned=y_num).groupby(buckets, observed=True)['churned'].agg(['mean', 'size'])
by_bucket.columns = ['churn_rate', 'customers']
print(by_bucket.assign(churn_rate=lambda d: (d['churn_rate'] * 100).round(1)).to_string())
churn_rate customers tenure (0, 3] 56.8 1051 (3, 6] 44.6 419 (6, 12] 35.9 705 (12, 24] 28.7 1024 (24, 48] 20.4 1594 (48, 72] 9.5 2239
fig, ax = plt.subplots(figsize=(9, 4.5))
monthly = churn_num.assign(churned=y_num).groupby('tenure')['churned'].mean()
ax.plot(monthly.index, monthly.values * 100, marker='o', ms=3, lw=1.4, color='#c4453c')
ax.set(xlabel='tenure (months as a customer)', ylabel='churn rate (%)',
title='Churn risk falls steeply in the early months, then levels off')
ax.axhline(y_num.mean() * 100, color='grey', ls='--', lw=1.2)
ax.text(50, y_num.mean() * 100 + 2, f'overall average {y_num.mean():.1%}', color='grey', fontsize=9)
plt.tight_layout()
plt.show()
Churn is 56.8% in a customer's first three months, 35.9% between six months and a year, and 9.5% after four years. Look at how fast it falls. About 21 points drop in the first nine months, and only 26 more over the next five years. The drop is steep early and flat later. That's a curve, not a straight line.
A linear model given raw tenure has to fit one straight slope through that curve, so it fits neither the steep part nor the flat part well. The next cells test whether that costs anything. Since the curve looks like a decay, log(tenure) should straighten it. Adding tenure² gives the model a way to bend.
A small precision note: logistic regression draws its straight line on the log-odds scale, not on the churn rate itself. A curve like this one is still a curve on that scale, so the same reasoning applies. The notebook also adds one business feature: how many services each customer has. Right now that information is spread across nine columns.
service_cols = ['PhoneService', 'MultipleLines', 'InternetService', 'OnlineSecurity',
'OnlineBackup', 'DeviceProtection', 'TechSupport', 'StreamingTV',
'StreamingMovies']
# 'No', 'No phone service' and 'No internet service' all mean "not subscribed"
not_subscribed = ['No', 'No phone service', 'No internet service']
num_services = (~churn_num[service_cols].isin(not_subscribed)).sum(axis=1)
baseline_features = churn_num[['tenure', 'MonthlyCharges', 'TotalCharges']]
engineered_features = baseline_features.assign(
log_tenure=np.log1p(churn_num['tenure']), # straightens the decay curve
tenure_squared=churn_num['tenure'] ** 2, # lets the boundary bend
num_services=num_services, # 9 columns condensed into 1 count
)
print('Baseline :', list(baseline_features.columns))
print('Engineered :', list(engineered_features.columns))
print()
print('Services subscribed, distribution:')
print(num_services.value_counts().sort_index().to_string())
Baseline : ['tenure', 'MonthlyCharges', 'TotalCharges'] Engineered : ['tenure', 'MonthlyCharges', 'TotalCharges', 'log_tenure', 'tenure_squared', 'num_services'] Services subscribed, distribution: 1 1260 2 857 3 846 4 965 5 921 6 906 7 674 8 395 9 208
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=RANDOM_STATE)
def cross_validated_auc(features):
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000))
return cross_val_score(model, features, y_num, cv=cv, scoring='roc_auc')
baseline_scores = cross_validated_auc(baseline_features)
engineered_scores = cross_validated_auc(engineered_features)
print('ROC AUC per fold')
print(f' baseline : {np.round(baseline_scores, 4)}')
print(f' engineered : {np.round(engineered_scores, 4)}')
print()
print(f'Baseline mean AUC : {baseline_scores.mean():.4f} (+/- {baseline_scores.std():.4f})')
print(f'Engineered mean AUC : {engineered_scores.mean():.4f} (+/- {engineered_scores.std():.4f})')
print()
print(f'Gain : {engineered_scores.mean() - baseline_scores.mean():+.4f} AUC')
print(f'Folds improved: {(engineered_scores > baseline_scores).sum()} out of {cv.get_n_splits()}')
ROC AUC per fold baseline : [0.8175 0.8075 0.8041 0.8031 0.8111] engineered : [0.833 0.8262 0.8168 0.8236 0.8206] Baseline mean AUC : 0.8087 (+/- 0.0052) Engineered mean AUC : 0.8241 (+/- 0.0055) Gain : +0.0154 AUC Folds improved: 5 out of 5
The gain is 0.0154 AUC. That looks tiny next to the 7-point jump on made-up data, but it's real. It improved the score in all 5 folds. A lucky feature would win some folds and lose others. The gain is also about three times bigger than the normal fold-to-fold variation.
And no new data was collected, but be clear about where the gain came from. Two of the three new features reshape tenure, a column the model already had. The third, num_services, summarizes nine service columns the baseline model never saw, so it brings in information the model didn't have before. Notebook 9 splits the gain: the tenure features add about 0.0072 and num_services about 0.0082.
What does ROC AUC mean? Pick one customer who churned and one who stayed. AUC is the chance the model gives the churner a higher risk score. 0.50 is random guessing and 1.00 is perfect. So a gain of 0.0154 means about 1.5 percentage points more of all churner-stayer pairs are now ranked correctly.
This is what real feature engineering usually looks like: small gains. Kaggle competitions are often won by a few thousandths. Companies rebuild models for gains this size, because ranking customers slightly better across millions of subscribers saves a lot of money. Two habits made this comparison trustworthy, and they matter more than the features themselves: using cross-validation instead of one split, and running the baseline on the same folds.
Section 4Types of features in real data#
Label every column before you change it, because everything after depends on it
Every transformation later in the series depends on a feature's type. Scaling makes sense for numbers but not for a yes/no flag. One-hot encoding suits unordered categories but throws away the order in ordered ones. If you get the type wrong now, every later step inherits the mistake.
Numerical features are quantities where math makes sense. Continuous values can be anything in a range, like MonthlyCharges at $70.70. Discrete values are counts, like tenure in whole months or the num_services feature we built. You can't have 2.5 services.
Some columns sit on the line. tenure is time, which is continuous, but it's recorded in whole months, so it's stored as a count. The automatic classifier below labels it continuous only because it has more than 50 different values. Either label is fine, as long as you choose on purpose.
Categorical features are labels, where math makes no sense. Nominal categories have no order. PaymentMethod has four options, and none is "higher" than another. Ordinal categories do have an order. Contract goes month-to-month, then one year, then two year, which reflects how committed the customer is.
Binary features have two values, like Partner or PaperlessBilling. They're a type of category, but a single 0/1 column is all they need, so they're usually handled separately.
Text, dates and time series each need their own methods and get their own notebooks later. Free text and timestamps both have to be turned into numbers before a model can use them.
The most common mix-up is nominal vs ordinal, because pandas can't tell them apart. Both are just strings to the computer. Only a person who understands phone contracts knows a two-year contract is a bigger commitment than a monthly one. If you don't write that into your code, the model never gets that information.
def classify_column(series):
# Rough first pass. Domain knowledge still overrides every line of this.
n_unique = series.nunique()
if pd.api.types.is_numeric_dtype(series):
if n_unique == 2:
return 'binary (numeric)'
return 'numerical - discrete' if n_unique < 50 else 'numerical - continuous'
if n_unique == 2:
return 'binary (text)'
return 'categorical - nominal' if n_unique <= 15 else 'high cardinality / identifier'
inventory = pd.DataFrame({
'column': churn.columns,
'dtype': [str(churn[c].dtype) for c in churn.columns],
'n_unique': [churn[c].nunique() for c in churn.columns],
'example': [churn[c].iloc[0] for c in churn.columns],
'auto_guess': [classify_column(churn[c]) for c in churn.columns],
})
inventory
column dtype n_unique example auto_guess 0 customerID str 7043 7590-VHVEG high cardinality / identifier 1 gender str 2 Female binary (text) 2 SeniorCitizen int64 2 0 binary (numeric) 3 Partner str 2 Yes binary (text) 4 Dependents str 2 No binary (text) 5 tenure int64 73 1 numerical - continuous 6 PhoneService str 2 No binary (text) 7 MultipleLines str 3 No phone service categorical - nominal 8 InternetService str 3 DSL categorical - nominal 9 OnlineSecurity str 3 No categorical - nominal 10 OnlineBackup str 3 Yes categorical - nominal 11 DeviceProtection str 3 No categorical - nominal 12 TechSupport str 3 No categorical - nominal 13 StreamingTV str 3 No categorical - nominal 14 StreamingMovies str 3 No categorical - nominal 15 Contract str 3 Month-to-month categorical - nominal 16 PaperlessBilling str 2 Yes binary (text) 17 PaymentMethod str 4 Electronic check categorical - nominal 18 MonthlyCharges float64 1585 29.85 numerical - continuous 19 TotalCharges str 6531 29.85 high cardinality / identifier 20 Churn str 2 No binary (text)
This automatic guess is a good start, but look at where it goes wrong. TotalCharges is guessed as text with many unique values. It's actually money and should be a number. A function that only looks at dtypes and unique counts can't catch that. Contract is guessed as nominal, but it's ordinal, and only domain knowledge tells you that.
SeniorCitizen is correctly flagged as binary. Notice it's stored as 0/1, while every other yes/no column uses 'Yes'/'No'. Mixed formats like this are normal in real data. customerID is correctly flagged as an ID, which we already handled.
One more thing the classifier can't see: six service columns, such as OnlineSecurity, have a third value, "No internet service". That value just repeats what InternetService = 'No' already says. So those columns are really yes/no questions that only apply to customers with internet, not three separate categories. Knowing that affects how you encode them later.
The dtype problem, and why .info() was misleading#
Earlier, .info() said TotalCharges had 7,043 non-null values. That was true but misleading. The column is text, and a single space ' ' is a valid, non-null string to pandas. The missing values were there all along, hiding as blank text.
blank_mask = churn['TotalCharges'].astype(str).str.strip() == ''
print(f'Non-null count reported by info() : {churn["TotalCharges"].notna().sum():,}')
print(f'Values that are actually blank : {blank_mask.sum()}')
print()
print('The offending rows:')
churn.loc[blank_mask, ['customerID', 'tenure', 'MonthlyCharges', 'TotalCharges', 'Contract', 'Churn']]
Non-null count reported by info() : 7,043 Values that are actually blank : 11 The offending rows:
customerID tenure MonthlyCharges TotalCharges Contract Churn 488 4472-LVYGI 0 52.55 Two year No 753 3115-CZMZD 0 20.25 Two year No 936 5709-LVOEQ 0 80.85 Two year No 1082 4367-NUYAO 0 25.75 Two year No 1340 1371-DWPAZ 0 56.05 Two year No 3331 7644-OMVMY 0 19.85 Two year No 3826 3213-VVOLG 0 25.35 Two year No 4380 2520-SGTTA 0 20.00 Two year No 5218 2923-ARZLG 0 19.70 One year No 6670 4075-WKNIU 0 73.35 Two year No 6754 2775-SEFEE 0 61.90 Two year No
# Convert to numeric; anything unparseable becomes a genuine NaN
churn['TotalCharges'] = pd.to_numeric(churn['TotalCharges'], errors='coerce')
print(f'dtype after conversion : {churn["TotalCharges"].dtype}')
print(f'Genuine NaN values now : {churn["TotalCharges"].isna().sum()}')
print()
missing_rows = churn[churn['TotalCharges'].isna()]
print('Tenure values among the missing rows:', missing_rows['tenure'].unique().tolist())
print('Churn values among the missing rows :', missing_rows['Churn'].unique().tolist())
dtype after conversion : float64 Genuine NaN values now : 11 Tenure values among the missing rows: [0] Churn values among the missing rows : ['No']
All eleven of these rows have tenure = 0, and none of them churned. So these aren't broken records. They're new customers who haven't been billed yet. TotalCharges is blank because there's no total yet. The blank isn't random. It actually tells you something.
That changes the fix. If the values were missing at random, filling in the median would be fine. Here, the median would claim a brand-new customer has already paid about $1,397, which isn't true. Filling with 0 makes sense, since they've been charged nothing. Dropping the eleven rows also makes sense. They're only 0.16% of the data, none of them churned, and customers with zero tenure haven't been around long enough to leave.
before = len(churn)
churn = churn[churn['tenure'] > 0].copy()
print(f'Rows before : {before:,}')
print(f'Rows after : {len(churn):,} ({before - len(churn)} removed, {(before - len(churn)) / before:.2%})')
print(f'Remaining missing values across all columns: {churn.isna().sum().sum()}')
# X and y were built back in section 1, from the unfiltered 7,043-row table.
# Dropping rows did not update them, so they are now the wrong length.
print()
print(f'Stale y from section 1 : {len(y):,} rows')
print(f'Current churn table : {len(churn):,} rows')
print(f'Still aligned? {len(y) == len(churn)}')
# Rebuild everything derived from the table, including the IDs we set aside
y = (churn['Churn'] == 'Yes').astype(int)
X = churn.drop(columns=['Churn', 'customerID'])
customer_ids = churn['customerID']
print(f'After rebuilding : X={X.shape}, y={y.shape}, ids={customer_ids.shape}')
Rows before : 7,043 Rows after : 7,032 (11 removed, 0.16%) Remaining missing values across all columns: 0 Stale y from section 1 : 7,043 rows Current churn table : 7,032 rows Still aligned? False After rebuilding : X=(7032, 19), y=(7032,), ids=(7032,)
The mismatch check in that output is more important than the dropped rows. X and y were built in section 1 from the full 7,043 rows. Filtering churn didn't update them, so for a moment the features and target had different lengths. Scikit-learn catches this case and raises an error, which is lucky.
The worse case is silent. Pandas lines up data by index, not row position, so an outdated column can quietly create NaN rows or wrong labels with no error. The fix: whenever you add or remove rows, rebuild everything that came from that table.
Telling pandas about order#
Contract is the only ordered column in this dataset. As plain text, pandas treats its three values as unrelated. Making it an ordered Categorical keeps the order we know exists.
contract_order = ['Month-to-month', 'One year', 'Two year']
churn['Contract'] = pd.Categorical(churn['Contract'], categories=contract_order, ordered=True)
print('dtype :', churn['Contract'].dtype)
print()
# Comparisons now work, which is impossible with plain strings
two_year = churn.loc[churn['Contract'] == 'Two year', 'Contract'].iloc[0]
monthly = churn.loc[churn['Contract'] == 'Month-to-month', 'Contract'].iloc[0]
print(f'Is "Two year" > "Month-to-month"? {two_year > monthly}')
print()
print('Sorting now respects commitment level rather than the alphabet:')
print(churn['Contract'].value_counts().sort_index().to_string())
dtype : category Is "Two year" > "Month-to-month"? True Sorting now respects commitment level rather than the alphabet: Contract Month-to-month 3875 One year 1472 Two year 1685
churn_rate_by_contract = churn.assign(churned=(churn['Churn'] == 'Yes').astype(int)) \
.groupby('Contract', observed=True)['churned'].mean()
print('Churn rate by contract type:')
for contract, rate in churn_rate_by_contract.items():
bar = '#' * int(rate * 100)
print(f' {contract:<16} {rate:>6.1%} {bar}')
Churn rate by contract type: Month-to-month 42.7% ########################################## One year 11.3% ########### Two year 2.8% ##
The order matters. Churn falls from 42.7% to 11.3% to 2.8% as the contract gets longer, so contract order matches risk order.
That affects how we encode it in the next notebook. One-hot encoding would turn Contract into three separate 0/1 columns and lose the order. The model would have to learn from data alone that two-year is further from month-to-month than one-year is. Encoding it as 0, 1, 2 keeps the order in one column, so the model has less to learn.
But an ordinal code also makes an assumption: that the steps are equal, so one-year sits exactly halfway between the other two. For logistic regression, the scale that matters is log-odds. There the two steps are about 1.8 and 1.5, close enough that 0, 1, 2 works well. For another column with very uneven steps, one-hot might be the better choice. Check before you assume.
Doing the same with PaymentMethod would be wrong. Coding it as 0, 1, 2, 3 would invent an order that doesn't exist, as if "Mailed check" sat halfway between two other payment methods. The same method is right for one column and harmful for the next. Only your knowledge of the data tells you which is which, so write it into your code as soon as you know it.
A note on the category dtype#
Categories also save memory. A normal text column stores a separate string for every cell, even when the same few values repeat thousands of times. A category column stores each unique value once and uses small numbers to point to them. On 7,000 rows this hardly matters. On tens of millions of rows, it can decide whether your data fits in memory.
text_cols = [c for c in churn.columns if churn[c].dtype == object or str(churn[c].dtype) == 'str']
memory_before = churn.memory_usage(deep=True).sum()
churn_compact = churn.copy()
for col in text_cols:
if col != 'customerID':
churn_compact[col] = churn_compact[col].astype('category')
memory_after = churn_compact.memory_usage(deep=True).sum()
print(f'Text columns converted : {len([c for c in text_cols if c != "customerID"])}')
print(f'Memory as text/object : {memory_before / 1024**2:.2f} MB')
print(f'Memory as category : {memory_after / 1024**2:.2f} MB')
print(f'Reduction : {(1 - memory_after / memory_before):.1%}')
Text columns converted : 15 Memory as text/object : 6.15 MB Memory as category : 0.77 MB Reduction : 87.4%
That's about an 87% saving, with the same values in every cell. The more repeated values a column has, the bigger the saving. PaymentMethod has 7,032 values but only four unique ones, so storing four strings and pointing to them is much cheaper. customerID is the opposite. Every value is unique, so a category would add a lookup table and save nothing. That's why it was left out.
Section 5Where feature engineering fits in the ML pipeline#
The one rule that breaks more models than anything else
Feature engineering comes after you understand and clean your data, and before you train a model:
The diagram doesn't show the most important rule, so here it is:
Split your data before fitting any transformer. Every statistic a transformer learns must come from training rows only.
Here's why. A StandardScaler learns a mean and standard deviation from the data you give it. A SimpleImputer learns a median. A one-hot encoder learns the list of categories. If any of these learn from the full dataset before you split, your test rows have already shaped how the training rows are transformed. The test set isn't truly unseen anymore. Your score will look better than it really is, and you'll only find out when the model does worse on real new data.
This mistake is called data leakage, and you'll see the term throughout the series.
Most people agree with this and still get it wrong, because the leak doesn't show up anywhere. No error, no warning. So the next cell makes it visible.
X_full = churn.drop(columns=['Churn', 'customerID'])
y_full = (churn['Churn'] == 'Yes').astype(int)
X_train, X_test, y_train, y_test = train_test_split(
X_full, y_full, test_size=0.2, random_state=RANDOM_STATE, stratify=y_full
)
numeric_cols = ['tenure', 'MonthlyCharges', 'TotalCharges']
comparison = pd.DataFrame({
'full_dataset_mean': X_full[numeric_cols].mean(),
'training_only_mean': X_train[numeric_cols].mean(),
'full_dataset_std': X_full[numeric_cols].std(),
'training_only_std': X_train[numeric_cols].std(),
}).round(4)
comparison['mean_differs'] = comparison['full_dataset_mean'] != comparison['training_only_mean']
comparison
full_dataset_mean training_only_mean full_dataset_std training_only_std mean_differs tenure 32.4218 32.5623 24.5453 24.5424 True MonthlyCharges 64.7982 64.9993 30.0860 30.1086 True TotalCharges 2283.3004 2301.8395 2266.7714 2275.5861 True
Every row in that comparison is different. MonthlyCharges averages 64.7982 over the full dataset and 64.9993 over the training rows. The gap is small, but that doesn't matter. What matters is that it isn't zero. The full-dataset number includes the 1,407 test customers. A scaler fitted on it uses information from data the model should never have seen.
Two things follow. First, the leak hurts most when you can least afford it: small datasets, rare classes and target-based encodings all make it worse. Second, you can't predict how bad it will be. Here it's a few cents, and notebook 9 measures how little it matters for a scaler: about a millionth of an AUC point. With a target-mean encoding fitted on the full dataset, the same mistake can basically hand the model the answers, and the score will look great until real data arrives.
"Just be careful" doesn't work, because people forget under pressure. The real fix is a Pipeline, which only lets you do things in the right order.
numeric_features = ['tenure', 'MonthlyCharges', 'TotalCharges']
categorical_features = [c for c in X_full.columns if c not in numeric_features]
# We dropped the only missing values earlier, so this imputer has nothing to do
# on today's data. It stays in because data arriving next month might not be clean.
print(f'Missing values in X_train: {X_train.isna().sum().sum()} (the imputer is insurance, not a fix)')
numeric_pipeline = Pipeline([
('impute', SimpleImputer(strategy='median')), # median learned from train only
('scale', StandardScaler()), # mean/std learned from train only
])
preprocessor = ColumnTransformer([
('numeric', numeric_pipeline, numeric_features),
# handle_unknown='ignore' keeps unseen categories from crashing prediction
('categorical', OneHotEncoder(handle_unknown='ignore', drop='if_binary'), categorical_features),
])
model = Pipeline([
('preprocess', preprocessor),
('classifier', LogisticRegression(max_iter=2000)),
])
model.fit(X_train, y_train) # every statistic above is fitted here, on X_train alone
test_probabilities = model.predict_proba(X_test)[:, 1]
print(f'Test accuracy : {accuracy_score(y_test, model.predict(X_test)):.4f}')
print(f'Test ROC AUC : {roc_auc_score(y_test, test_probabilities):.4f}')
Missing values in X_train: 0 (the imputer is insurance, not a fix)
Test accuracy : 0.8045 Test ROC AUC : 0.8359
Here's what that cell does. A Pipeline chains steps together. When you call .fit(), it fits each step in order. When you later call .transform() or .predict(), it reuses those fitted steps without refitting. A ColumnTransformer applies different steps to different columns: the numeric steps only see numeric_features, the one-hot encoder only sees categorical_features, and the results are joined back together for the model. Putting a small pipeline inside a column transformer, inside a bigger pipeline, is a normal pattern.
Don't skip handle_unknown='ignore' on the one-hot encoder. Without it, a category that shows up in real use but never appeared in training (say, a new payment method) would crash the prediction. With it, the new category just gets all zeros for that feature, and the model keeps working.
Two notes on the output. This model uses all 19 columns, so its test AUC of 0.8359 can't be compared directly with the 0.8087 and 0.8241 from earlier, which used only 3 to 6 numeric columns and five folds instead of one test split. Its accuracy of 80.5% is best compared with the 73.5% you'd get by always predicting "no churn".
And one warning about the imputer "insurance". If a brand-new customer arrives with a blank TotalCharges, this imputer fills in the training median, about $1,410. That's exactly the wrong value section 4 warned about, since a new customer has been charged nothing. For a real deployment, fill TotalCharges with 0 when tenure is 0, for example with SimpleImputer(strategy='constant', fill_value=0) on that column. An imputer only helps if it fills in the right value.
# Proof that the transformers only ever saw training data
fitted_imputer = model.named_steps['preprocess'].named_transformers_['numeric'].named_steps['impute']
learned = pd.DataFrame({
'learned_from_train': fitted_imputer.statistics_,
'would_be_from_full_data': X_full[numeric_features].median().values,
}, index=numeric_features)
learned['leaked_if_wrong'] = learned['learned_from_train'] != learned['would_be_from_full_data']
print('Medians stored inside the fitted pipeline:')
print(learned.to_string())
print()
n_features_out = model.named_steps['preprocess'].transform(X_train).shape[1]
print(f'Columns going in : {X_train.shape[1]}')
print(f'Columns coming out: {n_features_out} (after one-hot encoding)')
Medians stored inside the fitted pipeline:
learned_from_train would_be_from_full_data leaked_if_wrong
tenure 29.00 29.000 False
MonthlyCharges 70.60 70.350 True
TotalCharges 1410.25 1397.475 True
Columns going in : 19
Columns coming out: 40 (after one-hot encoding)The medians stored in the fitted pipeline come from the training rows, not the full dataset. For MonthlyCharges and TotalCharges the two are different, and that difference is the leak we avoided. For tenure they happen to be the same, but that proves nothing. What makes the pipeline correct is which rows it learned from, not whether the numbers match. When model.predict(X_test) runs, it uses the training medians on the test rows. It never refits on test data.
A fitted Pipeline also helps in other ways. It's one object to save and deploy, containing the imputer, scaler, encoder and model together. Shipping them separately and applying them by hand in the right order is a common source of bugs.
It also works correctly with cross_val_score: each fold refits the whole chain on that fold's training part only. If you preprocess by hand first, every fold is contaminated. And it handles unseen categories through handle_unknown='ignore'. A later notebook goes deeper into pipelines. For now, remember: split first, fit the pipeline on the training data only, then use it to transform everything else.
Turning scores into something a team can use#
This is where the ID we set aside in section 1 becomes useful. It was useless as a model input, but you need it to hand results to people. A retention team can't call a row number.
# customer_ids was kept aside in section 1; the index still lines up with X_test
at_risk = pd.DataFrame({
'customerID': customer_ids.loc[X_test.index],
'churn_probability': test_probabilities,
'actually_churned': y_test.map({0: 'No', 1: 'Yes'}),
'contract': X_test['Contract'],
'tenure': X_test['tenure'],
}).sort_values('churn_probability', ascending=False)
print('Ten customers the model considers most likely to leave:')
at_risk.head(10)
Ten customers the model considers most likely to leave:
customerID churn_probability actually_churned contract tenure 3380 5178-LMXOP 0.847312 Yes Month-to-month 1 3159 5150-ITWWB 0.845825 No Month-to-month 3 2631 6861-XWTWQ 0.826722 Yes Month-to-month 7 4585 1069-XAIEM 0.822486 Yes Month-to-month 1 2797 6023-YEBUP 0.816192 Yes Month-to-month 3 3727 9057-SIHCH 0.816097 Yes Month-to-month 3 582 2865-TCHJW 0.815908 Yes Month-to-month 4 933 4750-ZRXIU 0.810636 Yes Month-to-month 4 6240 6521-YYTYI 0.794302 Yes Month-to-month 1 2516 8245-UMPYT 0.791587 Yes Month-to-month 16
All ten highest-risk customers are on month-to-month contracts, and nine of the ten really did churn. That matches the 42.7% month-to-month churn rate we saw in section 4, so the model is finding the same pattern, on customers it hasn't seen. All ten have also been customers for 16 months or less, most for 7 or less, which matches the steep early drop in the tenure chart. Note that one of the ten didn't churn: a high score means high risk, not certainty.
Two habits here are worth keeping. The ID stayed next to each prediction without ever going into the model, which is what lets someone act on the results. And we used predict_proba instead of predict. A probability lets the team work down a ranked list until their budget runs out. A plain 0/1 answer throws that ranking away.
Wrap-upKey things to remember from this notebook#
Features are what you know at prediction time, and the target is what you want to know. customerID scored 100% on training data and then did worse than always guessing "no churn". Memorizing IDs isn't learning, and only unseen data can show the difference.
The problem type comes from the target, not the dataset. The same 7,043 customers gave us a classification problem, a regression problem and a clustering problem, depending on what we put in y.
Reshaping a feature doesn't add information. It makes existing information usable. Multiplying length by width raised accuracy from 91.3% to 98.4% because the true boundary was a curve. A ratio compared with a fixed number barely helped, because the model could already draw that line. Always ask: "Can my model already learn this?"
Expect small gains on real data. On the churn data, reshaping tenure and counting nine unused service columns raised ROC AUC from 0.8087 to 0.8241. It won all five folds and is worth real money at scale, even though it's much smaller than the jump on made-up data.
Label every column before you touch it. TotalCharges was money stored as text, hiding eleven blanks that .info() counted as present. They turned out to be new customers who hadn't been billed, so filling them with the median would have been wrong, both when training and later when the model is in use.
Split before you fit. Every statistic a transformer learns must come from the training rows only. A Pipeline does this for you, so you don't have to remember it every time.
Notebook 2 goes deeper into cleaning: the different kinds of missing data, finding outliers and when to remove them, duplicate records, and messy category labels.
This article goes with Notebook 1 of a feature engineering series built on the Telco Customer Churn dataset. Every code cell above runs exactly as shown in the notebook, with its real output.