Skip to main content

Command Palette

Search for a command to run...

My Final Fraud Model (Day 7)

Updated
•5 min read•View as Markdown
F
Learning cloud and data science from zero. Writing what I learn so it sticks; the wins, the confusion, and everything in between.

Day 1 was setup. Day 2 through 6 was the journey. Day 7 was the test set.

I had not touched it since Day 3. That was the point. The test set is the sealed envelope. The one number you only open at the end.

The problem it solves

For six days I had been making decisions on the validation set:

  • Day 5: choosing XGBoost over logistic regression

  • Day 6: choosing threshold 0.75 over 0.5 or 0.96

Every one of those decisions was informed by validation scores. That means my model had been tuned to validation. The validation number was no longer trustworthy as a final grade. It had been "used."

The test set exists for this exact reason. It is untouched data. No decisions were made on it. Whatever it says is real.

How it works

The test set was created in Day 3 with the same stratified split. I loaded it, checked it matched the training columns, and ran the final evaluation.

X_test = pd.read_csv('splits/X_test.csv')
y_test = pd.read_csv('splits/y_test.csv').squeeze()

print("X_test shape:", X_test.shape)
print("Fraud count in test:", y_test.sum(), "/", len(y_test))
print("Fraud rate:", y_test.mean())
print("Columns match train:", X_test.columns.equals(X_train.columns))
X_test shape: (42722, 30)
Fraud count in test: 74 / 42722
Fraud rate: 0.0017321286456626562
Columns match train: True

74 frauds. 0.173%. Same rate as train and validation. stratify did its job all the way through.

Then the final scoring:

y_test_proba = xgb.predict_proba(X_test)[:, 1]
y_test_pred  = (y_test_proba >= 0.75).astype(int)

Threshold 0.75, the one I chose in Day 6.

example output

Test AUC: 0.9571

Confusion matrix (TN, FP / FN, TP):
[[42635    13]
 [   15    59]]

Classification report:
              precision    recall  f1-score
           0     0.9996    0.9997    0.9997
           1     0.8194    0.7973    0.8082
Metric Validation (Day 5) Test (Day 7)
AUC 0.9765 0.9571
Fraud caught 60 / 74 59 / 74
Fraud missed 14 15
False alarms 45 13
Fraud precision 57.1% 81.9%
Fraud recall 81.1% 79.7%
F1-score 0.6704 0.8082

What I chose, and why

I chose a 0.75 threshold on Day 6 because, in fraud detection, the cost of missing fraud is higher than the cost of a false alarm. Missing a fraudulent transaction means losing money and potentially damaging customer trust. A false alarm, on the other hand, only requires the customer to complete an extra verification step. Given that trade-off, I would rather inconvenience 13 customers than miss 15 fraudulent transactions.

On the test set, this threshold caught 59 of 74 fraudulent transactions, while generating 13 false alarms. That seems like a manageable number of alerts for an investigation team to handle.

But there is an uncomfortable takeaway: the test AUC was 0.9571, exactly the same score logistic regression achieved on the Day 4 validation set.

The advantage XGBoost appeared to have over logistic regression during validation (0.9765 vs. 0.9571) which did not carry over to the test set. On test, they ended up with the same AUC.

What confused me at first

The drop in AUC initially confused me, but it made sense once I understood the limitations of the validation set.

There were only 74 fraud cases in the validation set, which makes AUC a relatively noisy estimate. A difference of 0.02 could easily be due to random variation rather than a genuine difference in model performance. Validation was useful for confirming that “XGBoost is not broken,” but it was not reliable enough to declare XGBoost the clear winner.

What mattered more in practice was the threshold. XGBoost and logistic regression performed similarly when it came to ranking transactions by fraud risk. The bigger difference came from where I chose to set the cut-off for triggering an alert.

Model choice was the smaller decision. Threshold choice was the bigger one.

Key takeaways

  • The test set is the only honest measure of performance. Validation helps guide model decisions, but the test set tells you how the final model actually performs on unseen data.

  • With only 74 fraud cases, AUC is noisy. A 0.02 difference can easily come from random variation, so small validation gains should not be overinterpreted.

  • The threshold mattered more than the model. XGBoost and logistic regression performed similarly in terms of ranking. The bigger impact came from deciding where to set the threshold.

  • The real question is the cost ratio. I prioritized recall because missing fraud is more costly than triggering a false alarm. However, I have not yet quantified exactly how much more costly it is. That is the next step.

What did not work

The XGBoost advantage from Day 5 did not hold. On validation, XGBoost looked 0.02 AUC better than logistic regression. On test, it landed at 0.9571, exactly what logistic got on validation.

That does not mean XGBoost was a waste. It does mean the validation edge was partly noise. With 74 frauds, everything is partly noise.

The end of the series

Seven days. One project. A working fraud detection model with a clear trade-off and honest evaluation.

If I learned one thing: in extreme imbalance, the honest number is the one from the data you never touched. And the real decision is not the model. It is where you set the threshold, and that depends on a cost ratio you probably have not calculated yet.

That is the next project.

The full project
Code, notebook, and everything from all 7 days:
github.com/Mwende-Fifi/fraud-detection

Fraud Project Log

Part 7 of 7

Following my credit card fraud detection project from absolute zero. Each post is one day of the build ; setup, data exploration, model training, and everything that breaks along the way. Written for anyone else starting from scratch.

Start from the beginning

How I Pulled a Kaggle Dataset Into Colab (Day 1 of My Fraud Project)

I'm building a credit card fraud detection project from scratch. But before I could touch any data, I had a smaller problem: how do I actually get a Kaggle dataset into my notebook? My old method was