RTRadu Todea ▚ / Case study ← All write-ups

Case studySCADA AI Detection

A near-perfect score is a question.

My M.Sc. project: a neural network that reads network flows, flags the attacks among them, and shows which signals it leaned on. On data it had never seen, it scored 99.3%. The more useful exercise came after that: working out what that number does and doesn’t prove.

Type
M.Sc. project · 2025
Built with
Python · TensorFlow · scikit-learn
Data
CIC-IDS 2017 (public)
Focus
Detection · explainability

01What I built

From raw flows to a reasoned verdict.

Industrial control systems sit behind ordinary IT networks, and many attacks reach them the ordinary way: floods, scans and web exploits at the perimeter. I trained on CIC-IDS 2017, a public dataset of labelled traffic: a normal working Monday, and days carrying DDoS, port scans, web attacks and denial of service. Each row is one network flow. I kept 11 of its 79 features: how long the flow lasted, how many packets went each way, how big they were, and how fast they came.

In

About a million flows Broken values dropped, outliers clipped at the 99th percentile, scaled, and balanced to half attack, half benign.

80% to learn from, 20% held back

Model

Two stacked LSTMs 64 then 32 units, dropout between them, one output: attack or benign.

ten passes over the training data

Out

Why it decided Shuffle one feature at a time and measure how much accuracy falls: the model’s reliance, measured.

plus a dashboard to replay a capture

02What it found

216,293 flows it had never seen.

Accuracy
99.3%
Precision
99.2%
Recall
99.5%
ROC AUC
0.9997
The flow wasGot it wrongGot it right
Benign (108,246) 895 false alarms 107,351 let through
An attack (108,047) 587 missed 107,460 caught

The explanation step was the part I trusted most, because it’s the part that could have embarrassed the model. It leaned hardest on what came back: shuffling the total size of the reply packets cut its accuracy by 34 points, the number of reply packets by 16, the byte rate by 13. The request side mattered least: shuffling the total size of what was sent cost just 3. That reads sensibly: floods and scans get small replies, or none.

03What the score doesn’t say

Notes to myself.

Balanced data flatters
The test set was half attacks. Real networks are nearly all benign. A 0.83% false-alarm rate sounds small until one flow in a thousand is an attack: then the same model raises about eight false alarms for every real one it catches. A detector is judged on the alerts it hands an analyst, not on a balanced test set.
Random splits are generous
Flows were split at random, so near-identical flows from the same attack sat on both sides of the line. And the captures differ by day: Monday is normal traffic only, and each attack type comes from a day of its own, so some of what the model learned may be the day’s background rather than the attack. Scores measured that way run optimistic.
An LSTM needs a sequence
Each flow went in on its own, a sequence one step long, so the network’s memory never had anything to remember. It worked as a per-flow classifier. Where an LSTM earns its place is a window of consecutive flows from one host: the rhythm of a scan, not a single probe.
IT traffic isn’t ICS traffic
CIC-IDS 2017 is enterprise traffic: the attacks that hit the perimeter around a control system, not the industrial protocols inside it. Watching the plant floor means reading Modbus and DNP3 themselves: which function codes, which registers, how often.

04What I’d do next

Make the number earn it.

The second version, designed to be doubted, is a short list:

The weak spotWhy it bitesWhat I’d do
A balanced test set Hides what false alarms cost Report precision and alerts per day at a realistic attack rate, and set the threshold for that.
A random split Near-duplicates inflate the score Split by capture day, and hold out one attack type entirely to see what the model does with something new.
Scaling before the split The test data shaped the preprocessing Fit the clipping and scaling on training data only, and ship them with the model so new traffic is treated the same way.
One flow at a time No context, so no rhythm Feed windows of flows per host, and let the recurrent layers do the job they’re there for.
Attack or not A verdict, but not which attack A multi-class output, so the dashboard’s attack breakdown comes from the model; in the demo it’s illustrative.
The honest part 99% is where the work starts The score proved the pipeline works end to end. Whether it’s a detector worth deploying gets measured on messy, mostly-benign traffic from the kind of network it would guard.