
Once sensors are installed and the dashboard is live, the next question is almost always the same: "can this system predict failures?" The honest answer β yes, but not automatically, and not in every case. Machine-learning anomaly detection delivers real value under certain conditions and wastes money under others. Telling the two apart early is the hardest part, and the one that saves the most budget.
Where threshold alarms run out
Threshold alarms are a correct and necessary foundation: if bearing temperature exceeds 80 Β°C, an alarm must sound. But thresholds carry three inherent limits:
- They fire after the fact. By the time a value genuinely crosses the limit, damage is often already underway.
- They are blind to combinations. Vibration slightly up, temperature slightly up, current slightly down β each still within limits, yet that combination is abnormal for this machine. Per-parameter thresholds never see the joint pattern.
- They ignore operating context. Vibration that is normal at 100% load may not be normal at 40%. A single threshold forces you to choose between nuisance alarms and late alarms.
This is where a data-driven approach comes in: instead of setting limits, the system learns what normal looks like for this machine across operating states, then flags departures from that pattern.
Prerequisites that get skipped
A model cannot learn what isn't in the data. Before discussing algorithms, four prerequisites decide whether the project succeeds:
- Sufficient history. Ideally several months covering normal variation: shift changes, product changeovers, seasons, and maintenance periods.
- Appropriate sampling rate. A 15-minute interval is fine for energy trends but far too sparse to capture mechanical vibration symptoms.
- Recorded operating context. Load, product type, shift state, machine mode. Without it, a model cannot separate "different because it's failing" from "different because we're running another variant".
- Event records. When the machine had problems and what maintenance was performed. This is what makes model output verifiable rather than merely believable.
If these are missing, the right first step is not building a model β it is fixing data collection.
A realistic staged approach
- Collect and clean. Bring sensor data and operating context into one consistent store, with timestamps aligned across sources.
- Build a statistical baseline. Before touching complex models, compute the normal range per operating mode. This step alone often reveals deviations that were hiding in plain sight β at a fraction of the cost.
- Apply anomaly detection. Once normal is well defined, models can flag parameter combinations that depart from historical patterns.
- Validate with the floor team. Every flagged anomaly must be judged by a technician: was something actually there, or is this normal variation? Skip this and the system quickly loses user trust.
- Only then discuss remaining useful life. Estimating when a component will fail requires a meaningful number of recorded failures. Most plants don't have that in year one β and that's fine.
Which cases fit, and which don't
Judging feasibility upfront is far cheaper than abandoning a project halfway. The distinguishing patterns:
- Good fit: rotating assets that run continuously. Pumps, compressors, blowers, and large motors produce stable vibration, current, and temperature signatures when healthy. Departures from that signature mean something, and data is plentiful because the machine rarely stops.
- Good fit: processes with interlinked parameters. Ovens, boilers, and cooling systems have relationships between temperature, pressure, flow, and energy draw. When that relationship shifts while every individual parameter still looks normal, something is usually moving.
- Poor fit: rarely operated machines. Equipment that runs a few hours a week takes a very long time to accumulate enough "normal" to learn from.
- Poor fit: processes that change every batch. If each order uses a different recipe and setpoints, "normal" becomes a moving target. A model is still possible, but it needs well-recorded batch context β which has to be fixed first.
- No model needed at all. If your biggest problem is machines failing because nobody was watching, well-designed threshold alarms already solve most of it, at a fraction of the cost.
Concluding that a case isn't ready yet is a legitimate and valuable outcome. What's expensive is spending six months building a model on data that could never support it.
The most expensive mistakes
Industrial analytics projects rarely fail on algorithm choice. The usual causes:
- Too many false alerts. A system that cries wolf daily is ignored within two weeks, and never trusted again afterwards.
- Output without explanation. Technicians need a reason β "vibration departs from the normal pattern at this load" β not a context-free score.
- No link to the workflow. Detection that doesn't generate a work order is just an interesting chart.
- Starting from the model, not the problem. Decide first which failure costs you the most, then build detection for that failure.
From data to decisions
The foundation stays the same: reliable, well-organised data. IncludeBox and IncludeGateways handle collection from machines and sensors, while the INCLUDE Smart Industry platform stores and presents it. On top of that, the AI & Machine Learning Development practice within INCLUDE services builds anomaly detection models fitted to your machines and processes β not a generic model forced onto them.
The order matters: measure first, understand second, predict last. Skipping the order is the fastest way to spend a budget without an operational result.
Sitting on sensor data you have not used?
Tell us about your machines and the data you already collect β the INCLUDE AI/ML team will assess whether anomaly detection is worth it.
Konsultasi Gratis via WhatsApp β See Services β