Enterprise AI · Asset decisions

Eleven weeks of vibration, five weeks to the outage, one call to make

One decision, worked end to end on an IBM Maximo and MAS estate: the signals that went in, the options they produced, the engineer who decided, and the record left behind.

Reliability engineer reviewing pump condition evidence alongside Maximo work history before an intervention decision

The case file

The asset, the week the trend was raised

Most operators already hold this much: the condition data exists, the history exists, the criticality is known. What is usually missing is the machinery that turns it into an action somebody owns.

  1. Asset and criticality

    Raw water pump, 400 kW, one of three units feeding a water treatment works. High criticality: one unit down constrains raw water into the works, so the failure is felt at the site boundary rather than in the pump house.

  2. Signals in

    Drive-end bearing vibration from Monitor, rising gradually across eleven weeks. No single reading breaches an alarm. Two prior years of stable readings on the same instrument.

  3. History in

    Maximo work history: coupling replaced seven months earlier, two short-notice restarts in the same period. Health scoring places the unit high on consequence.

  4. The window

    Planned outage in five weeks. The next one after that falls in the following maintenance year.

  5. The decision out

    Intervene now on an unplanned basis, or hold to the planned outage and accept the intervening risk. Both are decisions; only one of them usually gets recorded.

  6. Accountable for it

    The site reliability engineer, with the planner present. Named before the first signal was wired to anybody, which is organisational work no software performs.

Signal, person, system of record

How the eleven weeks ran

Five beats. Note where the person falls: not at the end as a rubber stamp, but at the point where the evidence has been assembled and has stopped being able to answer the question on its own.

  1. Weeks 1 to 11

    Vibration rises gradually. No step change, no alarm on any single reading, but a direction of travel outside the pattern the same pump held for two years.

    Written Nothing yet. On its own this is an observation, not a case for action.

  2. Week 11

    Condition trend, coupling history, the two restarts, criticality and the outage calendar are assembled into one view rather than sitting in two systems.

    Written A case raised against the asset, with the trend and the work history attached to it.

  3. Week 11, day 3

    The reliability engineer reads the trend against the coupling history, walks the pump, and holds the intervention to the planned outage five weeks out.

    Written A deferral recorded against the asset: the reasoning, the threshold that would overturn it, her name, the date.

  4. Week 11, day 3

    Two follow-on actions, so the decision is not a single word. An interim check with a stated threshold, and parts ordered against the window.

    Written A work order for a two-week vibration check carrying the threshold, the bearing set on order against the outage, and the alignment question raised as its own job.

  5. Week 16, outage

    The bearing is replaced and the alignment checked and corrected. As-found condition is consistent with the trend, which is what earns the capability its next decision.

    Written Closure against the deferral, so the engineer who made the call can explain it from the record rather than from memory.

A trend flagged into a weekly review pack gets four minutes among eleven other items and is carried forward, because nothing in the process forces it into a dated decision. The difference here is beat three.

Stated in the recommendation itself

Three things the evidence did not establish

Written down at the same time as what it did establish. Condition evidence read as a verdict rather than as one input among several is a governance failure rather than a modelling one.

Which bearing fails, and when
The trend establishes direction, and with enough history a rough probability. Decisions were taken on consequence and outage timing.
Often read as A failure date. Presenting one sets the capability up to be discredited by its first miss.
Bearing wear against residual misalignment
Both are consistent with the signature after coupling work, and the evidence could not separate them. That is why alignment was raised as its own job rather than assumed away inside the bearing replacement.
Often read as A cause, settled. The ambiguity was carried into the work rather than resolved on paper.
Whether five more weeks was safe
Nobody could say whether the unit would run five more weeks or five more months. An unknown is managed with a tripwire, which is what the two-week check with a stated threshold is.
Often read as A confident number. The interim check exists because the duration was unknown.

The record left behind

One decision, four fields, and the question it has to survive

The decision, as recorded

Decided by
The site reliability engineer, with the planner present. One named owner, not a committee position.
Decision
Hold to the planned outage. Interim vibration check at two weeks against a stated threshold. Bearing set ordered now.
Recorded as
A deliberate risk acceptance against the asset, with the reasoning, the threshold and the owner name attached.
Retained for
Later review, whichever way the pump behaved. Approve, defer and override are recorded identically.

What the record has to survive

The question afterwards
Why was the pump left running? The record answers it in four fields, from Monitor evidence and Manage history.
Left deliberately open
The alignment question, as a separate job, so it survives closure of the bearing work.
The failure mode designed against
Not an imperfect model. An unattributable decision.
Illustrative
Illustrative. The pump, timings and readings follow a typical engagement shape.

The deferral is the entry that matters most, because it is the one an investigation asks about.

Scope and boundaries

Where condition insight depends on things we do not control

Three limits worth agreeing before a first asset class is scoped, because each one changes what a realistic pilot looks like.

It will not give you a failure date

A trend establishes direction and, with enough history, a rough probability. It does not establish when a specific bearing on a specific pump will fail. Decisions are made on consequence and outage timing, and a capability sold on predicted dates is discredited by its first miss.

It will not create the authority to defer

Where nobody is clearly permitted to accept the risk of holding an intervention, better evidence produces better-informed indecision. Naming the owner and the escalation threshold happens before the first signal is wired to anybody.

It will not compensate for thin work history

Condition evidence is read against what has been done to the asset before. Where history is incomplete or spread across sites that recorded things differently, the trend cannot be interpreted with confidence and decision-ready data is the honest sequence.

Condition insight, frequently asked questions

What is condition insight in practical terms?
The discipline of turning a trend into a dated decision with a named owner, and recording what the evidence did not establish alongside what it did. The inputs are operating signals, maintenance history and asset criticality; the output is an intervention that somebody owns.
Does this replace planner and engineer judgement?
No. Named owners still decide what to execute, defer or escalate, and the deferral is recorded as a decision rather than as silence. If a page ever tells you a model decided an intervention on a critical asset, ask who signed it.
Where does IBM MAS fit?
Monitor, Predict and Health are the MAS components behind this capability. We implement them in a sequence set by the maturity of your Manage data, which advanced analytics sets out rung by rung.
What happens when the condition evidence turns out to have been wrong?
Because each recommendation states what it does not establish, and each decision records the reasoning and the owner, a wrong trend produces a reviewable decision rather than a loss of confidence in the capability.
How much condition data do we need before this is worth starting?
Less than most people assume for one asset class, and more than most have for a whole estate. The binding constraint is usually work and failure history rather than sensor coverage, because a trend without history behind it cannot be read. Where history is thin, decision-ready data is the honest first step.
Can this work in public-sector and safety-critical estates?
Where approval paths and evidence retention are clear, yes. Governance and assurance lists the questions an auditor or public-sector buyer asks, and the artefact that answers each.

Bring one asset you argued about last quarter.

Not a strategy conversation about predictive maintenance. One asset where the condition evidence and the intervention decision did not line up, and we walk it the way this page walks the pump: what the evidence established, what it did not, who should have owned the call, and what the record would need to contain for that call to stand up now.

Bring this to the first call

  • One asset where an intervention was taken late, or deferred and later questioned
  • The condition data you hold on it, however partial, and how far back the work history goes
  • The name of the role permitted to defer an intervention on that asset class
  • Your outage calendar for that asset class over the next two quarters