An automated line running unattended in a bright factory, a robotic arm mid-motion and a single status light showing green, because nobody had to be called

Product · Operations

A register of the conditions that recur, and what each one may trigger.

Autoheal runs the agreed steps for an agreed condition, verifies that the condition cleared, and writes the record while it happens. Six entries follow: four it acts on, and two where a person is woken every time.

The register

Six recurring conditions, and where each one stops

Each entry names the condition, what the playbook may do, the ceiling it stops at, and whether a person is still woken. Entries B1 and B2 are conditions Autoheal refuses to act on.

A1

Overnight preventive maintenance generation hung after the patch window

Automated, narrowly
What the playbook does
The five runbook steps engineers run at 03:00: clear the hung task, restart the job, watch it complete.
Where the ceiling is
One attempt, inside the monthly patch window.
Is a person woken
Only if verification fails

Ninety minutes by hand, four of them the fix. About five minutes by playbook. Illustrative times.

A2

Outbound integration queue that has stopped draining

Automated where the retry is agreed
What the playbook does
The agreed retry sequence for that named flow, on that environment only.
Where the ceiling is
The granted attempt count, and nothing outside the named flow.
Is a person woken
If the backlog is still static after the granted attempts
A3

Certificate approaching expiry on a named endpoint

Automated only where your policy permits it
What the playbook does
Where policy allows machine rotation, the agreed steps run and the new certificate is confirmed. Where it does not, the playbook gathers state and stops.
Where the ceiling is
Whatever your policy says, which is often the evidence and no further.
Is a person woken
Always, where policy reserves credential changes to a person
A4

Scheduled task that did not start inside its window

Automated, narrowly
What the playbook does
The agreed start sequence for that named task on that named environment.
Where the ceiling is
Inside the window only. A closed window is a decision about the day ahead.
Is a person woken
If the task will not start, or the window has already closed
B1

The same task hung, outside the agreed window, with an unfamiliar session count

No playbook applies
What the playbook does
Nothing. Two attributes fall outside the approved condition, so Autoheal stands down.
Where the ceiling is
Acting on a half-recognised condition destroys the state an engineer needs.
Is a person woken
Always
B2

Anything touching business data, anything irreversible, anything turning on judgement

Never a candidate
What the playbook does
None, and we say so during scoping. These belong with Sentinel, Diagnose and an engineer.
Where the ceiling is
Absolute, whatever else the register grows to hold.
Is a person woken
Always

Every event reaches the incident record, stand-downs included. Your register carries your own conditions.

Entry B1, on one night

02:51, the same condition, declined

Two attributes decide that this page is not entry A1.

  1. 02:51

    The overnight task is hung on the same environment as entry A1, and the page looks identical.

    Written The signal, and the entry it resembles.

  2. 02:51

    The patch window closed at 02:00, and the database session count sits above the agreed range.

    Written Both mismatched attributes, by name.

  3. 02:52

    Autoheal acts on nothing and pages on-call with the state it gathered.

    Written The stand-down, and the reason no playbook applied.

  4. 03:04

    The engineer takes it to Diagnose rather than the A1 runbook.

    Written The human decision, on the same incident record.

In early rollout the stand-down is the most common outcome. A register scoped too widely produces fewer of them.

Getting an entry onto the register

Agree, bound, match, verify, record

Automated remediation has to be narrower than an engineer. Each stage below is signed off before the next one runs.

  1. 01

    Agree the condition

    Named signal, environment, threshold and window, approved before anything runs unattended.

    Owner Your operations lead, with MaxIron Typically 1 to 2 weeks per entry

  2. 02

    Agree the action and its ceiling

    The steps your engineers already run by hand, the attempts granted, and the point at which the playbook stops and pages.

    Owner Your on-call engineers Typically Same window

  3. 03

    Match, or stand down

    Every live event is compared with the agreed condition. Anything unrecognised pages a person.

    Owner Autoheal Typically Seconds

  4. 04

    Act, then verify

    The steps run, then the condition is re-tested. A failure pages with the state before and after.

    Owner Autoheal Typically Minutes

  5. 05

    Record it, then widen it on evidence

    What fired, why it matched, the steps and the outcome go into the incident record at the time. Widening a playbook goes through Change Control.

    Owner Your change process Typically Per review cycle

Run this on one condition

Five statements about the page your team saw most often

Take the out-of-hours page your team handled most often last quarter and stop at the first statement that is false.

  1. 1

    The condition can be named as a signal on a named environment, with a threshold and a window.

  2. 2

    The steps your engineers run are the same every time, in the same order.

  3. 3

    A check proves the condition cleared, and reading it needs no judgement.

  4. 4

    Nothing in those steps touches business data or takes an irreversible action.

  5. 5

    Your change policy permits a machine to take those steps on that environment.

If you stopped early

The condition stays with a person, and the work is describing it precisely enough to reach statement five. Most start here.

If every statement held

It is a candidate for a register entry, with the condition, ceiling and verification written during scoping.

Scope and boundaries

Where Autoheal stops and a person takes over

The three limits that make unattended remediation acceptable in production.

It acts only on conditions agreed in writing

No learning on the fly, and no acting on a near match. A playbook is a written agreement about one condition, one action, one ceiling and one verification.

It does not work out what is wrong

Resolving repeat failures across Maximo and MAS estates MaxIron operates, closing the loop from Sentinel and Diagnose, with approvals through Change Control and Cloud Manager on managed hosting. An unexplained event is not a candidate for a register entry: detection belongs to Sentinel, cause to Diagnose, and the decision to an engineer.

Some conditions stay with a person permanently

Anything touching business data, anything with an irreversible step, and anything where the safe action turns on judgement. Higher-impact playbooks request approval through Change Control before they run.

MaxIron Autoheal, frequently asked questions

Which events does Autoheal act on?
Hung scheduled tasks, integration queue backlogs, certificate expiries and agreed retries. Each is a register entry with its own condition, ceiling and verification.
How many of our incidents will be automatable?
A minority at first. Start with the three conditions your team handled most often last quarter, where the steps are stable and the verification is unambiguous.
What happens if an automated action makes something worse?
The playbook takes only its granted attempts, touches only what its scope names, and must prove the condition cleared. Otherwise it stops, reverses its own action where the runbook defines a reversal, and pages a person with the state before and after.
Can we add our own playbooks?
Yes, alongside the curated ones. Playbooks are versioned, reviewed, and approved through your change process before they run unattended.
Do we need Sentinel and Diagnose?
Sentinel and Diagnose supply what to act on and why. Autoheal also acts on signals from existing monitoring where those signals name the condition precisely enough.

Bring three months of out-of-hours pages.

We sort them into conditions with stable steps, conditions that need Diagnose first, and conditions that stay with a person. Then we draft the condition, the ceiling and the verification for the first candidate as an entry on your register.

Bring this to the first call

  • Three months of out-of-hours pages, with the time each one landed
  • The runbook your engineers reach for at 03:00, however stale the page is
  • The actions your change policy will never allow a machine to take
  • Who signs off a playbook, and the evidence they need in order to sign