Skip to content

Operational decay

Keep useful systems useful.

Useful systems get harder to run as people, processes and requirements change. Recognising that drift early helps you decide what to keep, what to improve and what to replace.

How it develops

Each stage feels manageable in isolation.
  1. The build

    A team builds something to specification. It works. The demo is successful. Everyone is satisfied. The build team considers it done.

  2. The handover

    The system is handed over to the operation. Documentation is written but incomplete. Training happens once. The build team moves on to the next project.

  3. The drift

    Six months later, the system has drifted from its documented state. Workarounds have accumulated. The person who understood how everything connected has changed roles. Nobody is quite sure who owns it.

  4. The exposure

    Something breaks. Or someone asks a question the operation can't answer. Or the organisation tries to scale the system and discovers its architecture was never designed for that. The cost becomes visible all at once.

Where it shows up

Operational decay is medium-agnostic.

It happens in hardware deployments, software platforms and automated workflows. The manifestation is different. The root cause is the same.

Hardware

Installed, not integrated

  • Field devices calibrated once and never audited
  • Firmware versions diverging across sites with no tracking
  • No documented procedure for when a remote unit goes offline
  • Maintenance knowledge held by one person who is not always available
Software

Shipped, not sustained

  • A platform nobody wants to update in case they break it
  • Dependencies that are undocumented until they cause an outage
  • Manual steps that were meant to be temporary and became permanent
  • One engineer who understands the architecture and has been meaning to document it
Automation

Triggered, not trusted

  • Workflows built around one person's mental model of the operation
  • No runbook for unexpected behaviour
  • Processes that run correctly but that nobody can explain end-to-end
  • Changes that require the original author because nobody else is confident

Self-assessment

Is a system you depend on already drifting?

Six questions about one system you rely on. If more than two land, the decay is already priced into your operation — you just haven’t been invoiced for it yet.

  1. Only one person can safely change this system.
  2. The documentation hasn't been updated since handover.
  3. Nobody can say which firmware or version is running where.
  4. There is no runbook for the failure you've already seen once.
  5. A workaround introduced 'temporarily' is now part of the process.
  6. Scaling it would mean rebuilding it.

The alternative

Decay is preventable — and recoverable.

If you’re building, ownership belongs in the design. If you’ve already inherited something fragile, the first move is understanding what you have — not a rebuild.

Book a Fit CallRead the ownership standard

We’re honest about when rescue isn’t the answer.