Skip to content

Results

What changes for clients

Structured thinking shows up in ordinary numbers: shorter outages, fewer repeat incidents, decisions that survive the executive review, and less of everyone's time spent in rooms arguing about theories.

Four problems, and what actually caused them

Clients are not named. The details are as they happened.

Banking

Four days of "internet banking degraded"

A large Australian bank had made 27 changes over one weekend. By Monday, more than 60% of customers could not log in or use the service properly. Four days later the incident was still open. Changes had been backed out, which created new incidents, and every expert in the room had a theory worth chasing.

The words were the problem. "Internet banking degraded" has too many possible causes to investigate, so the team was picking theories rather than narrowing them.

We worked with them on a precise statement: the SSL handshake was not being received. That version had 3 possible causes. They backed out the most likely change and the service came back. Writing the statement took 90 minutes. Finding the cause took 45.

135 minutes

from a precise incident statement to a restored service, after 4 days of trial and error

Aviation

A global check-in glitch that hit only some passengers

A leading global airline had an intermittent fault in its check-in system. It crossed airports, flights and countries, but affected only a small number of passengers, many of them frequent flyers. Nothing in the system had changed.

The team worked through the obvious theories. Geography did not explain it. Flight type did not explain it. Frequent flyer status alone did not explain it either, because most frequent flyers were fine.

The pattern that held was in the passenger names. Travellers whose names contained the letters LOG hit the fault. Once that hypothesis was tested against other passengers, it held every time.

Cause found

after geography and flight-type theories were tested and eliminated

Banking

The ATM that rebooted for no reason

An ATM in a bank branch kept rebooting, sometimes several times a day, with no pattern anyone could predict. Engineers replaced the power supply, then other parts. It kept rebooting.

They swapped the machine with one from another branch. The replacement misbehaved in the same spot, and the original worked perfectly in its new home. That ruled out the ATM itself.

Sporadic faults usually come from something you do not control, so we pushed the team to look outside the machine. Behind the wall it backed onto, a specialist doctor had recently moved in, and the equipment in that office was being switched on through the day. That was the trigger.

Outside the box

the cause was not in the machine, or even in the room

Retail

Self-service checkouts that opened their own cash drawers

A supermarket had installed about 40 new self-service checkouts. Cash drawers started opening on their own, on unattended machines, with customers standing there. Only some checkouts were affected.

The sensible theories were tested first: a faulty batch of cash drawers, a firmware issue, a power spike. Each one was eliminated within a couple of days.

One staff member had suggested, hesitantly, that it was the flies. The drawers needed a 4-digit code, so it sounded absurd, but by then the structured analysis had ruled out everything else. Cameras confirmed it. Flies were landing on the touchpad and hitting the right sequence.

40 checkouts

and the cause was something nobody in the room wanted to say out loud

In their words

"Damn you Andrew! Wish I saw this stuff 20 years ago, I probably would have retired by now!"

Problem Manager

ANZ Bank

"I wanted my whole division to have a common company wide problem solving approach and got much more than I'd bargained for. We also reduced the trial and error costs in most of our divisions."

Liam Edwards

Past Executive Director, global investment bank

"The KF training introduced us to the concept of Technical Cause versus Root Cause and that unique insight made all the difference in all our Incident & Root Cause Analysis restorations."

David Pryde

EVP Product Support, SGX

Want a result like these?

Tell us what is going wrong and we will tell you whether the method fits.

Tell us about your situation or pick a time that suits