At Least Once
Incident 2041-0332 — Post-incident review (blameless)
Service: Next of Kin Notification (NKN)
Severity: 2
Author: R. Whitfield, engineering lead, on behalf of the on-call team
Status: Closed
Summary
On the morning of 4 March, the NKN service informed Carol Ellison of the death of her father, John Ellison, twice: once at 06:12 by push notification, and once at 23:19 the following day by telephone, forty-one hours later. Both notifications were correct. Both were delivered as designed. Mrs Ellison experienced the second as she experienced the first — as news.
No component failed. That is why this review exists.
Background
NKN exists because hospitals cannot ring every next of kin at the moment of death, and because before it existed, some people were never rung at all. The service receives a verified death event from the hospital or registrar, resolves the registered next of kin, and delivers notification through an escalating sequence of channels: application push first, then telephone contact by a trained caseworker. Delivery is repeated until the service receives a positive acknowledgment from a human being.
The escalation policy is not an accident of enthusiasm. It is the corrective from incident 2039-0117, which some of you will remember and which the rest of you should read. In 0117, the service sent one notification, received a delivery receipt from the push provider, marked the case complete, and stopped. The recipient’s phone was at the bottom of the Regent’s Canal. Michael Osei learned of his wife’s death nine days later, from a letter cancelling her direct debits. The 0117 review made one finding: the service conflated delivery with receipt. Its principal action item became policy: a delivery receipt is machine evidence that a message reached a device. It is not evidence that a person has been told. Only a positive human acknowledgment closes a case, and until one arrives, the service escalates.
This policy is correct. It caused this incident.
Timeline
All times GMT. Times marked † were established after the fact, from Mrs Ellison’s own account.
03:41 — John Ellison, 84, dies on Ward 9 of St Aldhelm’s Hospital. His daughter had been at the ward for six nights. She had gone home on the seventh.
05:52 — Verification of death is recorded in the hospital system. The NKN event is emitted at 06:04.
06:12 — Push notification delivered to Mrs Ellison’s phone. Per script v3.1, the payload does not announce the death; it reads: Please contact us urgently regarding John Ellison, with a number and an in-app link. The sending identity, displayed above the payload on the lock screen, is HM Bereavement Notification Service.
06:13† — Mrs Ellison, awake, reads her lock screen. She does not open the app. She does not ring the number. She later told the caseworker: “There was nothing I needed to ask.”
08:20† — She travels to her sister’s house in Frome. At some point that morning her phone is switched off. It remains off.
06:12 +24h — No acknowledgment received. The case enters the caseworker escalation queue, as designed. Three calls are attempted over the following twelve hours. All divert to voicemail. Per policy, no message content is left — a death is not announced to an answering machine — only a request to call back.
From the on-call channel, during this window:
jt — escalation picked up ellison at t+24 as designed. provider says delivered 06:12. no open event, no ack, calls going to voicemail
rw — delivered isn’t told. that’s the whole of 0117
jt — right. so by the book she doesn’t know
rw — by the book. the sender name says Bereavement Service and it sat on her lock screen. she’s his daughter
jt — we can’t act on what she probably saw
rw — no. we can’t
The on-call engineer was correct. So was the policy. The case remained open.
23:19 (+41h) — A caseworker reaches Mrs Ellison, whose phone is back on, at her sister’s kitchen table. Script v3.1 is an informing script. It opens: “I’m very sorry to have to tell you that your father, John Ellison, died in the early hours of Tuesday morning.”
From the caseworker’s contact log, verbatim: Recipient silent approx. 10 seconds. Then: “I know. I’ve known since Tuesday. It was on my phone before it was light. Why are you telling me as if it’s news?” Apologised. Completed acknowledgment workflow. Recipient’s sister took the handset to confirm details.
23:31 — Case closed. Positive acknowledgment recorded.
+3 days — Written complaint received. The caseworker had already flagged the contact for review. Incident raised.
Root cause
The service is required to know that Mrs Ellison knows. She knew at 06:13. The service had no way to learn this, because knowing is not an event a phone can observe. She did not open the app — the lock screen had already told her everything, through the one field our payload discipline cannot redact: the sender’s name. The envelope announced what the letter withheld.
From that moment the service and the world disagreed, and every mechanism we have runs on the service’s side of the disagreement. This is not a defect in our implementation. It is an old result, older than this service: over an unreliable channel, no finite exchange of messages produces certainty on both sides that both sides know. Every acknowledgment is itself a message that can be lost, and an acknowledgment of the acknowledgment resolves nothing, only recurses. You cannot engineer your way to exactly-once delivery. You choose at-most-once, and accept that some people are never told; or at-least-once, and accept that some people are told again. There is no third setting on the dial.
Incident 0117 was the first failure mode. Incident 0332 is the second. They are one incident, viewed from opposite ends.
The deeper assumption sits in the acknowledgment itself. The protocol asks the receiver to complete it — to tap, to ring back, to cooperate in her own notification. The protocol assumes the receiver survives receipt in a condition to act. Grief violates that assumption. Mrs Ellison did not fail to acknowledge. She sat on the edge of her bed in the dark, and then she went to her sister. The remainder was ours.
What went well
Every component behaved as specified. This section is usually a comfort.
Where we got lucky
She was not alone when the second call came. Her sister was at the table. The caseworker heard the handset being passed. We do not design for luck, and we had it anyway.
Action items
AI-1 — Script v4.0: confirm-first contact. (Accepted. Owner: caseworker operations.) The informing script assumed each contact was the first. Under at-least-once delivery, no contact may assume it is the first. The revised script opens: “I’m calling about John Ellison. Can I ask — do you already know why I’m calling?” If yes, the call becomes confirmation and support. If no, the caseworker proceeds to inform. A repeated question costs the listener a breath. A repeated announcement is the event itself, again. We cannot make the delivery happen exactly once, so we have made the sentence safe to deliver twice. The retry is now idempotent. The protocol is unchanged; the payload absorbs the fault.
AI-2 — One-tap acknowledgment on the push itself. (Accepted. Owner: NKN engineering.) The current acknowledgment requires opening the app or returning a call. The push will now carry a single response, available from the lock screen: “I have been told.” An acknowledgment should cost the sender’s machine everything and the receiver’s grief almost nothing. Some will still not tap it. See Lessons.
AI-3 — Treat OS read receipts as acknowledgment. (Rejected.) A lock-screen impression means the notification was rendered, not read; read, not understood; and a phone face-up on a kitchen table is read by whoever is standing over it. We will not close a death case on the evidence of a glance that may have been someone else’s. This is 0117’s finding restated, and it stands.
Lessons learnt
The reliability literature tells us to make operations idempotent because the network will deliver them again whether we like it or not. We have always read that as guidance about payments and database writes. It is guidance about sentences.
Incident 0117 taught us not to trust the machine’s word that a person knows. Incident 0332 taught us the cost of that mistrust, paid at a kitchen table in Frome. Both lessons are correct. They do not cancel; they bracket. Somewhere between never told and told twice is where this service lives, and every review we write moves us within that interval, never out of it.
Recommended reading for new joiners: this document, then 0117, then this document again.
Status: Closed.
Recurrence: expected.