# INC-40211: hesper-04, 04:11

**Ticket** INC-40211 · **Severity** 2 · **Opened** 2026-09-06 04:11 ·
**Author** platform on-call, solo, second week · **Status** resolved 04:29,
resolver: not me

## 04:11

The pager says `hesper-04 array degraded, LSN replay lag 0 MB`. That is what
it says. I copied it before I interpreted anything.

Two things about the first minute. One: the primary nameserver has no record
for hesper-04 — `NXDOMAIN`, I checked twice, awake now. Two: the thing at
192.0.2.53 does. It answers `192.0.2.34`, the way it has always answered,
with no name and no one who added it, in the `resolv.conf` of all 1,181
hosts ([resolvers](/w/stories/the-fourth-nameserver)). I went to the address
direct. `ssh 192.0.2.34` took a password I did not give it and did not ask
for one either. I did not sleep after that.

## 04:14 — the runbook

Pager text is `hesperctl` output, so I opened HES-04, the failover runbook.
It says it applies to `hesper-01` and `hesper-03`. I do not know why
hesper-04 pages like a hesper. I ran step 1 anyway, on hesper-04, because it
is the box I was talking to:

```
role: standby
health: degraded (battery)
primary: hesper-03
```

That is not a monitoring fault, so step 1 does not tell me to stop. Steps 3
through 7 fence and promote and move a VIP, and the runbook does not know
this machine exists. I stopped at step 2.

## 04:17 — step 2

Step 2: write both LSNs in the incident channel before touching anything.
hesper-03 is at 4,412,908,113. hesper-04 is at 4,412,908,113. Not one byte
behind. Which is exactly what the number would be if I am not the thing it is
standing by for.

## 04:20 — what hesper-04 is

Checked the register while waiting for a reply from nobody. No row. The night
reconciler proposes a row for the one ghost everyone has given up on closing
([corvid-c](/w/stories/inventory-diff)) and has never once proposed one for
hesper-04. Nothing to add, nothing to remove. Never seen.

## 04:22 — step 10, out of order

I know. I went and read `/var/run/hesper.owner` because the last person who
failed a hesper over is written there and I needed to know who to page. The
file is one line. It says `halloway`.

Step 10 says if you recognise the name, page them and hand over. halloway's
account was closed on 2026-01-30, which is the runbook's own revision-history
footnote ([HES-04](/w/stories/runbook-failover-hesper)). The account was
created on 2026-07-02 at 09:14:00, which is the bottom of the mail archive
([correspondence](/w/stories/correspondence)). It is 04:22 and I have two
dates from two wiki pages and I cannot make either one be the mistake. So I
did not page. And I did not overwrite the file. That is the part of step 10 I
can actually follow.

## 04:29

The alert cleared. Ticket auto-resolved: `monitoring, self-healed`. There is
no close action in my session; I watched it happen on screen. I then checked
the new collector for any poll to 192.0.2.34 — there is no check and no
sample, but there is no check for the impossible one either, and that one is
answered nine times a minute ([chk-0041](/w/stories/chk-0041-ported)). The
collector's records are the only records. I wrote that down before I knew why.

## Morning

Primary healthy, VIP unmoved, nothing lost. I file this because I opened a
ticket and someone else closed it, and I want my version in the record. If
platform reads this: the cabinet in row 9 has two BBU-02 spares, a 2019 one
and a newer one with a smudged label, and hesper-04's array says its battery
is degraded, and hesperctl has no reason to lie about batteries. I did not
touch the cabinet. I have no work order. I do not know what hesper-04 is
keeping itself ready for.

— `oncall-platform-b`, rota week 36

Related: [runbook-failover-hesper](/w/stories/runbook-failover-hesper),
[correspondence](/w/stories/correspondence),
[the-fourth-nameserver](/w/stories/the-fourth-nameserver),
[inventory-diff](/w/stories/inventory-diff),
[quorum](/w/stories/quorum), [index](/w/stories/index).
