Skip to content

The DNS Server That Was Quietly Doing Nothing

  • by

Two DNS servers were supposed to be splitting the work on a network we manage — one primary, one backup, standard redundancy. During an unrelated infrastructure project, we checked both, expecting a routine “yes, both fine.” One of them had an entirely empty configuration. No zones, no records, nothing. It had been that way for months.

And nobody had noticed, because the network never actually needed it to notice — the second resolver had been quietly doing all the work by itself the entire time.

Redundancy that isn’t tested isn’t redundancy

The whole point of having two DNS servers is that if one fails, the other keeps the network running while it gets fixed. That only works if the second one is actually configured to do the job. On paper, this network had that safety net. In practice, it had one working resolver and one empty shell that happened to still be reachable on the network — which is exactly the kind of gap that looks fine on every day-to-day check, because day-to-day, it genuinely was fine. The single working resolver never went down.

That’s the trap. A backup that’s silently broken is indistinguishable from a backup that’s working perfectly, right up until the moment you actually need it — and by definition, that’s the worst possible moment to find out.

How a gap like this survives a migration

The most likely explanation here traced back to an earlier network migration. When the primary DNS server was rebuilt or reconfigured as part of that project, its configuration apparently never got carried over completely — leaving it live and reachable on the network, but with nothing actually loaded into it. Every device on the network was pointed at both resolvers as designed. They just never needed the second one, so nothing forced the gap into view.

This is a pattern worth recognizing beyond DNS specifically: migrations are exactly when redundant systems are most likely to end up quietly incomplete, because attention naturally goes to “does the new thing work,” not “does the backup still work too.”

What actually catches this

Not a “is it online” check — the broken resolver was fully online and reachable the entire time. It takes an actual functional test: does this specific server correctly answer a real DNS query for a real internal name, not just “does it respond to a ping.” Once found, the fix here took minutes — resync the configuration from the working resolver. The months-long gap wasn’t a hard problem. It was an unmonitored one.

Questions worth asking about your own backups

  • For every system you consider redundant, has anyone actually tested the backup doing the job alone, recently — not just confirmed it’s reachable?
  • After your last infrastructure migration, did anyone specifically verify that backup/secondary systems came through intact, or just the primary one?
  • Would “the backup is broken” show up as a clear alert, or would it just sit quietly until the day the primary actually fails?

A backup you haven’t tested is a hope, not a plan. The only way to know it works is to actually check — before you need it to.

Leave a Reply

Your email address will not be published. Required fields are marked *