Skip to main content
RunBook Academy

Backup & DRII · RPO, RTO and Recovery SequencingObjectives

Service tiers and what a tier actually buys

Intermediate⏱ ~26 minawk

What you'll learn

  • Name the levers a tier has to change before it is operationally real rather than documentary
  • Derive a tier assignment from business impact rather than from technical interest or ownership pressure
  • Detect tier inversions, where a high-tier service depends on a lower-tier system
  • Explain why a service cannot complete recovery before its slowest required dependency does

Prerequisites

Verified against restic 0.19.1 · BorgBackup 1.4.5 · rclone 1.75.0 · MinIO (S3-compatible object storage) RELEASE.2025-09-07T16-13-09Z · OpenZFS 2.4.1 · LVM2 2.03.31(2) · btrfs-progs 6.17.1 · PostgreSQL 18.6 · pgBackRest 2.59.1 · Kubernetes (k3s) and etcd k3s v1.36.3+k3s1, etcd 3.7.1 · Velero 1.18.2 · Docker Engine 29.7.2 · Proxmox Backup Server (documentation only) 4.0.10-1 · Ubuntu (host baseline) 26.04 LTS · 2026-08-28

Not yet marked complete on this device.

Once restore duration is estimated from evidence rather than hope, an uncomfortable arithmetic appears. The estate contains more systems than the recovery capacity can carry at once: people, bandwidth, target hardware and attention are all finite while an incident is running, and something has to be restored second. Tiering is the decision about which something, made in advance and in daylight instead of at three in the morning by whoever is loudest. Done well it is a deliberate allocation of scarce recovery capacity. Done badly it is a colour-coded spreadsheet that changes nothing at all.

A tier is a set of differences, or it is a label

Assigning a system to a tier is a claim that it will be treated differently from a system in the tier below. The claim is testable, and the test is short: name the difference. If nothing about how the system is protected or recovered changes when the label changes, then the label describes somebody’s feelings about the system rather than a property of the recovery architecture.

There are only so many levers a tier can pull, and each of them costs something, which is exactly what makes a tier an allocation rather than an opinion.

Backup frequency sets the size of the loss window the tier is capable of delivering. A tier that promises fifteen minutes of loss while its members are captured once a night is not a tier; it is a wish with a number in it.

Retention and the count of restore points decide how far back recovery can reach. Corruption discovered on day forty is unrecoverable from a fourteen-day retention no matter how urgently the incident is prioritised.

Storage placement decides how fast a copy can be read back. A copy in a local repository, a copy in an object store across a wide-area link, and a copy in a cold archive class with a retrieval delay before the first byte moves are three different recovery times built from identical bytes.

Independent copies and replication decide whether recovery means restoring or means switching. This is the most expensive lever on the list and the one most often promised in a tier document without being bought.

Rehearsal cadence decides how stale your evidence is. A restore-duration estimate is worth no more than the age of the measurement behind it, and rehearsal is the only thing that keeps it young.

Position in the recovery order decides which system gets the people, the link and the target capacity first. It is the one lever that costs no money at all, and it is the one most estates never actually set.

Write the tiers as rows and the levers as columns, then read across. Two adjacent tiers whose rows are identical have a decorative boundary between them. Notice also what is absent from that list: the RPO and RTO figures written in the tier document. Those are the commitments. The levers are what make the commitments true, and a register that states the numbers without setting any lever has recorded an intention and filed it as a design.

Impact decides the tier; technical interest does not

The input to a tier assignment is what happens to the organisation while the system is gone, and that is not a fact engineers hold. It is elicited, from the people who carry the consequence, with questions concrete enough to be answered by someone who does not know what a snapshot is.

What stops, and who notices first — a customer, a regulator, or only the team that owns the system? Is there a manual fallback, and how many hours can it carry the real load rather than a demonstration of itself? Does a clock start that is not yours: a settlement window, a payroll run, a statutory reporting deadline, a contractual service credit? Is the damage reversible — orders taken on paper can be keyed in afterwards, whereas an audit trail with a hole in it stays holed. And does harm grow smoothly with the length of the outage, or step at a threshold? A system that is merely annoying for six hours and career-ending at seven does not belong in a tier described by an average.

Answers to those questions rank the estate differently from the way engineers rank it unprompted. Technical interest tracks novelty, scale and difficulty: the large database, the new cluster, the service somebody enjoys operating. Impact tracks dependence, and dependence accumulates under the least interesting systems in the building. Internal DNS, DHCP, the private certificate authority, the identity provider, the licence server, the configuration source of truth and the backup catalogue itself are nobody’s favourite system and sit beneath everything that is. A tier register derived from impact usually looks boring, names infrastructure nobody has thought about in two years near the top, and is correct.

A service cannot recover faster than its slowest dependency

A tier assigned to a service is an assertion about everything that has to be working before that service is. Almost no production service recovers alone. It needs name resolution to find anything, a certificate to be trusted, an identity provider to authenticate the first request, a secret store to hand it a database password, an image registry or package mirror to supply its code, a configuration source to tell it what it is, and — usually forgotten entirely — the monitoring that will tell the operator whether the restore worked.

Each of those is itself a system with its own tier, its own backup frequency and its own place in the recovery order. When a Tier 1 service depends on a Tier 3 identity provider, the tier register contains an arithmetic claim that is simply false: the service is promised a recovery time shorter than the recovery time of something it cannot start without. Nothing about the label changes the arithmetic. The service recovers at the pace of the dependency, which is to say it recovers at Tier 3, and the only thing the Tier 1 label achieves is that nobody finds out until the day it matters.

Every system is Tier 1, which is the same as no tiers

Tiering fails much more often through the label than through the arithmetic. The register goes out for review, and no owner accepts Tier 2, because the label is read as a judgement of the system’s worth and by extension of its owner’s. The pressure is asymmetric: arguing a system up costs an email, while arguing it down costs a meeting nobody wants. After two review rounds every system is Tier 1.

The result is not a cautious estate. It is an estate with a uniform recovery architecture — the same backup frequency, the same retention, the same storage, the same rehearsal cadence for everything — which is byte for byte the architecture you would have had with no tiers at all, plus a document promising recovery times that the uniform architecture cannot deliver for more than the first few systems in the queue. Having no tiers is honest. Having tiers that are all the same tier is the identical architecture with a broken promise attached.

The fix is not firmer language in the policy. It is to make Tier 1 rivalrous, so that adding a system to it visibly removes something from another system. Fix the capacity first and allocate it afterwards: a fixed number of rehearsal slots per quarter, a fixed number of positions before position ten in the recovery order, a fixed replication budget derived from the bandwidth that actually exists. Then Tier 1 membership is a subtraction from a pool rather than an adjective, and the conversation changes from “is this system important?” — to which the answer is always yes — to “should this system be restored before that one?”, which has a real answer that owners can argue about productively.

Two supports make that hold. Price the tier, so the frequency, the retained copies, the warm storage and the rehearsal hours land in the owner’s budget rather than in an infrastructure line nobody reads. And force a total order for recovery, because a ranking cannot be gamed the way a label can: positions one through N exist, exactly one system occupies each, and every promotion is somebody else’s demotion.

A register you can audit, and the inversion it must never contain

A tier register earns its keep when it is machine-readable, because the most valuable question you can ask of it is one nobody answers by eye. Keep it beside the estate inventory, one row per service, with its tier and the systems it cannot start without.

service,tier,dependencies
checkout-api,1,identity;orders-db;dns-internal
orders-db,1,dns-internal;secret-store
identity,3,dns-internal;identity-db
dns-internal,1,
secret-store,2,dns-internal
identity-db,3,dns-internal

The audit is a search for inversions: any service whose dependency sits at a lower tier than the service itself.

REGISTER=tiers.csv
awk -F, 'NR > 1 { tier[$1] = $2; deps[$1] = $3 }
END {
  for (svc in deps) {
    n = split(deps[svc], d, ";")
    for (i = 1; i <= n; i++)
      if (tier[d[i]] != "" && tier[d[i]] + 0 > tier[svc] + 0)
        printf "%s is tier %s but depends on %s at tier %s\n", svc, tier[svc], d[i], tier[d[i]]
  }
}' "$REGISTER"

Every line it emits is one of exactly two things, and both need a decision. It is either a dependency that has been mis-tiered, in which case the dependency’s levers have to move and someone has to pay for them, or it is a commitment on the dependent service that cannot be met, in which case the register is promising the business something the architecture will not do. Silence is not an available option, because the inversion resolves itself during the incident by lengthening the outage.

Production discipline

  1. Express every tier as a diff. For each boundary, state which of frequency, retention, storage placement, replication, rehearsal cadence and recovery order is set differently. A boundary with no differing lever is documentation and should be deleted rather than defended.
  2. Take impact from the people who carry the consequence. Ask what stops, who notices, whether a manual fallback exists and for how long, whether an external clock starts, and whether the damage is reversible. Engineers rank by interest; interest is uncorrelated with dependence.
  3. Tier the dependency closure, not the service. Identity, DNS, certificate issuance, secret storage, image supply, configuration and the monitoring you will use to judge the restore are all in the closure, and the service’s achievable recovery time has a floor set by them.
  4. Make Tier 1 rivalrous before you publish it. Fix the number of rehearsal slots, early recovery-order positions and replicated targets first, then allocate. A tier that cannot run out will not stay meaningful through two review cycles.
  5. Audit for inversions on every change, not once a year. Treat each inversion as either a dependency to re-tier and fund, or a commitment to withdraw from the business, and record which one you chose.

Cross-course references

  • Observability for Production Sysadmins — Part XXII (SLO-Based Alerting) turns a promise about a service into a measured budget that is spent and reported, which is the same discipline this lesson asks of a tier: a tier nobody measures against decays into a label exactly as an unmeasured objective does, and the rehearsal cadence in a tier definition is what produces the measurement.
  • Kubernetes for Production Sysadmins — Part CX (Priority and Preemption) is the clearest example in an estate of a tier that buys something concrete, because under node pressure a higher-priority Pod evicts a lower-priority one and the label therefore changes behaviour under contention; a recovery tier has to be specified with the same precision about what happens when capacity runs short.
  • Secrets, PKI & Certificate Management for Infrastructure Engineers — Part XII (Secret Management Platforms) covers the system most frequently left at a low tier underneath high-tier services, and its material on how workloads obtain credentials tells you whether a restored Tier 1 service can authenticate at all before that platform is back, which is exactly the dependency floor this lesson makes you compute.

Quiz

Knowledge check · 5 questions

  1. Q1. A three-tier policy is signed off by the business. Tier 1 and Tier 2 systems are captured on the same nightly schedule into the same repository with the same 30-day retention, are rehearsed on the same annual cadence, and appear in the recovery runbook in alphabetical order. What has the tiering accomplished?

  2. Q2. A checkout service is Tier 1 with a four-hour recovery commitment. Every request it serves is authenticated against an identity provider classified Tier 3, whose rehearsed recovery takes most of a working day and which must be serving before checkout can take traffic. What is true of the four-hour figure?

  3. Q3. A Tier 1 checkout service cannot start without an identity provider currently classified Tier 3. Editing the register to promote that identity provider to Tier 1, with no other change, shortens the checkout service recovery floor.

  4. Q4. Which of these, when set differently either side of a tier boundary, make that boundary operationally real? Select all that apply.

  5. Q5. After two review rounds, every system in the estate is Tier 1. State what has been lost, and name one mechanism that restores the distinction.

Passing score: 75%. Answers are checked in this browser.