Cloud

Azure PostgreSQL: diagnose read replica lag before promotion

A production runbook to qualify Azure Database for PostgreSQL read replica lag, prove the achievable RPO, and choose repair, planned switchover, or forced promotion.

30 Aug 2026 azurepostgresqlflexible-serverread-replicareplicationwaldisaster-recoveryobservabilityautomationrunbookrollbackproduction

A cross-region read replica is ready for disaster recovery, but its lag starts growing while the primary is still serving writes. The first reaction is often to promote it quickly, scale it, or recreate it. Each option changes a different part of the system. A forced promotion can expose a writable server sooner, but any transaction still missing from the replica defines real data loss.

The production case is an Azure Database for PostgreSQL Flexible Server primary with an asynchronous read replica used for reporting or regional recovery. This runbook qualifies the replication path, proves the recovery point available on the replica, and ends with one explicit decision: let it catch up, repair capacity or workload pressure, perform a planned switchover, force promotion with an accepted RPO, or rebuild the replica without promoting it.

Freeze the promotion contract

Do not start with the Promote button. Record which topology exists and what the application expects after the operation. Promoting to a standalone server and switching the replica into the primary role are different changes.

yaml postgres-replica-incident.yml
incident:
started_at_utc: 2026-08-30T05:40:00Z
primary_server: pg-orders-prod-we
candidate_replica: pg-orders-dr-ne
primary_region: westeurope
replica_region: northeurope
workload: orders-api
role: reporting-and-disaster-recovery

observed:
replica_lag_seconds: <measured-value>
max_physical_lag_bytes: <measured-value>
primary_transaction_log_storage: <measured-value>
primary_write_available: true-or-false
replica_read_available: true-or-false
replication_state: <state>

promotion_contract:
mode: switchover-or-standalone
option: planned-or-forced
maximum_accepted_rpo: <seconds-and-business-marker>
maximum_accepted_rto: <minutes>
writer_virtual_endpoint: <fqdn-or-none>
reader_virtual_endpoint: <fqdn-or-none>
decision_owner: <role>
application_validation: <probe>

stop_conditions:
- replication evidence is unavailable
- replica data point is older than accepted RPO
- target authentication or network policy is untested
- client connection target after promotion is ambiguous

A promotion is not successful merely because the target becomes writable. The application, identities, DNS or virtual endpoints, firewall rules, parameters, and monitoring must all point to a coherent primary.

Read lag in seconds and bytes

Azure exposes Read Replica Lag on the replica in seconds and Max Physical Replication Lag on the primary in bytes. These signals answer different questions. Seconds estimate how old the last replayed transaction is. Bytes estimate the WAL distance to the most delayed connected replica.

A quiet database can show an increasing time since the last replayed transaction even when nothing new needs replaying. Conversely, a write-heavy workload can generate a large byte gap in a short time. Read both dimensions with write throughput and the replication state.

bash 01-replica-metrics.sh
PRIMARY_ID="<primary-resource-id>"
REPLICA_ID="<replica-resource-id>"
START="2026-08-30T05:30:00Z"
END="2026-08-30T06:30:00Z"

az monitor metrics list --resource "$REPLICA_ID" --metric physical_replication_delay_in_seconds --interval PT1M --start-time "$START" --end-time "$END" --aggregation Maximum Average --output json

az monitor metrics list --resource "$PRIMARY_ID" --metric physical_replication_delay_in_bytes storage_percent --interval PT1M --start-time "$START" --end-time "$END" --aggregation Maximum Average --output json

Also retain Transaction Log Storage Used, CPU, IOPS and disk bandwidth around the same window. A growing byte gap with saturated replica I/O points toward replay capacity. A growing WAL footprint on the primary shows that waiting has a cost even when primary queries still succeed.

Verify the replication state inside PostgreSQL

Platform metrics show the trend. PostgreSQL views show whether WAL is being sent, received, flushed and replayed. Run the first query on the primary with a suitably privileged diagnostic identity.

sql 02-primary-replication-state.sql
select
application_name,
client_addr,
state,
sync_state,
sent_lsn,
write_lsn,
flush_lsn,
replay_lsn,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn)) as replay_gap,
write_lag,
flush_lag,
replay_lag
from pg_stat_replication
order by application_name;

On the candidate replica, prove that it is still in recovery and capture its receive and replay positions.

sql 03-replica-replay-state.sql
select
pg_is_in_recovery() as is_replica,
pg_last_wal_receive_lsn() as receive_lsn,
pg_last_wal_replay_lsn() as replay_lsn,
pg_size_pretty(
  pg_wal_lsn_diff(
    pg_last_wal_receive_lsn(),
    pg_last_wal_replay_lsn()
  )
) as local_replay_gap,
pg_last_xact_replay_timestamp() as last_replayed_transaction,
now() - pg_last_xact_replay_timestamp() as replay_age;

Lag columns can be null during idle periods, and timestamps alone do not prove missing data. Correlate the LSN positions with a known business marker or controlled write before making an RPO decision.

Locate the pressure before scaling

Separate four failure families:

  • the primary suddenly produces more WAL because of a batch, index operation, bulk import, vacuum pressure, or release;
  • the replica cannot receive or replay fast enough because compute, IOPS, disk bandwidth, or network path is constrained;
  • reporting queries on the replica compete with recovery or create repeated conflicts;
  • the replication link or slot has become unhealthy and the service is using a slower catch-up path.

Compare the first rise in lag with deployments, jobs, schema changes and reporting schedules. Scaling the primary can increase WAL generation. Scaling only the replica helps only when replay capacity is the limiting factor. Cancel a reporting query only when evidence ties it to recovery pressure and the business owner accepts the interruption.

Do not reduce parameters on a replica while it is behind. Configuration drift between primary and replica also matters for promotion: authentication mode and server parameters do not become correct automatically because roles change.

Protect the primary from retained WAL

Read replicas rely on WAL retention. A persistently lagging or inactive replica can cause transaction logs to accumulate on the primary. Storage pressure can then threaten the writable service that the replica was intended to protect.

The containment order is:

  1. stop or reduce the identified batch or reporting pressure when operationally acceptable;
  2. preserve metrics, LSNs, active sessions and the replication state;
  3. increase replica capacity only when saturation is proven;
  4. watch primary storage and transaction log usage while the replica catches up;
  5. delete and rebuild a broken replica only after deciding that its recovery point is no longer useful.

Deleting a replica releases the relationship but also removes it as a promotion candidate. It is containment for primary health, not a transparent repair.

Prove the target connection path

A planned switchover promotes the read replica to primary and demotes the former primary. It requires the writer and reader virtual endpoints to be configured, with the candidate included as the reader target. A standalone promotion instead detaches the replica and leaves application redirection to the team.

Before either action, test the target server’s effective controls:

  • PostgreSQL and Microsoft Entra authentication expected by the application;
  • firewall or private networking path from production clients;
  • server parameters required by the workload;
  • extensions, roles and databases expected after cutover;
  • writer and reader endpoint resolution from the application runtime;
  • dashboards and alerts attached to the future primary.

A green read-only query does not validate write identity, transaction behavior, or connection pooling after promotion.

Use a marker to prove the achievable RPO

When the primary is still available, create a harmless, auditable marker through the normal application write path. Do not invent a production write directly in the database if the application contract does not allow it.

yaml replica-promotion-gate.yml
marker:
id: dr-check-20260830-0615
written_through: orders-api
committed_at_utc: 2026-08-30T06:15:00Z
contains_personal_data: false

candidate_replica:
marker_visible: true-or-false
observed_at_utc: <timestamp>
replay_lsn: <lsn>
lag_seconds: <value>
lag_bytes: <value>

allow_planned_switchover_when:
- marker is visible on the replica
- byte gap converges to the accepted boundary
- both servers are Ready
- virtual endpoints target the expected servers
- write-path canary and rollback owner are ready

allow_forced_promotion_when:
- primary region is unavailable or cannot safely recover in the RTO
- last replicated business marker is known
- data loss up to the observed lag is explicitly accepted
- reconciliation procedure exists for missing writes

abort_when:
- replica continues to diverge
- primary storage reaches the incident stop threshold
- target identity or network path fails
- application owners cannot state the accepted RPO

The marker turns “lag looks low” into a business recovery point. Keep its identifier with the promotion record and use it to reconcile writes after a forced operation.

Decide repair, switchover, forced promotion, or rebuild

Let the replica catch up when the primary is healthy, the byte gap is shrinking, storage remains safe and no outage requires immediate cutover. Repair or scale when a measured bottleneck explains the lag and a canary proves convergence.

Use planned switchover when both servers are Ready, pending WAL can synchronize, virtual endpoints are valid, and the application has passed the target checks. Use forced promotion only for an outage where availability is worth the measured data loss. The lag at detachment is the approximate loss boundary, not a cosmetic metric.

After a switchover, validate writes, reads, identities, endpoints, replication direction, HA requirements and monitoring. The former primary becomes a replica; switching back is another controlled promotion, not an undo button. After standalone or forced promotion, retain the old topology until write reconciliation and ownership of the new endpoint are complete.

Conclusion

A delayed PostgreSQL read replica is not automatically a promotion target. First determine whether time lag represents real missing WAL, measure the byte gap and business marker, locate pressure, and protect primary storage.

Promote only when the target connection path is proven and the accepted RPO is explicit. Prefer a planned switchover when synchronization is possible. When a regional outage forces the decision, preserve the last replicated marker, acknowledge the loss boundary, validate the new writer, and keep reconciliation as part of the recovery rather than an afterthought.