I am running a self-hosted deployment of Temporal OSS (v1.29+), packaged internally. We are validating and hardening Cross-Cluster Replication observability as part of DR readiness.
Which metric can help in setting up alert when replication is broken for self hosted Temporal Clusters?
Take a look at server and standby dashboard in repo here if helps
There are a few metrics, but these 3 are probably the most important.
- replication_tasks_lag - A rising trend in lag indicates replication cannot keep up or the remote cluster is not applying tasks.
- replication_latency - Watch for abnormal spikes or long periods of high latency.
- replication_tasks_fetched - Monitors whether poller is still pulling tasks. A drop to 0 while the source is taking writes usually means replication stopped.