# 250ms latency for a workflow with 2 empty activities

**URL:** <https://community.temporal.io/t/250ms-latency-for-a-workflow-with-2-empty-activities/16771>\
**Category:** Community Support\
**Tags:** go-sdk, helm, postgresql\
**Created:** [March 12, 2025, 8:12pm UTC](https://community.temporal.io/t/250ms-latency-for-a-workflow-with-2-empty-activities/16771 "2025-03-12T20:12:05Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![abhishek.more](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abhishek.more](https://community.temporal.io/u/abhishek.more)\
**Post date:** [March 12, 2025, 8:12pm UTC](https://community.temporal.io/t/250ms-latency-for-a-workflow-with-2-empty-activities/16771/1 "2025-03-12T20:12:05Z")

</div>

I did load testing of self-hosted temporal deployment  
temporal version: 1.25  
postgres17 m7g8xlarge single az  
deployed on aws managed k8s with 12 pods for frontend, matching and history service each  
Have configured HPA as well to scale horizontally and ensured it doesn’t hit max replicas  
20 pods for worker with 20k max concurrent activity/workflow and 200 max poller for activity/workflow support

load testing, single workflow with 2 empty activities for 50rps, 100rps, 200rps, 300rps  
On an average got 250ms latency for workflow completion and 80ms for workflow execute-call-to-schedule. Both of these metrics are custom and not from metrics emitted temporal.

1. Is it good latencies considering workflow with 2 empty activities?
2. Could it be improved by using temporal cloud? If yes, then what would be latency with temporal cloud?

Thanks 🙂

---

<div class="post-metadata">

**Author:** ![tihomir](https://sea2.discourse-cdn.com/flex016/user_avatar/community.temporal.io/tihomir/32/6580_2.png) [@tihomir](https://community.temporal.io/u/tihomir)\
**Post date:** [March 16, 2025, 10:54pm UTC](https://community.temporal.io/t/250ms-latency-for-a-workflow-with-2-empty-activities/16771/2 "2025-03-16T22:54:56Z")

</div>

Can you show  
shard lock latency  
`histogram_quantile(0.99, sum by (le) (rate(semaphore_latency_bucket{operation="ShardInfo",service_name="history"}[1m])))`

``  
service latency  
histogram\_quantile(0.95, sum(rate(service\_latency\_bucket{service=“frontend”}[1m])) by (operation, le))  
`

db latency  
`histogram_quantile(0.95, sum(rate(persistence_latency_bucket{}[1m])) by (operation, le))`

resource exhausted errors  
`sum(rate(service_errors_resource_exhausted{}[1m])) by (operation, resource_exhausted_cause)`

> Is it good latencies considering workflow with 2 empty activities?

i think its hard to tell. understanding your db latency especially is imo important

> Could it be improved by using temporal cloud? If yes, then what would be latency with temporal cloud?

docs [here](https://docs.temporal.io/cloud/service-availability#:~:text=What%20kind%20of%20latency%20can,SLO%20of%20200ms%20per%20region.&text=As%20Temporal%20continues%20working%20on,these%20numbers%20will%20progressively%20decrease) can help, but cloud would give you higher throughput and lower latencies in general, see [Benchmarking Latency: Temporal Cloud vs. Self-Hosted Temporal | Temporal](https://temporal.io/blog/benchmarking-latency-temporal-cloud-vs-self-hosted-temporal) for example

---

<div class="post-metadata">

**Author:** ![abhishek.more](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abhishek.more](https://community.temporal.io/u/abhishek.more)\
**Post date:** [March 17, 2025, 9:35am UTC](https://community.temporal.io/t/250ms-latency-for-a-workflow-with-2-empty-activities/16771/3 "2025-03-17T09:35:35Z")

</div>

db latency

 ![Screenshot 2025-03-17 at 3.00.08 PM](https://us1.discourse-cdn.com/flex016/uploads/temporal/original/2X/7/77db7f7abbfe1c79f83e02eb3ee4e2eec84ebbf2.png)

service latency

 ![Screenshot 2025-03-17 at 3.02.28 PM](https://us1.discourse-cdn.com/flex016/uploads/temporal/original/2X/7/7002cbeb786c4e00aafa47f5926466c2ad17d121.png)

shard lock latency

 ![Screenshot 2025-03-17 at 3.03.47 PM](https://us1.discourse-cdn.com/flex016/uploads/temporal/original/2X/0/0a1f1740ef1e1c21717968f68071a26271fad012.png)

resource exhausted errors

 ![Screenshot 2025-03-17 at 3.04.56 PM](https://us1.discourse-cdn.com/flex016/uploads/temporal/original/2X/8/8a3de538498d3defe4f5698e5549b1ed01882401.png)

---

<div class="post-metadata">

**Author:** ![tihomir](https://sea2.discourse-cdn.com/flex016/user_avatar/community.temporal.io/tihomir/32/6580_2.png) [@tihomir](https://community.temporal.io/u/tihomir)\
**Post date:** [March 17, 2025, 2:33pm UTC](https://community.temporal.io/t/250ms-latency-for-a-workflow-with-2-empty-activities/16771/4 "2025-03-17T14:33:18Z")

</div>

exclude PollWorkflow/ActivityTaskQueue operations from your service latency graph (they are long-poll operations so can take up to 70s)

---

<div class="post-metadata">

**Author:** ![abhishek.more](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abhishek.more](https://community.temporal.io/u/abhishek.more)\
**Post date:** [March 17, 2025, 7:49pm UTC](https://community.temporal.io/t/250ms-latency-for-a-workflow-with-2-empty-activities/16771/5 "2025-03-17T19:49:27Z")

</div>

Updated service latency graph

 ![Screenshot 2025-03-18 at 1.18.31 AM](https://us1.discourse-cdn.com/flex016/uploads/temporal/original/2X/4/40cc2f9244f7563f1220bb43ac6738de6f8d3ebb.png)

---

<div class="post-metadata">

**Author:** ![abhishek.more](https://avatars.discourse-cdn.com/v4/letter/a/278dde/32.png) [@abhishek.more](https://community.temporal.io/u/abhishek.more)\
**Post date:** [March 26, 2025, 8:26am UTC](https://community.temporal.io/t/250ms-latency-for-a-workflow-with-2-empty-activities/16771/6 "2025-03-26T08:26:19Z")

</div>

@tihomir bumping this up
