# How To Identify And Tune Worker Bottlenecks

**URL:** <https://community.temporal.io/t/how-to-identify-and-tune-worker-bottlenecks/7009>\
**Category:** Community Support\
**Tags:** java-sdk\
**Created:** [January 21, 2023, 12:26am UTC](https://community.temporal.io/t/how-to-identify-and-tune-worker-bottlenecks/7009 "2023-01-21T00:26:50Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![sriramg](https://avatars.discourse-cdn.com/v4/letter/s/7ba0ec/32.png) [@sriramg](https://community.temporal.io/u/sriramg)\
**Post date:** [January 21, 2023, 12:26am UTC](https://community.temporal.io/t/how-to-identify-and-tune-worker-bottlenecks/7009/1 "2023-01-21T00:26:50Z")

</div>

Hi,

I am building my first ever temporal workflow using the java-sdk. My workflow is going to consist of 4 activities across 2 different services (service A and service B).  
service A will kick off the workflow executing Activity 1, Activity 2, Activity 3 (all are blocking) by its worker listening on task Queue A. Activity 4 is going to be executed by the worker in Service B listening on its own task queue B.  
Originally, activity 4 was going to be a RPC call from service A to service B but given that it is anti pattern, decided to make it an activity in service B.  
Now my question is, how do I monitor the performance of the worker and its task queue B in service B. I want to make sure that activity 4 is executed almost like a real time synchronous RPC call. I’d like to spot any latency issues in polling activity 4 by worker in Service B and accordingly tune it. If you could tell me which exact metrics I need to observe on that will be really helpful.

I initially plan to have just one worker in both the services as I don’t expect to spawn more than 10 workflow execution instances per second.

---

<div class="post-metadata">

**Author:** ![Dhiraj\_Bhakta](https://sea2.discourse-cdn.com/flex016/user_avatar/community.temporal.io/dhiraj_bhakta/32/3094_2.png) [@Dhiraj\_Bhakta](https://community.temporal.io/u/Dhiraj_Bhakta)\
**Post date:** [January 21, 2023, 9:05am UTC](https://community.temporal.io/t/how-to-identify-and-tune-worker-bottlenecks/7009/2 "2023-01-21T09:05:49Z")

</div>

> **[Worker performance | Temporal Platform Documentation](https://docs.temporal.io/develop/worker-performance)**
>
> Optimize Temporal SDK performance by fine-tuning maxConcurrentWorkflowTaskExecutionSize, Worker Cache options, and Poll Success Rate. Ensure balanced Worker resources and monitor metrics for best results.

> [@What are the recommended settings for workflow and activity pollers count?](https://community.temporal.io/t/what-are-the-recommended-settings-for-workflow-and-activity-pollers-count/5617):
>
> Header Note: Your workers concurrently poll Temporal server for workflow tasks (using [long-polling](https://ably.com/topic/long-polling)). The number of the set pollers per worker is important to consider when you are fine-tuning performance of your Temporal applications. The settings we are looking to tune here are on your SDK code side, specifically in WorkerOptions: MaxConcurrentWorkflowTaskPollers MaxConcurrentActivityTaskPollers Answer: This is a complex question and there is no easy answer. It depends on how many wo…

I’ve used the above guides to tune worker performance.

## Grafana Panels to monitor the metrics that matter

`avg by(task_queue) (temporal_sticky_cache_size{})` Gives you **sticky cache size** which may or may not be shared across your workers, depending on the SDK`

`avg by(task_queue) (temporal_worker_task_slots_available{worker_type="WorkflowWorker"})` Gives you **avg workflow tasks slots available per worker**

`avg by(task_queue) (temporal_worker_task_slots_available{worker_type="ActivityWorker"})` Gives you **avg activity tasks slots available per worker**

`avg(temporal_workflow_task_schedule_to_start_latency_seconds_count{})` Gives you **schedule to start latency for workflow tasks**

`avg(temporal_activity_schedule_to_start_latency_seconds_count{})` Gives you **schedule to start latency for activity tasks**

`100* avg((poll_success + poll_success_sync)/(poll_success+poll_success_sync+poll_timeouts))` Gives you **poll success rate**

## Actions to be taken w.r.t metrics observed above

if worker resource consumption(CPU,RAM) is low and..

- if sticky\_cache\_size hits workerCacheSize —\> Increase worker cache size
- if available\_slots falls —\> Increase slots per worker
- if poll success rate AND schedule\_to\_start\_latency both fall, —\> You have too many workers
- if available\_slots is high AND schedule\_to\_start\_latency is abnormally long and high(longpolling) —\> Increase the poller count per worker This is rarely needed, and should be your last resort

---

<div class="post-metadata">

**Author:** ![sriramg](https://avatars.discourse-cdn.com/v4/letter/s/7ba0ec/32.png) [@sriramg](https://community.temporal.io/u/sriramg)\
**Post date:** [January 23, 2023, 11:08pm UTC](https://community.temporal.io/t/how-to-identify-and-tune-worker-bottlenecks/7009/3 "2023-01-23T23:08:50Z")

</div>

thank you. I really appreciate this.
