# Matching service start/stop loop in production deployment

**URL:** <https://community.temporal.io/t/matching-service-start-stop-loop-in-production-deployment/907>\
**Category:** Community Support\
**Created:** [November 4, 2020, 12:03am UTC](https://community.temporal.io/t/matching-service-start-stop-loop-in-production-deployment/907 "2020-11-04T00:03:32Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![benmays](https://avatars.discourse-cdn.com/v4/letter/b/8c91f0/32.png) [@benmays](https://community.temporal.io/u/benmays)\
**Post date:** [November 4, 2020, 12:03am UTC](https://community.temporal.io/t/matching-service-start-stop-loop-in-production-deployment/907/1 "2020-11-04T00:03:32Z")

</div>

Hi all,

I’m having trouble getting Temporal run in our production environment. I’m currently running all 4 services as a single container as a k8s pod, with 2 pods running. I’ve opened the same exact ports as the docker container. I’m using RDS for the backing datastore.

The temporal-server containers simply loop, seemingly stuck with the following error message:

```auto
{"level":"error","ts":"2020-11-03T23:41:56.067Z","msg":"Error updating timer ack level for shard","service":"history","shard-id":4,"address":"127.0.0.1:7234","shard-item":"0xc0013dcd00","component":"timer-queue-processor","cluster-name":"active","error":"Failed to update shard. Previous range ID: 5717; new range ID: 5718","logging-call-at":"timerQueueAckMgr.go:391","stacktrace":"go.temporal.io/server/common/log/loggerimpl.(*loggerImpl).Error\n\t/temporal/common/log/loggerimpl/logger.go:138\ngo.temporal.io/server/service/history.(*timerQueueAckMgrImpl).updateAckLevel\n\t/temporal/service/history/timerQueueAckMgr.go:391\ngo.temporal.io/server/service/history.(*timerQueueProcessorBase).internalProcessor\n\t/temporal/service/history/timerQueueProcessorBase.go:319\ngo.temporal.io/server/service/history.(*timerQueueProcessorBase).processorPump\n\t/temporal/service/history/timerQueueProcessorBase.go:194"}

```

I also see persistence failures:

```auto
{"level":"error","ts":"2020-11-03T23:41:10.869Z","msg":"Persistent store operation failure","service":"matching","component":"matching-engine","wf-task-queue-name":"/_sys/temporal-sys-tq-scanner-taskqueue-0/2","wf-task-queue-type":"Activity","store-operation":"update-task-queue","error":"Persistence Max QPS Reached.","logging-call-at":"taskReader.go:166","stacktrace":"go.temporal.io/server/common/log/loggerimpl.(*loggerImpl).Error\n\t/temporal/common/log/loggerimpl/logger.go:138\ngo.temporal.io/server/service/matching.(*taskReader).getTasksPump\n\t/temporal/service/matching/taskReader.go:166"}

```

Also seeing the following error rarely:

```auto
{"level":"error","ts":"2020-11-03T23:18:01.139Z","msg":"Error looking up host for shardID","service":"history","component":"shard-controller","address":"127.0.0.1:7234","error":"Not enough hosts to serve the request","operation-result":"OperationFailed","shard-id":1,"logging-call-at":"shardController.go:343","stacktrace":"go.temporal.io/server/common/log/loggerimpl.(*loggerImpl).Error\n\t/temporal/common/log/loggerimpl/logger.go:138\ngo.temporal.io/server/service/history.(*shardController).acquireShards.func1\n\t/temporal/service/history/shardController.go:343"}

```

The frontend API isn’t responding to tctl requests so I don’t have much to debug with right now. The frontend does show the grpc listener coming up though.

Not sure how to proceed without digging into the code, any ideas? I’d be curious if there is a way to directly inspect the cluster/gossip output from the database?

---

<div class="post-metadata">

**Author:** ![Wenquan\_Xing](https://sea2.discourse-cdn.com/flex016/user_avatar/community.temporal.io/wenquan_xing/32/362_2.png) [@Wenquan\_Xing](https://community.temporal.io/u/Wenquan_Xing)\
**Post date:** [November 4, 2020, 4:48am UTC](https://community.temporal.io/t/matching-service-start-stop-loop-in-production-deployment/907/2 "2020-11-04T04:48:03Z")

</div>

> [@benmays](#):
>
> `Previous range ID: 5717; new range ID: 5718`

This huge range ID may indicates that your membership (frontend / matching / history) ring is not stable, one host may steal the shard from another one, like a ping pong game.

> [@benmays](#):
>
> `Persistence Max QPS Reached.`

this means that the persistence layer is under high load.

can you shard the service config? also make sure the servers (frontend / matching / history) can talk to each other.

---

<div class="post-metadata">

**Author:** ![benmays](https://avatars.discourse-cdn.com/v4/letter/b/8c91f0/32.png) [@benmays](https://community.temporal.io/u/benmays)\
**Post date:** [November 5, 2020, 6:26am UTC](https://community.temporal.io/t/matching-service-start-stop-loop-in-production-deployment/907/3 "2020-11-05T06:26:51Z")

</div>

Thanks for the quick turnaround! It looks like the issue was caused because my temporal-server peers could reach each other. I was able to resolve it by setting the broadcast address to each host’s public ip and then binding on 0.0.0.0. I also opened up additional ports for peer-to-peer communication.

I found some useful info deep in the support tickets, linking here for others:

- Ports required for peers: [Communication between multiple instances of temporal server](https://community.temporal.io/t/communication-between-multiple-instances-of-temporal-server/228/14)
- Background info on broadcast address: [What is the equivalent of cadence:bootstrapHosts in temporal](https://community.temporal.io/t/what-is-the-equivalent-of-cadence-bootstraphosts-in-temporal/524)
