#13062 Input Needed: RelEng Monitoring Requirements for Zabbix.
Opened by amedvede. Modified

Hello everyone,

As you may know, we are planning to migrate our monitoring from Nagios to Zabbix.

Following a productive discussion with Greg, we concluded that the RelEng team has specific monitoring requirements that would be incredibly useful for managing both our planned and unplanned work.

To ensure the new Zabbix implementation serves our team effectively, we need your expertise. We are gathering a list of the specific metrics and triggers you want to have for RelEng services and tools.

How to Contribute
This ticket will remain open for some time, as we expect to gather this information incrementally. As you think of new requirements, please add them here using the following format:

Server/Tool Name: (The name of the server or tool)
Metrics: (The specific metrics or resources you want to track, e.g., job queue length, disk I/O, API endpoint latency)
Triggers: (The conditions that should generate an important notification)
Additional information: (Any other context, justification, or thoughts on this item)

We will gradually add the collected monitoring items to the project configurations.

Thanks for your input!


Server/Tool Name: Bodhi
Triggers: We potentially may make a trigger similar to one that comes to email with `Bodhi error title`
Additional information: Now, when a specific error happens that affects `bodhi-celery` service, we usually restart the service and resume the queue. Zabbix allows us to take actions based on triggers, so we can automate this process.

Metadata Update from @jnsamyak:
- Issue untagged with: sprint-4
- Issue tagged with: sprint-5

Metadata Update from @jnsamyak:
- Issue untagged with: sprint-5

Metadata