As a customer with a large foot-print of on-prem PI System software across multiple sites, I need a solution for health monitoring and diagnostics that allows me to identify issues (potential or current) with any component of the PI System or its data flow. PI System is a mission-critical application in my organization and any downtime has serious impact to the operations.
As a PI Administrator, I have been asked to use the PI System Monitoring AF toolkit, which is a service, and not a licensed product by OSIsoft/AVEVA. My organization has invested time and resources on implementing the toolkit and found very limited value (e.g. no capabilities to monitor PI Vision, PI Connectors, limited PI AF monitoring, no upgrade process, primitive analytic rules). Furthermore, the onus is on the customer to determine the health of the system, forcing us to become experts on the inner workings of the PI System. We would rather invest resources on building business value on top of the PI System.
What customers need and request is a way to be alerted to bad actors in the system, whether that is the update manager backing up, expensive queries from PI Vision or lagging analytics. There should be a way to be proactively notified about these issues with some indication of the root cause or at least the location of the issue (e.g., data archive versus PI Vision versus AF analytics). It would also be useful to have metrics around how many such issues exist and how frequently.