Daily Inspection
Inspections should cover end-to-end business logic, not just process survival. The following is the minimum baseline that can be directly incorporated into the on-call manual.
flowchart LR
A[Each shift: online rate, streams, recordings, disks] --> B[Weekly: sample real workflows and backups]
B --> C[Monthly: recovery drill, security, capacity]
C --> D[Turn manual findings into automatic alerts]
Per Shift or Daily
- WebUI/API liveness and readiness are normal.
- Main service and NodeAgent are online; host daemons are not restarting repeatedly.
- Heartbeats, versions, ports, and stream counts for all required MediaServers are normal.
- Database connections, lock waits, slow queries, and disk space show no anomalies.
- Device online rate shows no sudden mutations compared to yesterday's baseline.
- Protocol command timeout rates, authentication failures, and registration fluctuations are within the baseline.
- Active recording sessions are continuously producing files; scanning and trimming are not backlogged.
- Recording disk and system disk bytes and inodes are above the safety threshold.
- WebHook, asynchronous command, and RTC release failures are not growing continuously.
- Certificates have sufficient time before expiration; the most recent renewal task was successful.
Weekly
- Spot-check one real stream each for ONVIF, GB28181, and RTSP.
- Spot-check the retrieval, Range playback, and download of a new recording segment.
- Check for configuration version drift, pending restarts, and unclosed failed tasks.
- Review newly added administrators, roles, Tokens, and abnormal logins.
- Review capacity growth trends, rather than just current remaining space.
- Confirm backup jobs were successful and spot-check that backup files are readable.
Monthly or Before Major Changes
- Perform a database and configuration recovery in an isolated environment.
- Verify host process recovery and alert notifications during a maintenance window.
- Review port exposure, reverse proxy, TLS, and TURN configurations.
- Update device model/firmware compatibility records.
- Clean up expired Tokens, orphaned devices, long-term failed tasks, and expired trimmed files.
- Evaluate version upgrades and prepare rollback materials.
Key Metric Recommendations
Layer metrics by host, AK process, Agent, MediaServer, MySQL, disk, protocol, recording, and RTC. Metrics must at least include node and environment tags. Control the cardinality of device and channel metrics to avoid treating every short-lived stream as a permanent high-cardinality tag.
Handling Inspection Anomalies
- Record the time, metric value, threshold, and scope of impact.
- Compare with the baseline from the same time yesterday or last week.
- Confirm if there are planned changes or capacity growth explanations.
- Use the TraceId, device identifier, stream identifier, or task ID to access the corresponding logs.
- If users are already affected, use the Problem Locator.
- After resolution, verify metric recovery and the end-user path, and record the root cause and preventive actions.
Inspection Completion Standards
Every item must have a clear owner, data source, threshold, handling action, and escalation path. Issues discovered during manual inspections should be gradually converted into automated alerts; alerts that are long-term ignored should have their thresholds adjusted or be deleted, rather than becoming noise.