Operations and Troubleshooting
The goal of operations is to ensure that issues are detected by monitoring before users notice them, to enable rapid localization using a consistent set of evidence after a failure occurs, and to ensure that every upgrade or recovery has a fallback path.
flowchart LR
A[Inspection and monitoring] --> B[Detect anomaly]
B --> C[Fix time and impact scope]
C --> D[Find first failed layer]
D --> E[Change one variable]
E --> F[Verify user path and record root cause]
Establishing an Operational Baseline
- Use Daily Inspections to define shift, daily, and weekly checks.
- Use Monitoring, Logging, and Alerting to cover the host, control plane, protocols, media, database, and disk.
- Manage accounts, tokens, keys, networks, and auditing according to the Security Baseline.
- To notify third parties about devices, RTC, MediaServer failures, streams, and recordings, configure the Main-System Third-Party Webhook.
- Complete Backup, Recovery, Upgrade, and Rollback before making production changes.
When a Failure Occurs
If the layer where the problem exists is unknown, first refer to Locating Problems by Symptom. Common specialized areas include:
- Device Offline or Not Discovered
- No Video or Playback Failure
- Recording Session Successful but Recording File Not Found
- RTC or Browser Connection Failure
Effective Troubleshooting
Do not start by "restarting all services." First, preserve the scene:
- Start and recovery time of the failure;
- Affected users, devices, channels, and nodes;
- Whether it is a total failure or a failure of a specific protocol, browser, or network;
- Recent changes to configuration, certificates, network, version, or devices;
- Command ID, Task ID, TraceId, Call-ID, stream identifier, or recording index/file ID;
- Application, Agent, MediaServer, database, and proxy logs within the same time window.
Check backward from the user's final result to find the first inconsistency. Change only one variable at a time and record the results before and after the change.
Restart Principles
Restarting can restore state, but it also destroys evidence and may affect more users. Execute a restart only after necessary logs have been saved and the scope of impact and rollback plan have been clearly defined. The root cause must still be identified after a restart; "it worked after restarting" is not a valid conclusion to close an issue.