AKStream.Next · 文档中心

Operations and Troubleshooting

Operations and Troubleshooting

The goal of operations is to ensure that issues are detected by monitoring before users notice them, to enable rapid localization using a consistent set of evidence after a failure occurs, and to ensure that every upgrade or recovery has a fallback path.

flowchart LR
    A[Inspection and monitoring] --> B[Detect anomaly]
    B --> C[Fix time and impact scope]
    C --> D[Find first failed layer]
    D --> E[Change one variable]
    E --> F[Verify user path and record root cause]

Establishing an Operational Baseline

  1. Use Daily Inspections to define shift, daily, and weekly checks.
  2. Use Monitoring, Logging, and Alerting to cover the host, control plane, protocols, media, database, and disk.
  3. Manage accounts, tokens, keys, networks, and auditing according to the Security Baseline.
  4. To notify third parties about devices, RTC, MediaServer failures, streams, and recordings, configure the Main-System Third-Party Webhook.
  5. Complete Backup, Recovery, Upgrade, and Rollback before making production changes.

When a Failure Occurs

If the layer where the problem exists is unknown, first refer to Locating Problems by Symptom. Common specialized areas include:

Effective Troubleshooting

Do not start by "restarting all services." First, preserve the scene:

  • Start and recovery time of the failure;
  • Affected users, devices, channels, and nodes;
  • Whether it is a total failure or a failure of a specific protocol, browser, or network;
  • Recent changes to configuration, certificates, network, version, or devices;
  • Command ID, Task ID, TraceId, Call-ID, stream identifier, or recording index/file ID;
  • Application, Agent, MediaServer, database, and proxy logs within the same time window.

Check backward from the user's final result to find the first inconsistency. Change only one variable at a time and record the results before and after the change.

Restart Principles

Restarting can restore state, but it also destroys evidence and may affect more users. Execute a restart only after necessary logs have been saved and the scope of impact and rollback plan have been clearly defined. The root cause must still be identified after a restart; "it worked after restarting" is not a valid conclusion to close an issue.