AKStream.Next · 文档中心

Monitoring, Logging, and Alerting

Monitoring, Logging, and Alerting

Video failures span devices, signaling, media, and storage. Monitoring must cover the entire chain and be able to link a single operation using correlation identifiers.

flowchart LR
    A[User request] --> B[Device and protocol]
    B --> C[MediaServer]
    C --> D[Playback or recording]
    D --> E[Final user result]
    T[TraceId / CommandId / StreamId] -.correlates.-> A
    T -.correlates.-> E

Monitoring Layers

Layer Key Observations
Host CPU, Memory, Swap, File Descriptors, NIC, Clock, Disk
AKStream.Next Request Errors, GC, Thread Pools, Configuration Version, Task Status
NodeAgent Heartbeat, Provider, Version, Recent Operations, Offline Back-filling
MediaServer Process, Version, Node ID, Streams, Players, Recording, WebRTC
Database Connections, Slow Queries, Locks, Space, Backups, Schema
Protocol Device Online Status, Registration, Command Success/Timeout, Subscription, Signaling Errors
Recording Active Sessions, File Output, Scan/Trim Backlog, Failures
RTC Rooms, Members, Media Sessions, ICE, Release and Event Backlog

Log Sources

Unified collection of application structured logs, NodeAgent operation logs, MediaServer stdout/stderr and module logs, protocol message summaries, configuration/lifecycle/security audits, and channel runtime timelines.

Log timestamps must be synchronized and include environment, node, and component information. Do not log database passwords, API Tokens, MediaServer secrets, plain-text device credentials, full Cookies, or authorization private keys.

Correlation Identifiers

Record based on scenario:

  • HTTP: TraceId, User/Token summary, Route, Status Code.
  • Asynchronous Commands: commandId/jobId, Target Resource, Final Status.
  • GB28181: Device ID, Channel ID, Call-ID, SN, RTP Port.
  • ONVIF: Device ID, Profile/VideoSource token, Operation Name.
  • Media: mediaServerId, vhost, app, stream.
  • Recording: sessionId, fileId, Physical Node, Time Range.
  • RTC: roomId, participantId, mediaSessionId.

Once support personnel obtain any stable identifier, they should be able to trace backward to the request and forward to the business result.

  • Main service health check fails consecutively or readiness is blocked.
  • NodeAgent or MediaServer exceeds the heartbeat window.
  • Configuration version drift or prolonged pending restart.
  • Sudden spikes in protocol timeout rates, authentication failure rates, or device offline rates.
  • Active recording sessions with no new files produced beyond the shard cycle.
  • Bytes, inodes, or write probes fall below the threshold.
  • Growth in backlogs for WebHooks, recording scans, trimming, or asynchronous tasks.
  • Continuous growth of RTC session release failures.
  • TLS certificates approaching expiration or renewal failure.

Alerts should specify the "affected object, current value, threshold, duration, recent changes, and the primary troubleshooting document."

Retention and Sampling

Retention periods for security audits, authorization, and deletion operations are typically longer than for general debug logs. High-frequency success events may be sampled, but failures, timeouts, permission denials, and state transitions must not be unconditionally sampled out.

Alert Closed-Loop

After an alert is triggered, record the acknowledgment, handler, impact, associated ticket, recovery time, and root cause. Acknowledgments, ignores, and recoveries in batch alert handling must have their results preserved separately; do not treat device arming states such as ONDUTY/OFFDUTY as new alert events.