Monitoring, Logging, and Alerting
Video failures span devices, signaling, media, and storage. Monitoring must cover the entire chain and be able to link a single operation using correlation identifiers.
flowchart LR
A[User request] --> B[Device and protocol]
B --> C[MediaServer]
C --> D[Playback or recording]
D --> E[Final user result]
T[TraceId / CommandId / StreamId] -.correlates.-> A
T -.correlates.-> E
Monitoring Layers
| Layer | Key Observations |
|---|---|
| Host | CPU, Memory, Swap, File Descriptors, NIC, Clock, Disk |
| AKStream.Next | Request Errors, GC, Thread Pools, Configuration Version, Task Status |
| NodeAgent | Heartbeat, Provider, Version, Recent Operations, Offline Back-filling |
| MediaServer | Process, Version, Node ID, Streams, Players, Recording, WebRTC |
| Database | Connections, Slow Queries, Locks, Space, Backups, Schema |
| Protocol | Device Online Status, Registration, Command Success/Timeout, Subscription, Signaling Errors |
| Recording | Active Sessions, File Output, Scan/Trim Backlog, Failures |
| RTC | Rooms, Members, Media Sessions, ICE, Release and Event Backlog |
Log Sources
Unified collection of application structured logs, NodeAgent operation logs, MediaServer stdout/stderr and module logs, protocol message summaries, configuration/lifecycle/security audits, and channel runtime timelines.
Log timestamps must be synchronized and include environment, node, and component information. Do not log database passwords, API Tokens, MediaServer secrets, plain-text device credentials, full Cookies, or authorization private keys.
Correlation Identifiers
Record based on scenario:
- HTTP: TraceId, User/Token summary, Route, Status Code.
- Asynchronous Commands: commandId/jobId, Target Resource, Final Status.
- GB28181: Device ID, Channel ID, Call-ID, SN, RTP Port.
- ONVIF: Device ID, Profile/VideoSource token, Operation Name.
- Media: mediaServerId, vhost, app, stream.
- Recording: sessionId, fileId, Physical Node, Time Range.
- RTC: roomId, participantId, mediaSessionId.
Once support personnel obtain any stable identifier, they should be able to trace backward to the request and forward to the business result.
Recommended Alerts
- Main service health check fails consecutively or readiness is blocked.
- NodeAgent or MediaServer exceeds the heartbeat window.
- Configuration version drift or prolonged pending restart.
- Sudden spikes in protocol timeout rates, authentication failure rates, or device offline rates.
- Active recording sessions with no new files produced beyond the shard cycle.
- Bytes, inodes, or write probes fall below the threshold.
- Growth in backlogs for WebHooks, recording scans, trimming, or asynchronous tasks.
- Continuous growth of RTC session release failures.
- TLS certificates approaching expiration or renewal failure.
Alerts should specify the "affected object, current value, threshold, duration, recent changes, and the primary troubleshooting document."
Retention and Sampling
Retention periods for security audits, authorization, and deletion operations are typically longer than for general debug logs. High-frequency success events may be sampled, but failures, timeouts, permission denials, and state transitions must not be unconditionally sampled out.
Alert Closed-Loop
After an alert is triggered, record the acknowledgment, handler, impact, associated ticket, recovery time, and root cause. Acknowledgments, ignores, and recoveries in batch alert handling must have their results preserved separately; do not treat device arming states such as ONDUTY/OFFDUTY as new alert events.