Backup, Recovery, Upgrade, and Rollback
Upgrading is not simply copying new files. A reliable change must treat the program version, configuration, database schema, security materials, MediaServer, and recording index as a single recoverable entity.
Before executing a formal package upgrade, read Online upgrade: complete guide to akn upgrade and page-based upgrade. This page covers the broader governance principles for backup, recovery, and rollback.
Preserve deployment overrides
Node addresses and policies saved in the UI may reside in Config/akstream.next.Deployment.json; it is not a disposable installer file. Installation and upgrade retain existing values, including false, zero, empty strings, null, arrays, and unknown fields. Only new installation parameters missing from both the main configuration and the override are added. New defaults do not shadow managed paths, flags, or FFmpeg settings already defined in the main configuration.
The managed ZLMediaKit ZLMediaKit/config.ini is also mutable runtime configuration, not a program file that can be replaced by package defaults. Upgrade uses the new package file as the structural baseline and adds newly introduced sections and keys, while every existing value, custom key, and custom section from the installed file keeps precedence. User-managed native latency, track-wait, RTSP/RTP, HLS, and other tuning therefore stay unchanged across an upgrade. Official ZLMediaKit files may split one logical section across multiple headers; the installer merges those headers, while duplicate keys or malformed lines stop the upgrade before the old services are stopped rather than silently choosing one value. After startup, media identity, Secret, hooks, security policy, managed listeners, RTC, and other keys explicitly owned by AKStream.Next are still reconciled from the preserved AKStream.Next configuration; that is not a package-default overwrite.
Malformed JSON and duplicate or case-conflicting keys are rejected before stopping the old service. Changes use an atomic same-directory replacement with an akstream.next.Deployment.json.upgrade-backup.* backup. A retired bootstrap override is not recreated when no new override is needed. Runtime/GC merging still uses the new runtime structure while preserving existing configProperties values.
If an older upgrade already lost an address, inspect the upgrade transaction's Config backup and restore the affected fields through the target node's configuration page. Do not overwrite the entire new configuration with an old file or publicly share backups containing credentials.
flowchart LR
A[Consistent backup] --> B[Recovery test in isolation]
B --> C[Production upgrade]
C --> D{Critical paths pass?}
D -->|Yes| E[Observe and complete]
D -->|No| F[Rollback the consistent state]
Backup Scope
| Content | Why it is Required |
|---|---|
Config |
Main configuration, node and media configurations |
Data/Security |
Tokens, signatures, and trusted materials |
| Database | Devices, channels, tasks, recording index, permissions, and audits |
| Recordings and Clip Files | Actual media evidence; can be fully backed up or stored independently based on business needs |
| Reverse Proxy and Certificates | Domain names, TLS, WebSocket, and media proxies |
| Installation Packages and Checksums | To roll back to the exact previous version |
| License Files/Public Keys | To restore authorization status and verification chains |
Databases and files must share a consistent point-in-time. Backing up only the index without the files, or vice versa, will not allow for a complete recovery of recording services.
Recovery Drills
Perform the following regularly in an isolated environment:
- Prepare a target system and database compatible with the production environment.
- Restore configurations, security directories, and the database.
- Restore or remount the recording root.
- Start using a fixed older version and verify the connection between nodes and the MediaServer.
- Spot-check login, devices, live streams, recording retrieval, and downloads.
- Record the RPO, RTO, missing data, and manual steps required.
"Backup task successful" does not equal "recoverable."
Pre-Upgrade
- Read the target version notes and known limitations.
- Verify the target operating system/RID package, version, and SHA-256.
- Back up and record the recovery point.
- Check that there are no abnormal backlogs in the database, recording scans, clipping, or asynchronous tasks.
- Perform regression testing in pre-production using real devices and critical user paths.
- Define the downtime window, observation period, rollback decision-maker, and commands.
Upgrade Execution
- Stop adding new high-risk tasks and wait for critical recordings or clipping to finish.
- Stop/upgrade services according to the deployment script sequence; do not directly overwrite files currently in use.
- Apply database and configuration migrations, and save the migration logs.
- After startup, first check health, schema, configuration versions, and nodes.
- Then verify devices, the first stream, recordings, RTC, and third-party integrations.
- Observe error rates, resource usage, protocol timeouts, and file output.
Rollback
Rollback trigger conditions should be defined before the upgrade, such as failure of core user paths, data migration errors, persistent recording interruptions, or uncontrollable performance degradation.
Program rollbacks must be compatible with the database and configuration. If the new version performed an irreversible data migration, you cannot simply swap back to the old binary; you must restore a complete consistent backup or execute verified downgrade steps.
Multi-node Sequence
Upgrade in batches to maintain a protocol compatibility window. Upgrade verification nodes first, then expand the scope. Lossless migration of existing media sessions across nodes is not guaranteed; notify users of the impact and choose a low-traffic window before operating.
Completion Criteria
- The target version is consistent across all components, with no health or configuration drift.
- Critical user paths have passed, and no new high-priority anomalies occurred during the observation period.
- Backups, upgrade logs, test evidence, and actual changes have been archived.
- Rollback materials are retained until the end of the observation period.