AKStream.Next · 文档中心

Choosing Chinese-CLIP or SigLIP2 for recording search

Choosing Chinese-CLIP or SigLIP2 for recording search

Choose Chinese-CLIP RN50 for constrained CPUs and RK3588, compare Chinese-CLIP ViT-B/16 first for Chinese retrieval on capable PCs and GPUs, and evaluate SigLIP2 Base 224 for multilingual or broader semantic tasks. A newer model does not automatically retrieve Chinese surveillance scenes more accurately.

This guide separates public retrieval benchmarks, device measurements and unverified deployment paths. Final selection requires the same labeled recordings, queries and sampling policy.

Model size and purpose

Property Chinese-CLIP RN50 Chinese-CLIP ViT-B/16 SigLIP2 Base Patch16-224
Image input 224×224 224×224 224×224
Vision architecture ResNet50 ViT-B/16 ViT-B/16 backbone with pooling head
Full vision tower About 38M parameters About 86M About 92.9M; backbone about 86M
Text tower About 39M About 102M About 282.3M
Complete model About 77M About 188M About 375.2M
Product embedding dimensions 1024 512 768
Main reason to evaluate Smaller inference budget Stronger public Chinese retrieval results Multilingual and broader semantic tasks
Main cost Fine semantic matching needs validation Heavier image and text towers Much larger text tower and resident footprint

Chinese-CLIP sizes come from the official model table. Tensor shapes in the Google SigLIP2 checkpoint give 92,884,224 vision parameters, 282,303,744 text parameters and two scalars: 375,187,970 total. The often cited 86M backbone is not the complete vision tower including its pooling head.

Continuous indexing primarily runs the image encoder. Queries run the text encoder, but infrequent queries do not eliminate the memory of warmed or cached text models. Separate image-only compute from the complete resident model.

Public Chinese retrieval results

The following are official zero-shot R@1 percentages, not fine-tuned results. Improvement is measured in percentage points.

Dataset and direction RN50 ViT-B/16 Improvement
Flickr30K-CN, text to image 48.8 62.7 13.9
Flickr30K-CN, image to text 60.0 74.6 14.6
MUGE, text to image 42.6 52.1 9.5
COCO-CN, text to image 48.1 62.2 14.1
COCO-CN, image to text 51.6 57.0 5.4

See official Results.md. This MUGE table only publishes text-to-image results. The scores favor evaluating ViT-B/16 for Chinese retrieval, but cannot be converted into surveillance accuracy or missed-recording rates.

The SigLIP2 paper covers multilingual semantic understanding, localization and dense features. It does not establish a controlled Chinese surveillance comparison against these Chinese-CLIP models. Do not label SigLIP2 a proven higher-quality surveillance option solely because it is newer.

Use labeled examples to distinguish wrong objects, colors, actions and relationships. Whole-image embedding is not a small-object detector; larger models alone do not solve distant objects, occlusion or poor night imagery.

Hardware and product support

Platform Model selection Product validation boundary
NVIDIA CUDA RN50 for lower cost; compare ViT-B/16 for Chinese retrieval; SigLIP2 for multilingual needs Three models tested on Linux V100, including shared-frame paths; no universal NVIDIA throughput guarantee
Intel OpenVINO GPU All three are candidates subject to resources and dataset results Linux B580 basic retrieval tested; Windows B580 RN50 recording/search and concurrent queries passed; other models remain unverified
OpenVINO CPU RN50 is a reasonable candidate to evaluate before larger towers The current product Cpu backend uses ONNX Runtime CPU, not an accepted OpenVINO CPU implementation
RK3588 / RKNN RN50 is the current edge preference; larger models need stricter budgets All three converted models passed native Linux RK3588 DMA-path checks; new Docker acceptance is tracked separately
AMD GPU Validate the actual inference stack before comparing models Windows x64 DirectML is integrated; RN50 towers passed a Radeon 610M probe. Other models/devices remain unverified; Linux AMD is not integrated

CUDA is NVIDIA-specific. Intel GPU inference is not AMD acceleration. Windows AMD can be researched through DirectML; Linux AMD can be researched through MIGraphX. Windows x64 DirectML is now integrated; Linux MIGraphX is not. The legacy ROCm provider documentation also describes its removal from ORT 1.23. Video decoding acceleration is a separate capability.

Transformer conversion requires checking attention, normalization, graph operations, CPU assistance and transfers. Do not infer performance from NPU TOPS. Existing converted model versions work, but future checkpoints require new conversion and acceptance checks.

Measured warm image latency

These samples are from accepted model versions around 2026-10-02. They are warm single-image path medians, not cold startup or promised whole-system recording capacity.

Path, 50 samples per model RN50 ViT-B/16 SigLIP2 Base 224
CUDA V100, shared-frame image path 5.6991 ms 11.4017 ms 13.6378 ms
RK3588, DMA image path About 48.7 ms About 188.7 ms About 188.2 ms

The larger vision models were close in the RK3588 sample. Different hardware, precision, graph conversion and core allocations can change their relative behavior. An RN50 probe on Windows Radeon 610M measured about 36.9ms for image inference and 9.2ms for text after warm-up. These are model latencies, not end-to-end recording FPS or a three-model ranking.

All three passed assembly and basic retrieval checks. Controlled accuracy on labeled Chinese surveillance footage remains outstanding. Similarity values from different embedding spaces and agreement with reference output are not quality rankings.

Memory and utilization

Parameter storage is parameter count multiplied by bytes per parameter. Decimal MB below is a theoretical weight budget, not file size, process RSS, peak VRAM or a promise of accepted INT8 support.

Model and scope FP32 FP16 INT8
RN50 vision, about 38M 152 MB 76 MB 38 MB
ViT-B/16 vision, about 86M 344 MB 172 MB 86 MB
SigLIP2 complete vision, about 92.9M 372 MB 186 MB 93 MB
RN50 complete towers, about 77M 308 MB 154 MB 77 MB
ViT-B/16 complete towers, about 188M 752 MB 376 MB 188 MB
SigLIP2 complete towers, about 375.2M 1,501 MB 750 MB 375 MB

Real peaks include activations, workspaces, arenas, graph caches, buffers and parallel instances. Quantization also requires calibration and accuracy checks.

Execution CPU work Accelerator work Memory to measure
ONNX CPU Decode, preprocessing, inference and retrieval support None for image/text GPU inference RSS and system availability
CUDA / Intel GPU Decode, preprocessing, submissions and postprocessing Image/text inference and workspaces RSS, peak VRAM and per-instance allocations
RK3588 NPU Decode, conversion, submissions and auxiliary operators Inference and shared-memory buffers System RAM, DMA buffers and runtime allocations
Vector indexing Writes, index construction and queries Separate from encoder utilization Vectors, graph index, payload and caches

Low-rate sampling can create short busy bursts followed by idle periods. A process reporting 100% CPU may occupy one core. Measure latency, sustained throughput, peak memory, power and application response together.

Sampling and capacity

Requested image inference rate is participating channels multiplied by sampling FPS, not camera FPS. Image embeddings populate the index; text embeddings query the corresponding model space.

flowchart LR
  V[Recording stream] --> S[Budgeted frame sampling] --> I[Image encoder] --> Q[Model vector index]
  T[Query text] --> E[Text encoder] --> Q
  Q --> R[Locate and play recordings]

Sixteen channels at two sampled frames per second request 32 inferences/second. The RK3588 RN50 median suggests a single-chain ideal ceiling near 20.5/second before other work, so that load is not automatically comfortable. At one sample every five seconds, demand is 3.2/second. Requested rate multiplied by latency gives about 15.6% single-chain inference busy time for RN50 and 60.4% for ViT-B/16. Those are not CPU/NPU utilization measurements or linear-scaling guarantees.

The product also limits total node inference rate; requested sampling above that budget is deferred. Raw FP32 vectors alone cost approximately 4.096 GB, 2.048 GB and 3.072 GB per million entries for RN50, ViT-B/16 and SigLIP2 respectively, before indexing and metadata. The lightest encoder does not necessarily produce the smallest vector store.

Selection and evaluation

Goal Initial choice Verify before rollout
Weak CPU, RK3588, tight memory RN50 End-to-end resource budget, stability, coverage
Chinese retrieval, ample GPU compute Compare ViT-B/16 Same-data accuracy and latency against RN50
Multilingual tasks Evaluate SigLIP2 Real Chinese footage, text memory and cold start
AMD GPU deployment Windows x64 DirectML RN50 Other devices/models still need driver and operator checks

Use RN50 for constrained edges, compare ViT-B/16 on Chinese workloads with enough compute, and evaluate SigLIP2 for multilingual tasks. Keep the same recordings, sampling, queries and ground truth. Record R@1/R@5, false hits, missed segments, time localization, cold/warm latency, indexing load and concurrent queries.

Weights, tokenizers and preprocessing define the embedding space. Models use separate indexes; never mix RN50 vectors with ViT-B/16 or SigLIP2 queries. RN50 remains the product default.

Download and offline installation

Model weights and large acceleration runtimes are separate downloads. First-run installation can authorize downloads; dependencies can also be prepared later. The website offers R2 and local-proxy links, with SHA256 validation. Offline users download the correct model/backend bundle on another computer, verify and transfer it, extract into the model directory, install platform runtimes and restart as instructed. Download permission does not activate search or replace feature licensing.

Windows x64 and ARM64 packages bundle their native DirectML and ONNX Runtime libraries. No R2 upload or DirectML download is needed; ONNX model weights remain separate. AMD RN50 on x64 was validated on hardware. ARM64 native libraries and Qualcomm GPU detection are integrated; Adreno inference still needs hardware validation. A virtual machine may not expose a compatible DirectX 12 GPU. Restart services after changing backends.