Files
sub2api/backend/internal
shunwang-cryptoandClaude Opus 4.8 cd05772e91 fix(ops): avoid mixing cgroup and host memory metrics
In Docker + cgroup v2 with no memory limit set, /sys/fs/cgroup/memory.current
returns a small container number while /sys/fs/cgroup/memory.max is "max".
readCgroupMemoryBytes then returned (used=<container>, total=0, ok=true).

collectSystemStats used that container "used" but, being unable to derive a
cgroup total, filled the total from the host via gopsutil. The dashboard then
computed container_used / host_total, e.g. ~60MB / 23GB ≈ 0.3% — wildly
understating real usage.

Fix: introduce resolveMemoryStats, which picks a single self-consistent
(used, total, percent) trio from ONE source. cgroup metrics are used only when
the cgroup exposes both a current usage AND a concrete limit (memory.max != max,
so total > 0); otherwise used/total/percent all fall back to the host reading.
The two sources are never mixed.

- memory.current valid + memory.max = "max"  -> all host metrics
- memory.current = 512MiB + memory.max = 2GiB -> ~25% from cgroup
- no cgroup (bare metal)                       -> all host metrics

CPU metric behavior is unchanged (cgroup attempt then host fallback).

Adds ops_metrics_collector_memory_test.go covering all branches.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-22 19:35:31 +08:00
..