In Docker + cgroup v2 with no memory limit set, /sys/fs/cgroup/memory.current
returns a small container number while /sys/fs/cgroup/memory.max is "max".
readCgroupMemoryBytes then returned (used=<container>, total=0, ok=true).
collectSystemStats used that container "used" but, being unable to derive a
cgroup total, filled the total from the host via gopsutil. The dashboard then
computed container_used / host_total, e.g. ~60MB / 23GB ≈ 0.3% — wildly
understating real usage.
Fix: introduce resolveMemoryStats, which picks a single self-consistent
(used, total, percent) trio from ONE source. cgroup metrics are used only when
the cgroup exposes both a current usage AND a concrete limit (memory.max != max,
so total > 0); otherwise used/total/percent all fall back to the host reading.
The two sources are never mixed.
- memory.current valid + memory.max = "max" -> all host metrics
- memory.current = 512MiB + memory.max = 2GiB -> ~25% from cgroup
- no cgroup (bare metal) -> all host metrics
CPU metric behavior is unchanged (cgroup attempt then host fallback).
Adds ops_metrics_collector_memory_test.go covering all branches.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ConfigManager.refreshLoop reloads the Prompt Guard config every 5s and
Reload logged config_loaded on every successful load, so an unchanged
config produced up to ~17k identical lines per instance per day and
buried real configuration changes.
Log the event only when the reload carries news: the first snapshot, a
new config version (every admin save bumps it under the advisory lock in
UpdateConfig), a flip of the global risk control gate (a separate setting
that leaves the version untouched), or a recovery from a failed reload so
the degraded-to-healthy transition stays visible.
This mirrors logInvalidTokenEndpoints, which already warns once per
change rather than on every refresh.
CN-provider accounts explicitly configured with api_protocol=anthropic
(the frontend offers per-provider presets such as GLM Anthropic) fell
through the account-test router to the generic Claude tester, which:
- appended ?beta=true, which the native passthrough deliberately avoids
for third-party Anthropic endpoints, and
- defaulted a missing base_url to https://api.anthropic.com — sending
the provider's API key to Anthropic instead of the provider's own
Anthropic-compatible endpoint (guaranteed 401 plus cross-provider
credential exposure).
Route them to a dedicated probe that mirrors the adaptive Anthropic
check: resolve the endpoint via GetAnthropicProtocolBaseURL (same
resolution as real /v1/messages forwarding, including per-platform
defaults), use the shared API-key auth header, and mark 401/403 as the
account error state.
Also fail fast with an actionable hint when the anthropic-protocol
base_url still points at an OpenAI-compatible endpoint (paas path,
version segment, or chat/completions suffix): the naive {base}/v1/messages
join would 404 (e.g. /api/paas/v4/v1/messages) without any pointer to
the actual misconfiguration.