mirror of
https://github.com/Wei-Shaw/sub2api.git
synced 2026-10-07 17:27:58 +08:00
Cache Hit Rate was calculated as cache_read / (cache_read + cache_creation), which always yields 100% for OpenAI models since cache_creation is never reported by the OpenAI API. The denominator should include all prompt tokens (input_tokens + cache_read_tokens + cache_creation_tokens) so the rate reflects the actual percentage of input tokens served from cache. Fixes #2291