Model loading took 8.01 GiB memory and 66.190265 seconds GPU KV cache size: 867,999 tokens, Maximum concurrency for 8,192 tokens per request: 105.96x