--- Benchmarking Context: qwen2-128k-gpu (8192 tokens) --- Context Response: The first sentence of the given text is "This is a context verification test." Prompt Eval Rate (Context Parsing): 56.30 tokens/s Generation Rate: 17.79 tokens/s VRAM Residency: NAME ID SIZE PROCESSOR CONTEXT UNTIL qwen2-128k-gpu:latest 2ba64284f4e5 547 MB 100% CPU 32768 4 minutes from now