LLM Inference (256K Context)
Supports a 256K ultra-long context window capable of processing entire book-length texts in a single pass. Benchmark performance: first-token latency of ~1.4s on short inputs (~63s on full long contexts) and generation speed of ~70 to 86 tokens/second.