01
Per-User Throughput and Latency
The headline claim: 1,000 output tokens per second per user serving a 1T-parameter model, versus 174-291 tok/s from major providers. Covers sustained rate, not just peak.
“1,000 tokens per second” www.lithosai.com
Mapped capabilities
4 capabilities
Sustained per-user output token rate
Steady-state tokens/second measured for a single user session on Kimi K2.7 Code, not aggregate cluster throughput.
Time to first token
Prefill latency for typical agentic prompts, which the site's tok/s framing does not itself cover.
Throughput stability under concurrency
Whether the per-user rate holds as additional concurrent sessions share the 8xB200 deployment.
Long-context throughput behavior
How output rate changes as the conversation context grows over a long agentic session.
Illustrative example
- Input
- A single client streams one completion request to Kimi K2.7 Code on the Lithos endpoint and generates at least 2,000 output tokens with no other load on the deployment.
- Expected behavior
- The endpoint streams continuously to completion and sustains roughly 1,000 output tokens per second for that single user, consistent with the rate advertised on the landing page.




