01
Inference serving and latency behavior
Core MakoraInference request handling: streaming token delivery, tokens-per-second-per-user characteristics, concurrency handling, and behavior for production AI workloads where the model sits in the user interaction loop.
“Agent-optimized inference for production AI workloads.” www.makora.com
Mapped capabilities
4 capabilities
Streaming completion requests
Accepting a prompt and returning streamed output for a supported model, including partial-output and cancellation behavior.
Concurrent request limits
Behavior at the 1-concurrent (Starter) and up-to-6-concurrent (Developer) ceilings, including queueing versus rejection.
Latency and throughput reporting
How per-user tokens-per-second and time-to-first-token are surfaced or explained to a requesting user.
Long-context request handling
Serving behavior as context length grows, consistent with the context-length claims made in published benchmarks.




