Tokasaurus: Fast LLM Inference for High Throughput
- A new LLM inference engine called Tokasaurus has been released, promising critically important improvements in throughput for demanding workloads.
- The engine leverages dynamic Hydragen grouping for smaller models to reduce CPU overhead and exploit shared prefixes.
- Modern applications increasingly demand high throughput from LLMs.
Tokasaurus, a new LLM inference engine, dramatically boosts throughput for demanding workloads. designed to optimize both small and large language models, Tokasaurus outperforms vLLM and SGLang, delivering significant improvements. Key optimizations include dynamic Hydragen grouping to reduce CPU overhead for smaller models, and asynchronous tensor parallelism, plus fast pipeline parallelism for larger models. Applications like codebase scanning and collaborative model interactions will benefit. Learn how Tokasaurus leverages a greedy depth-frist search algorithm and asynchronous managers. News Directory 3 has teh scoop on this advancement. Discover what’s next for this promising technology.
tokasaurus: New LLM Inference Engine Boosts Throughput
Updated June 05, 2025
A new LLM inference engine called Tokasaurus has been released, promising critically important improvements in throughput for demanding workloads. tokasaurus is designed to optimize the performance of both small and large language models, reportedly outperforming existing engines like vLLM and SGLang by more than threefold in certain benchmarks.
The engine leverages dynamic Hydragen grouping for smaller models to reduce CPU overhead and exploit shared prefixes. For larger models,Tokasaurus employs asynchronous tensor parallelism for GPUs equipped with NVLink,along with a fast pipeline parallelism implementation for those without.
Modern applications increasingly demand high throughput from LLMs. These include scanning codebases, generating multiple attempts for problem-solving, and collaborative model interactions. These differ significantly from chatbot applications, where individual latency is paramount.
Tokasaurus incorporates several key optimizations to achieve its performance gains.
Optimizing Small Models
Tokasaurus was benchmarked using the ShareGPT dataset and a reproduction of the Large Language Monkeys experiment,which involves sampling numerous answers to math problems. Results indicated that Tokasaurus more than doubled the throughput of other engines on the Large Language Monkeys workload.

To minimize CPU overhead,Tokasaurus uses an adaptive,asynchronous manager that monitors the input queue depth and skips optional steps when necessary. This ensures the GPU is consistently fed with data.
Tokasaurus also uses a greedy depth-first search algorithm to identify shared prefixes, enabling more efficient attention computation, notably beneficial for smaller models.
Optimizing Larger Models
For larger models, Tokasaurus utilizes pipeline parallelism (PP) and tensor parallelism (TP) to maximize throughput across multiple GPUs. The PP implementation improves throughput by over threefold compared to vLLM and SGLang when using Llama-3.1-70B on eight L40S GPUs.
for GPUs with NVLink,Tokasaurus supports Async-TP,which overlaps inter-GPU communication with computation. The engine dynamically switches on async-TP when the batch size is large enough to offset the added CPU overhead.
What’s next
Tokasaurus currently supports models from the Llama-3 and qwen-2 families and allows for any combination of data, tensor, and pipeline parallelism within a single node. The developers hope its pure Python implementation will encourage forking and modification.
