Skip to main content
News Directory 3
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
Menu
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
Tokasaurus: Fast LLM Inference for High Throughput - News Directory 3

Tokasaurus: Fast LLM Inference for High Throughput

June 5, 2025 Catherine Williams Tech
News Context
At a glance
  • A new LLM inference engine called Tokasaurus has been released, promising⁣ critically important improvements in throughput ⁣for demanding workloads.
  • The engine ‍leverages dynamic⁤ Hydragen grouping for smaller‍ models to ⁢reduce CPU overhead and exploit shared prefixes.
  • Modern applications increasingly demand ‍high throughput from LLMs.
Original source: scalingintelligence.stanford.edu

Tokasaurus, a new LLM inference‍ engine, dramatically boosts throughput for demanding workloads. designed to optimize both small and large language models,⁤ Tokasaurus outperforms vLLM and SGLang, delivering significant improvements. Key optimizations include⁣ dynamic Hydragen grouping to reduce CPU overhead for smaller models, and asynchronous tensor parallelism, plus fast ⁢pipeline parallelism for larger models. ⁤Applications like codebase scanning and collaborative model interactions will benefit. Learn how Tokasaurus leverages a greedy depth-frist search algorithm ⁤and⁤ asynchronous managers. News Directory 3 has teh scoop on this advancement. Discover what’s next for this promising technology.

Key Points

  • Tokasaurus is a⁢ new LLM inference ⁢engine designed⁣ for high throughput.
  • It optimizes performance for both⁣ small and large language models.
  • Tokasaurus outperforms vLLM and SGLang in ‍throughput-focused⁢ benchmarks.
  • Key optimizations include minimizing CPU overhead and async tensor parallelism.

tokasaurus: New ⁢LLM Inference Engine Boosts Throughput

⁢ Updated June 05, 2025

A new LLM inference engine called Tokasaurus has been released, promising⁣ critically important improvements in throughput ⁣for demanding workloads. tokasaurus⁤ is⁤ designed to optimize the performance of both small and large language⁣ models, reportedly outperforming existing ‍engines like vLLM and⁢ SGLang by more than threefold in certain benchmarks.

The engine ‍leverages dynamic⁤ Hydragen grouping for smaller‍ models to ⁢reduce CPU overhead and exploit shared prefixes. For⁢ larger models,Tokasaurus employs asynchronous tensor parallelism for GPUs equipped with NVLink,along with a fast pipeline parallelism implementation for those without.

Modern applications increasingly demand ‍high throughput from LLMs. These include scanning ⁢codebases, generating multiple attempts for ⁢problem-solving, and collaborative model interactions. These differ significantly from chatbot ⁣applications, where individual latency is paramount.

Tokasaurus incorporates several key optimizations to ⁢achieve its⁢ performance gains.

Optimizing Small Models

Tokasaurus was benchmarked using the ShareGPT dataset and a ⁢reproduction of ⁣the Large Language ⁢Monkeys experiment,which involves sampling numerous answers‍ to math problems. Results indicated that Tokasaurus⁤ more than doubled the throughput of ⁤other engines on ⁣the Large Language Monkeys workload.

Tokasaurus small models performance
Tokasaurus large batch ⁤sampling results

To minimize CPU overhead,Tokasaurus uses an adaptive,asynchronous manager that monitors the input queue depth and skips⁣ optional steps when necessary.‍ This ensures the GPU is consistently fed with data.

Tokasaurus also ⁢uses a greedy ‍depth-first search algorithm to identify shared prefixes, enabling⁢ more efficient attention computation, notably beneficial for smaller models.

Optimizing Larger Models

For larger models, Tokasaurus utilizes pipeline parallelism (PP) ‍and tensor parallelism (TP) to maximize throughput across multiple GPUs. The PP implementation improves throughput by over threefold compared to⁢ vLLM and⁤ SGLang when using Llama-3.1-70B on eight L40S GPUs.

Tokasaurus⁣ pipeline parallelism performance

for⁣ GPUs with NVLink,Tokasaurus supports ⁤Async-TP,which overlaps inter-GPU communication with computation. ‍The engine dynamically switches on async-TP when the⁤ batch size is large enough to offset the added CPU overhead.

Tokasaurus async tensor parallelism performance

What’s next

Tokasaurus currently supports models from the Llama-3 and qwen-2 families and allows for ⁤any combination ⁣of data, tensor, ⁣and pipeline parallelism within a single node. The developers⁣ hope ⁤its pure Python implementation will encourage forking and modification.

Share this:

  • Share on Facebook (Opens in new window) Facebook
  • Share on X (Opens in new window) X

Related reading

  • Samsung Galaxy S27 Ultra leaks point to a redesigned camera island
  • Yadullah Abidi compares Claude and Gemini on Android apps

Related

Search:

News Directory 3

News Directory 3 catalogs US newspapers, news services, newsstands and digital news outlets across all 50 states. Browse local publishers by city, state, or topic, and follow current headlines linked back to their original sources.

Quick Links

  • Disclaimer
  • Terms and Conditions
  • About Us
  • Advertising Policy
  • Contact Us
  • Cookie Policy
  • Editorial Guidelines
  • Privacy Policy

Browse by State

  • Alabama
  • Alaska
  • Arizona
  • Arkansas
  • California
  • Colorado

© 2026 News Directory 3. All rights reserved.
For contact, advertising, copyright, issues email: office@newsdirectory3.com