Skip to main content
News Directory 3
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
Menu
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
TensorRT & Stable Diffusion 3.5: Faster RTX Performance - News Directory 3

TensorRT & Stable Diffusion 3.5: Faster RTX Performance

June 16, 2025 Catherine Williams Tech
News Context
At a glance
  • NVIDIA and Stability AI are collaborating⁣ to ‍enhance the performance and efficiency of Stable Diffusion, a popular generative⁢ AI image model.
  • The base Stable Diffusion 3.5 Large ⁤model demands over 18GB of VRAM, limiting its accessibility.
  • NVIDIA stated that quantizing Stable Diffusion (SD) 3.5‍ Large to FP8⁢ reduces VRAM⁤ consumption by 40%.
Original source: blogs.nvidia.com

Harness the power of NVIDIA TensorRT and RTX to supercharge your Stable Diffusion 3.5 image generation. This integration⁣ reduces VRAM consumption by an impressive 40% while delivering a notable⁣ performance boost, with some benchmarks showing speeds more than doubling, thanks to the ⁤proprietary NVIDIA TensorRT SDK. Through ⁤quantization, NVIDIA’s innovations make running‍ Stable ⁤Diffusion, a primary_keyword, more⁣ accessible than ever,⁤ optimizing model weights for the latest RTX GPUs ⁣for rapid ⁢results. The new standalone SDK streamlines AI deployments for desktop ⁢AI PCs too,⁢ making it a key ⁤secondary_keyword.⁢ discover ⁣how News Directory 3 is staying ahead! Eager to learn about the upcoming NIM microservice release and its expected impact? ‍Discover what’s⁤ next …


NVIDIA Reduces Stable Diffusion VRAM Use with TensorRT, RTX










key Points

Table of Contents

    • key Points
  • NVIDIA Optimizes stable Diffusion, Reduces VRAM Use with TensorRT and RTX
    • What’s next
    • Further reading
  • NVIDIA and Stability AI optimize ‍Stable Diffusion 3.5.
  • TensorRT and RTX reduce VRAM consumption by 40%.
  • Performance doubles with NVIDIA ‍TensorRT SDK.
  • TensorRT for RTX now a standalone SDK.

NVIDIA Optimizes stable Diffusion, Reduces VRAM Use with TensorRT and RTX

Updated June 16, 2025
‍

NVIDIA and Stability AI are collaborating⁣ to ‍enhance the performance and efficiency of Stable Diffusion, a popular generative⁢ AI image model. By leveraging NVIDIA TensorRT and RTX technologies, they’ve achieved significant reductions ⁢in VRAM (video random access memory) requirements and substantial performance gains.

The base Stable Diffusion 3.5 Large ⁤model demands over 18GB of VRAM, limiting its accessibility. NVIDIA’s⁤ solution ⁢involves quantization, a process that removes non-critical layers or reduces their precision. NVIDIA GeForce RTX 40 Series GPUs and⁤ Ada Lovelace generation NVIDIA RTX PRO GPUs support FP8 quantization,while the latest NVIDIA⁢ Blackwell GPUs support FP4.

NVIDIA stated that quantizing Stable Diffusion (SD) 3.5‍ Large to FP8⁢ reduces VRAM⁤ consumption by 40%. Further optimizations using the NVIDIA TensorRT⁣ SDK double the model’s performance. TensorRT for RTX, now a standalone SDK, combines performance with on-device engine building, streamlining AI deployment to RTX AI PCs.


Comparison of Stable Diffusion 3.5 image generation with FP16 (left)⁤ and FP8 (right) showing similar quality with faster generation on FP8.
Stable Diffusion 3.5 quantized FP8 (right) generates images in half ‍the time with similar quality as ⁣FP16 (left). Prompt: ‍A serene mountain lake at sunrise,⁢ crystal clear water reflecting snow-capped ⁢peaks, lush pine‍ trees along the‍ shore, soft morning mist, photorealistic, vibrant colors, high‍ resolution.

Quantization with TensorRT to FP8 reduces the VRAM requirement for SD3.5 Large by 40% to 11GB. this allows more GeForce RTX GPUs to run the model from ‍memory.

TensorRT ⁣optimizes model weights and instructions‍ specifically for RTX GPUs. FP8 ⁣TensorRT delivers a 2.3x performance boost on ⁣SD3.5 Large ⁢compared to⁣ running the original models ‍in BF16 PyTorch, while using 40% ⁢less memory. For SD3.5 Medium, BF16 TensorRT provides a 1.7x performance increase compared to BF16‍ PyTorch.


Graph showing performance boost⁤ of Stable⁤ Diffusion 3.5 with FP8 TensorRT.
FP8 TensorRT boosts ‍SD3.5 Large performance by 2.3x vs. BF16 PyTorch, with 40% less memory use. For SD3.5 Medium, BF16 TensorRT⁢ delivers a 1.7x speedup.

The optimized models are available on Stability AI’s Hugging Face page. NVIDIA and Stability AI are also collaborating to release ⁣SD3.5⁣ as an ⁢NVIDIA NIM microservice.

What’s next

The NIM microservice is expected to be released in July, making it easier for creators and developers to access and deploy the model‍ for various applications, ‍further accelerating generative AI workflows.

Further reading

  • NVIDIA TensorRT for RTX Introduces ⁣An Optimized Inference AI ‍Library on Windows

Share this:

  • Share on Facebook (Opens in new window) Facebook
  • Share on X (Opens in new window) X

Keep reading

  • vivo X500 Series and Dimensity 9600 Leaks: Specs, 2nm Chip, and September Launch
  • The Origin of Electronic Health Records: The Legacy of MIS-I

Related

Search:

News Directory 3

News Directory 3 catalogs US newspapers, news services, newsstands and digital news outlets across all 50 states. Browse local publishers by city, state, or topic, and follow current headlines linked back to their original sources.

Quick Links

  • Disclaimer
  • Terms and Conditions
  • About Us
  • Advertising Policy
  • Contact Us
  • Cookie Policy
  • Editorial Guidelines
  • Privacy Policy

Browse by State

  • Alabama
  • Alaska
  • Arizona
  • Arkansas
  • California
  • Colorado

© 2026 News Directory 3. All rights reserved.
For contact, advertising, copyright, issues email: office@newsdirectory3.com