TensorRT & Stable Diffusion 3.5: Faster RTX Performance
- NVIDIA and Stability AI are collaborating to enhance the performance and efficiency of Stable Diffusion, a popular generative AI image model.
- The base Stable Diffusion 3.5 Large model demands over 18GB of VRAM, limiting its accessibility.
- NVIDIA stated that quantizing Stable Diffusion (SD) 3.5 Large to FP8 reduces VRAM consumption by 40%.
Harness the power of NVIDIA TensorRT and RTX to supercharge your Stable Diffusion 3.5 image generation. This integration reduces VRAM consumption by an impressive 40% while delivering a notable performance boost, with some benchmarks showing speeds more than doubling, thanks to the proprietary NVIDIA TensorRT SDK. Through quantization, NVIDIA’s innovations make running Stable Diffusion, a primary_keyword, more accessible than ever, optimizing model weights for the latest RTX GPUs for rapid results. The new standalone SDK streamlines AI deployments for desktop AI PCs too, making it a key secondary_keyword. discover how News Directory 3 is staying ahead! Eager to learn about the upcoming NIM microservice release and its expected impact? Discover what’s next …
NVIDIA Optimizes stable Diffusion, Reduces VRAM Use with TensorRT and RTX
Updated June 16, 2025
NVIDIA and Stability AI are collaborating to enhance the performance and efficiency of Stable Diffusion, a popular generative AI image model. By leveraging NVIDIA TensorRT and RTX technologies, they’ve achieved significant reductions in VRAM (video random access memory) requirements and substantial performance gains.
The base Stable Diffusion 3.5 Large model demands over 18GB of VRAM, limiting its accessibility. NVIDIA’s solution involves quantization, a process that removes non-critical layers or reduces their precision. NVIDIA GeForce RTX 40 Series GPUs and Ada Lovelace generation NVIDIA RTX PRO GPUs support FP8 quantization,while the latest NVIDIA Blackwell GPUs support FP4.
NVIDIA stated that quantizing Stable Diffusion (SD) 3.5 Large to FP8 reduces VRAM consumption by 40%. Further optimizations using the NVIDIA TensorRT SDK double the model’s performance. TensorRT for RTX, now a standalone SDK, combines performance with on-device engine building, streamlining AI deployment to RTX AI PCs.

Quantization with TensorRT to FP8 reduces the VRAM requirement for SD3.5 Large by 40% to 11GB. this allows more GeForce RTX GPUs to run the model from memory.
TensorRT optimizes model weights and instructions specifically for RTX GPUs. FP8 TensorRT delivers a 2.3x performance boost on SD3.5 Large compared to running the original models in BF16 PyTorch, while using 40% less memory. For SD3.5 Medium, BF16 TensorRT provides a 1.7x performance increase compared to BF16 PyTorch.

The optimized models are available on Stability AI’s Hugging Face page. NVIDIA and Stability AI are also collaborating to release SD3.5 as an NVIDIA NIM microservice.
What’s next
The NIM microservice is expected to be released in July, making it easier for creators and developers to access and deploy the model for various applications, further accelerating generative AI workflows.
