Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
IT之家·Sep 12, 2026, 2:28 AM

DLSS 5 Hybrid Precision Mod Yields Just 1–2% Render Time Drop on RTX 50 Series

Original title:英伟达 DLSS 5 混合精度 Mod 测试:RTX 50 系列显卡性能仅提升 1% 至 2%

Products64

IT Home, September 12 — Following the leak of NVIDIA's DLSS 5 neural rendering DLL file, various community enthusiasts have stepped in to tweak and optimize it, with modders attempting to boost gaming performance by reducing model inference costs. According to information shared by X user SwurvGaming, a mod developer applied FP8 and NVFP4 mixed precision to the DLSS 5 neural rendering model and ran it on an RTX 50-series graphics card.

However, at 4K resolution, this approach only reduced the execution time of the neural rendering stage by about 1% to 2%, with the developer calling the improvement on the Blackwell architecture "very limited."

Since the fifth-generation Tensor Cores in NVIDIA's Blackwell architecture natively support NVFP4, developers attempted to switch part of the DLSS 5 computations from FP8 to NVFP4 to lower inference overhead.

It is worth noting that the 1% to 2% figure only represents the reduction in execution time for the DLSS 5 neural rendering phase, not an overall frame rate increase for the entire game.

This experiment was introduced by the OptiScaler-DLSSNR-PreSR-Multipass project, with the relevant update appearing in version 0.7.1. The project's release notes also pointed out that unsupported model shapes will still fall back to FP8, meaning the mixed-precision approach does not cover the entire computational process.

NVFP4 is a low-precision data format introduced by NVIDIA for neural network inference. Compared to FP8, it reduces memory requirements and leverages the native acceleration capabilities of Blackwell Tensor Cores.

IT Home Note: FP8 and NVFP4 refer to the 8-bit floating-point format and NVIDIA's 4-bit floating-point format, respectively; the latter requires quantization and optimization techniques to function effectively.

The project developer subsequently further optimized the mixed-precision path in version 0.7.2, listing the NVFP4 mixed configuration as the recommended option for Blackwell graphics cards. In a 120-second benchmark run in Baldur's Gate 3, the mixed-precision solution delivered a rendering frame rate of 55.16 FPS compared to 54.57 FPS with the pure FP8 solution—an improvement of roughly 1.08%.

DLSS 5 remains NVIDIA's most computationally demanding DLSS model to date, making the reduction of neural rendering costs a key focus of optimization. NVIDIA previously told Wccftech that DLSS 5 already runs about five times faster than its initial demo at GTC 2026, with further optimizations underway and official support planned for RTX 40-series graphics cards in the future.

Judging from current test results, partially converting the existing FP8 model to NVFP4 has not yet delivered the anticipated significant performance boost. The overall performance of DLSS 5 still depends on factors such as model architecture, quantization schemes, and graphics drivers, and community developers continue to explore more effective optimization paths.

Related Reading: "NVIDIA DLSS 5 Runs on AMD Graphics Card for the First Time: ~30 FPS at 1080P in Cyberpunk 2077" "Foreign Media Tests NVIDIA DLSS 5: Higher-End GPUs See Greater Power Consumption Increases, RTX 5090 Jumps 34% at 4K" "Apple M5 Pro Unofficially Ports NVIDIA DLSS 5: Improved Image Quality, Latency Reaches 240ms" "DLSS 5 Launch Limited to RTX 50, but NVIDIA Confirms Eventual Expansion to RTX 40 Series" "NVIDIA DLSS 5 Announced: RTX 50 Series Support, Debuts with NBA 2K27" "NVIDIA DLSS-NR DLL Surfaces in NBA 2K27 Game Files for the First Time; Massive File Size Hints at Major Feature Upgrade"

Why it's worth reading

It offers an early real-world baseline for grassroots NVFP4 neural rendering optimizations on Blackwell hardware, showing where low-precision gains stall.

Tags

NvidiaDLSS 5BlackwellNVFP4OptiScalerNeural RenderingGaming

Score breakdown

  • Novelty68
  • Impact55
  • Practicality62
  • Credibility70
  • Timeliness65