In the rapidly evolving landscape of PC gaming technology, few innovations have captured the imagination of hardware enthusiasts quite like Nvidia’s Deep Learning Super Sampling (DLSS). Since the recent emergence of DLSS 5—the latest iteration of Nvidia’s neural rendering technology—the community has been abuzz with speculation regarding its overhead costs and performance ceilings. However, a groundbreaking new development has bypassed traditional limitations, showcasing a technical feat that allows users to offload the heavy lifting of neural rendering to a secondary graphics card. This experimental approach, pioneered by developer Marcelo Guibout, represents a potential paradigm shift for power users. By decoupling the rendering process from the neural post-processing phase, Guibout has demonstrated that it is possible to achieve significantly higher frame rates while simultaneously reducing the thermal load on the primary GPU. While this is far from a standard consumer feature, it provides a fascinating look at the future of distributed computational power in gaming. The Genesis of the "MGPU Bridge" The project, which has been released publicly on GitHub under the name "Neural-coprocessor," functions as a ReShade add-on that enables a multi-GPU workflow. In traditional gaming scenarios, a single GPU is tasked with both the heavy lifting of raw rendering and the subsequent application of neural upscaling. As resolutions and frame targets climb, the neural processing component begins to consume an increasingly large share of the GPU’s total available compute, creating a bottleneck that limits the potential gains of AI-driven upscaling. Guibout’s solution, dubbed "MGPU Bridge," changes this dynamic. The add-on creates a dedicated D3D12 device on a secondary graphics card. Once the primary GPU finishes rendering a frame, the data is sent across the system to the second card. The second GPU then takes over the task of performing the neural rendering—the process that cleans, scales, and enhances the image—before outputting the final result to a display connected directly to that card. A Chronology of the Discovery The demonstration surfaced in the wake of the recent, widespread interest in DLSS 5’s capabilities. Initially appearing as a series of technical proofs-of-concept, the videos shared by Guibout highlighted the technology in two distinct environments: the cinematic, high-fidelity world of The Blood of Dawnwalker and the notoriously demanding streets of Cyberpunk 2077. In the early stages of the experiment, critics questioned whether the latency introduced by transferring frames between two GPUs would negate the performance benefits. Guibout addressed these concerns directly, admitting that the current iteration of the project does indeed double display latency. This makes the technique better suited for cinematic experiences or slower-paced titles rather than competitive e-sports where millisecond-level responsiveness is non-negotiable. Following the initial viral reception of the video demonstrations, Guibout provided the project files to the public, inviting the community to test the configuration. This transition from a private technical showcase to an open-source GitHub project has accelerated the discussion, with users reporting varying levels of success depending on their PCIe bandwidth configurations and multi-GPU setups. Supporting Data: The Efficiency Gap To understand why this method is gaining traction, one must examine the performance metrics. In tests conducted at 1080p resolution, the performance gains are not just marginal—they are transformative. When running DLSS 5 natively on a single card, the neural rendering component "saturates" the device. As the rendering mode shifts from high-quality DLAA down to "Ultra Performance," the raw rendering work decreases significantly, but the cost of the neural post-processing remains relatively static because it must always run at the final output resolution. Consequently, the primary GPU ends up dedicating a disproportionate amount of its resources to the upscaling process rather than the rendering process. By moving the neural workload to a second GPU, the primary card is freed from the upscaling burden. In Guibout’s testing, this resulted in the primary GPU running approximately 21 degrees Celsius cooler. Furthermore, the performance retention rate is significantly improved; where a single-card setup might keep only about a third of the potential gains from a performance mode switch, the dual-GPU configuration retains roughly 86% of the theoretical performance increase. Performance Table (TBOD Demo, 1080p) DLSS Mode DLSS 5 Off (FPS) DLSS 5 on Render Card (FPS) DLSS 5 on 2nd Card (FPS) DLAA 67-70 44 67-70 Quality 98-99 54-55 91 Performance 127-131 59 106-107 Ultra Perf. 172 69-71 157 As evidenced by the data, the "Neural-coprocessor" approach allows the second card to handle the heavy AI lifting with near-negligible performance penalties compared to the baseline of running without DLSS entirely. The Technical Infrastructure: What You Need The hardware requirements for this experiment are non-trivial. Guibout’s setup utilizes a Ryzen 7 7800X3D processor paired with 32GB of DDR5 memory. The graphics configuration involves two Nvidia RTX 5060 Ti 16GB cards, both operating on PCIe 5.0 x8 lanes. Crucially, the setup requires a display to be connected directly to the second GPU, which acts as the display output for the final rendered frames. This requirement is why the technique is currently limited to users who have the physical space and power supply headroom to accommodate two GPUs. Interestingly, the second demo utilized a more modest Ryzen 5 5600 on a DDR4 system, proving that the technique is not restricted to high-end, cutting-edge hardware, provided the PCIe connectivity is sufficient to handle the frame data transfer. Official Stances and Industry Context It is important to note that Nvidia has not officially endorsed or supported this method. In fact, Guibout has been careful to manage expectations, explicitly stating that this is not a revival of SLI (Scalable Link Interface). While SLI was designed to split the rendering of a single frame across multiple GPUs, this new method is a post-processing offload technique. Nvidia’s official vision for DLSS 5 focuses on deep integration within the game engine and the driver, aiming to minimize latency through proprietary hardware-level optimizations. Guibout’s project, by contrast, operates at the application level via ReShade, meaning it lacks the sophisticated, low-level integration that Nvidia provides. This leads to the aforementioned latency issues, as the system must bridge the gap between two independent hardware devices. The Broader Implications: Is This the Future? The implications of this experiment are two-fold. First, it serves as a testament to the sheer potential of AI-driven rendering. If neural processing can be offloaded to a secondary card—or even potentially a dedicated NPU (Neural Processing Unit)—it suggests that future gaming hardware could move toward a more specialized, modular architecture. Second, it highlights a growing frustration among power users regarding the "tax" that modern upscaling features place on the primary GPU. As games become more reliant on AI-based frame generation and upscaling, the power consumption and heat output of top-tier graphics cards have hit an all-time high. A distributed approach, where a secondary, low-power card handles the AI post-processing, could provide a pathway for gamers to achieve high frame rates without requiring massive, power-hungry cooling solutions for a single card. However, the barrier to entry remains high. Most modern gaming motherboards and chassis are not designed to accommodate dual high-end GPUs, and the requirement for a secondary display output makes this a niche solution for enthusiasts. Furthermore, the reliance on third-party software like ReShade means that this method could be broken by future driver updates or anti-cheat implementations in online games. Conclusion Marcelo Guibout’s "MGPU Bridge" project is an impressive display of technical ingenuity. By demonstrating that the neural rendering workload of DLSS 5 is distinct and transportable, he has opened a new conversation about the future of PC gaming performance. While it is unlikely that this will become a mainstream feature adopted by the masses, it provides a valuable blueprint for how developers and hardware enthusiasts can maximize the efficiency of their systems. As we look toward the next generation of gaming, the ability to offload AI tasks to specialized hardware—whether it be a second GPU, an integrated NPU, or a dedicated AI co-processor—may well become a standard feature. For now, the "Neural-coprocessor" remains a playground for those who possess the hardware, the patience, and the curiosity to push the boundaries of what their PCs can do. Whether or not Nvidia chooses to incorporate similar concepts into their official roadmap, the message is clear: the era of offloaded neural rendering has arrived, and it is here to stay. Post navigation The Resurrection of Vice City: How GTA Is Defying Barriers in the Modern Browser