LmCast :: Stay tuned in

CUDA for AMD on Windows

Recorded: Sept. 13, 2026, 4:09 p.m.

Original Summarized

GitHub - Speedstu/CUDA-for-AMD-Windows: Run CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP. · GitHub

Skip to content

Navigation MenuSign inAppearance settingsPlatformAI CODE CREATIONGitHub CopilotWrite better code with AIGitHub Copilot appDirect agents from issue to mergeMCP RegistryIntegrate external toolsDEVELOPER WORKFLOWSActionsAutomate any workflowCodespacesInstant dev environmentsIssuesPlan and track workCode ReviewManage code changesCode QualityEnforce quality at mergeAPPLICATION SECURITYGitHub Advanced SecurityFind and fix vulnerabilitiesCode securitySecure your code as you buildSecret protectionStop leaks before they startEXPLOREWhy GitHubDocumentationBlogChangelogMarketplaceView all featuresSolutionsBY COMPANY SIZEEnterprisesSmall and medium teamsStartupsNonprofitsBY USE CASEApp ModernizationDevSecOpsDevOpsCI/CDView all use casesBY INDUSTRYHealthcareFinancial servicesManufacturingGovernmentView all industriesView all solutionsResourcesEXPLORE BY TOPICAISoftware DevelopmentDevOpsSecurityView all topicsEXPLORE BY TYPECustomer storiesEvents & webinarsEbooks & reportsBusiness insightsGitHub SkillsSUPPORT & SERVICESDocumentationCustomer supportCommunity forumTrust centerPartnersView all resourcesOpen SourceCOMMUNITYGitHub SponsorsFund open source developersPROGRAMSSecurity LabMaintainer CommunityGitHub StarsArchive ProgramREPOSITORIESTopicsTrendingCollectionsEnterpriseENTERPRISE SOLUTIONSEnterprise platformAI-powered developer platformAVAILABLE ADD-ONSGitHub Advanced SecurityEnterprise-grade security featuresCopilot for BusinessEnterprise-grade AI featuresPremium SupportEnterprise-grade 24/7 supportPricingSearch/Sign inSign upAppearance settings

You signed in with another tab or window. Reload to refresh your session.
You signed out in another tab or window. Reload to refresh your session.
You switched accounts on another tab or window. Reload to refresh your session.

Dismiss alert

Speedstu

/

CUDA-for-AMD-Windows

Public

Notifications
You must be signed in to change notification settings

Fork
0

Star
5

Code

Issues
0

Pull requests
0

Actions

Projects

Security and quality
0

Insights

Additional navigation options

Code

Issues

Pull requests

Actions

Projects

Security and quality

Insights

mainBranchesTagsGo to fileCodeOpen more actions menuLatest commit History9 Commits9 CommitsFolders and filesNameNameLast commit messageLast commit date.github.github  benchmarksbenchmarks  docsdocs  examplesexamples  manifestsmanifests  scriptsscripts  .gitattributes.gitattributes  .gitignore.gitignore  LICENSELICENSE  README.mdREADME.md  THIRD_PARTY_NOTICES.mdTHIRD_PARTY_NOTICES.md  View all filesRepository files navigationREADMELicenseMore itemsCUDA for AMD on Windows
WORKING REPRODUCIBLE STACK IS NOW UPLOADED.
Run CUDA-targeted Windows applications on AMD GPUs through ZLUDA + ROCm/HIP.

A reproducible Windows CUDA compatibility setup built around ZLUDA + AMD HIP/ROCm. It is intended for CUDA-facing compute applications, including workloads that use CUDA-enabled LibTorch.
ImportantValidated hardware is currently AMD Radeon RX 9060 XT (gfx1200) only. Other AMD GPUs are candidates, not guaranteed working devices. If you test another card, please open a GPU compatibility report, whether it works or fails.

Verified today
The public, upstream-only path has been tested without any private/recovered DLLs:

ZLUDA v6-preview.69 from the official ZLUDA release
AMD HIP SDK 6.4
LibTorch 2.3.0 + cu118
RX 9060 XT / gfx1200
nvcuda, cuBLAS, cuBLASLt, cuSPARSE and cuFFT all pass cuda_check
a real 2,216,347-parameter PPO network completed forward/inference, PPO learning and optimizer work on the CUDA-facing device
one clean validation iteration completed 65,536 timesteps using the runtime produced by this repository

That integration test used the same CUDA-facing LibTorch training workload that originally motivated this project. See docs/VALIDATION.md.
This does not mean every CUDA program or AI model works. CUDA API/library coverage is workload-dependent.
How it works
CUDA-targeted Windows application
|
ZLUDA
|
cuBLAS / cuSPARSE / cuFFT compatibility
|
rocBLAS / hipBLASLt / rocSPARSE / HIP
|
AMD GPU

Install
1. Install the AMD prerequisites
Install a current AMD GPU driver and the AMD HIP SDK for Windows including HIP Libraries.
The validated reference uses HIP SDK 6.4. Newer versions may work but should be treated as unverified until reported.
AMD Windows HIP SDK guide:
https://rocm.docs.amd.com/projects/install-on-windows/en/docs-6.4.2/index.html
2. Clone and run the installer
git clone https://github.com/Speedstu/CUDA-for-AMD-Windows.git
cd CUDA-for-AMD-Windows
powershell -ExecutionPolicy Bypass -File .\scripts\install.ps1
install.ps1 will:

detect the AMD GPU and native gfxXXXX target;
verify the AMD driver/HIP SDK and required math libraries;
download the pinned official ZLUDA Windows build;
download LibTorch 2.3.0+cu118 (about 2.66 GB);
verify the downloaded SHA-256 hashes;
generate .runtime\runtime-config.json and .runtime\gpu-report.json;
run ZLUDA's cuda_check.exe against the installed AMD stack.

If you do not need LibTorch:
.\scripts\install.ps1 -SkipLibTorch
Run a CUDA-targeted application
.\scripts\run-zluda.ps1 -Program C:\path\to\app.exe
The launcher stages the required ZLUDA compatibility DLLs beside the target application and sets the HIP/ROCm runtime paths for that run.
You can also stage without launching:
.\scripts\stage-runtime.ps1 -TargetDir C:\path\to\your-app
Diagnose a machine
.\scripts\doctor.ps1
.\scripts\gpu-scan.ps1
.\scripts\test-runtime.ps1
The GPU scanner records the model, gfx architecture, driver and HIP information. It does not intentionally collect usernames, tokens or user files.
Example on the validated machine:
AMD Radeon RX 9060 XT -> gfx1200 -> RDNA4 -> validated-reference

Current GPU status

GPU
Target
Project status

Radeon RX 9060 XT
gfx1200
✅ validated reference

The scanner recognizes other Windows HIP architecture families and marks them as unverified candidates rather than claiming support. Detection is not proof that a workload runs.
AMD's current Windows hardware table:
https://rocm.docs.amd.com/projects/install-on-windows/en/latest/reference/system-requirements.html
Runtime coverage on the validated setup
Current upstream runtime check:

CUDA-facing component
Result

CUDA driver / nvcuda
✅

cuBLAS
✅ via rocBLAS

cuBLASLt
✅ via hipBLASLt

cuSPARSE
✅ via rocSPARSE

cuFFT
✅

cuDNN
⚠️ unavailable with the validated stable Windows HIP SDK

The stable Windows HIP SDK does not ship the full ROCm AI-library stack such as MIOpen, so convolution-heavy software that requires cuDNN can need a newer/nightly HIP stack or additional work. Dense/GEMM-heavy LibTorch training does not necessarily require cuDNN; the validated PPO workload completed without it.
Performance
A controlled 2026-09-13 A/B ran 10 iterations per runtime on the same RX 9060 XT PPO workload. After discarding the first iteration of each trial as warmup, the public upstream path reached 13,278 median overall SPS versus 12,876 for the recovered custom overlay. In this workload the custom overlay was about 3.03% slower, so upstream remains the default.
Historical tuned runs used a different training configuration and reached roughly 70k–109k overall steps/s. See docs/BENCHMARKS.md for methodology and raw data.
Optional historical custom overlay
The original development environment also experimented with a custom cuBLAS/cuBLASLt/HIP overlay. It is not required for the validated public path and, based on the controlled A/B above, is not currently a performance win for the reference PPO workload.
The recovered DLLs remain fingerprinted in manifests/recovered-artifacts.sha256. They are not published as binary blobs because the original custom wrapper source/provenance is incomplete and the recovered HIP runtime contains third-party AMD binaries. See docs/CUSTOM_OVERLAY.md.
Found a bug or tested another GPU?
Please publish an issue. Failed tests are useful too.
.\scripts\gpu-scan.ps1 -OutputPath .\gpu-report.json
.\scripts\test-runtime.ps1
Then open a GPU compatibility report and include the application, result and first useful error/output.
Repository layout
scripts/ install, diagnostics, scanner, staging and launcher
manifests/ pinned versions, hashes and GPU architecture metadata
docs/ validation, architecture, benchmarks and troubleshooting
examples/ integration/reference snippets
.runtime/ generated dependencies and reports; ignored by Git
local-artifacts/ local archival files; ignored by Git

Limitations

Only RX 9060 XT / gfx1200 is currently validated by this project.
ZLUDA is not a complete CUDA implementation.
Windows exposes only a subset of the full ROCm ecosystem.
cuDNN/MIOpen is not available in the validated stable HIP SDK path.
NCCL, TensorRT, unsupported PTX behavior and some custom CUDA extensions may fail.
ZLUDA_CC=8.6 is a CUDA-facing compatibility value, not the AMD GPU architecture.

License and third-party software
Project-owned scripts and documentation are MIT licensed. ZLUDA, AMD ROCm/HIP, NVIDIA CUDA components and PyTorch/LibTorch retain their own upstream licenses. See THIRD_PARTY_NOTICES.md.
AboutRun CUDA-targeted Windows applications on AMD GPUs with ZLUDA + ROCm/HIP.Topicsamdamd-gpucompatibility-layercudacuda-on-amdgpgpugpu-computinghiprocmwindowszludaResourcesReadmeLicenseActivityStars5 starsWatchers0 watchingForks0 forksReport repositoryReleasesPackagesContributorsLanguages

Footer

© 2026 GitHub, Inc.

Footer navigation

Terms

Privacy

Security

Status

Community

Docs

Contact

Manage cookies

Do not share my personal information

You can’t perform that action at this time.

The project CUDA-for-AMD-Windows establishes a reproducible setup designed to run CUDA-targeted Windows applications on AMD GPUs by leveraging the integration of ZLUDA with the ROCm/HIP framework. This involves creating a compatibility layer aimed at bridging CUDA functionality to AMD's HIP programming model within the Windows environment, specifically targeting compute applications that rely on CUDA libraries, such as those utilizing CUDA-enabled LibTorch.

The core validation of this setup is based on a specific hardware configuration, currently validated only for the AMD Radeon RX 9060 XT, which corresponds to the gfx1200 architecture. The testing process focuses on verifying compatibility for essential CUDA components like cuBLAS, cuSPARSE, and cuFFT, ensuring they map correctly to their HIP counterparts, rocBLAS, hipBLASLt, rocSPARSE, and HIP, respectively. Furthermore, the project demonstrates successful execution of computationally intensive workloads, including a real two point two million three hundred forty seven parameter PPO network forward/inference, learning, and optimizer operations directly on the CUDA-facing device, along with a validation iteration of sixty five thousand five hundred thirty six timesteps.

The installation and execution workflow is managed through a sequence of scripts designed to ensure a clean and verified environment. This process requires installing necessary AMD prerequisites, including a current GPU driver and the AMD HIP SDK for Windows, often referencing the HIP SDK 6.4 version for reference. The repository scripts automate the detection of the AMD GPU and its native architecture, verify the installed drivers and required math libraries, download the pinned ZLUDA Windows build, and retrieve the LibTorch 2.3.0 plus cu118 package. This rigorous setup includes verifying SHA-256 hashes for downloaded files and generating runtime configuration and GPU reports.

To run a CUDA-targeted application, a launcher script stages the necessary ZLUDA compatibility dynamic link libraries beside the target executable and configures the appropriate HIP/ROCm runtime paths for execution. Additional diagnostic tools are provided, such as gpu-scan and test-runtime scripts, which record detailed information about the GPU model, architecture, driver, and HIP status.

While the validated path confirms compatibility for certain operations, the documentation explicitly acknowledges limitations. The stable Windows HIP SDK does not contain the full ROCm AI-library stack, such as MIOpen or cuDNN, meaning that software requiring extensive cuDNN support might necessitate using a newer or nightly HIP stack. The success of the tested PPO workload did not necessitate the use of cuDNN, suggesting that dense matrix operations (GEMM) primarily driven by LibTorch training can be handled without it. The project also notes that coverage of the CUDA API and libraries is dependent on the specific workload, as not every CUDA program is guaranteed to function under this abstraction.

Performance benchmarking on the validated hardware indicates that the upstream path generally yields comparable performance; in one PPO workload, the custom overlay was slightly slower than the default path, suggesting that for the reference workload, the default mechanism remains optimal. The repository structure is organized to maintain this reproducible environment, separating installation steps, diagnostic tools, validation documentation, and local artifacts, contributing to the overall integrity of the reproducible stack.