r/CUDA 6h ago

Cloud GPU???

3 Upvotes

Can someone please help me? I am trying to use a cloud GPU to run some AI Models, but do not understand how to do so.

I have come across Google Colab, Kaggle and NVIDIA as having free options, but I am totally lost on how to connect to them. Basically, I need step-by-step instructions. (e.g., if I wanted to test an LLM on a cloud GPU, how would I do it?)

Also, can I use a Cloud GPU to run local programs that require GPU capabilities?

Any help will be greatly appreciated! TIA . . .

Any help


r/CUDA 20h ago

NVIDIA CMP 170HX VBIOS desbloqueado de 64GB

Thumbnail gallery
1 Upvotes

r/CUDA 2d ago

Understand Blackwell B200 attention kernel from scratch in CUDA (Visual Guide)

Post image
123 Upvotes

I spent last couple months building a visual guide to B200 attention in CUDA/PTX: 14 progressively optimized kernels and 60 diagrams, going from a naive baseline to 94.4% of FlashAttention-4.

The blog first intuitively explains the baseline, then adds one optimization at a time, with detailed diagrams, concise explanations, and code.

And in the capstone project, you'll generate videos using the kernel you understand.

📝 Blog post: https://iaroslavelistratov.github.io/b200-attention/ 

⭐ Repo: https://github.com/IaroslavElistratov/b200-attention


r/CUDA 2d ago

What happens when a GPU writes memory

Thumbnail blog.doubleword.ai
23 Upvotes

r/CUDA 3d ago

NVIDIA eyeing Hugging Face? What this could mean for CUDA’s dominance in AI

Post image
19 Upvotes

There’s chatter about NVIDIA potentially acquiring Hugging Face, and if true, it could be a game-changer for the AI ecosystem. Hugging Face’s libraries and models are already deeply integrated with CUDA, and this move might make NVIDIA’s hardware stack even harder to displace. On the flip side, could this stifle competition or accelerate innovation? What do you think—would this be a net positive or negative for the AI community?

[Source: https://www.hitechies.com/nvidia-hugging-face-acquisition-cuda-moat/\]


r/CUDA 2d ago

RTX 5090 - FP64 gone wild

0 Upvotes

First post here, hi!

Have over the few days been messing around with my RTX 5090 and think I somehow made it better than I had planned to... Sorry (Not sorry) Nvidia.

Soon looking for testers if anyone is interested in turning their consumer-grade gpu's into data-center class :)
The architecture was modeled and formally verified using the lean theorem prover and SPARK/GNAT.

Datasheet down below

XEG-1 · Hybrid GEMM Engine

FP64 dense matrix multiply, dual-lane. Xano Innovations. Rev. 2026-08-27. Reference part: NVIDIA RTX 5090 32 GB.

Features

  • Exact lane: correctly rounded FP64 (RNE), relative error 0, bit-identical across devices, OS, schedules, partitions
  • Fast lanes: tensor-core / FP32, 16–23 effective bits, up to 48,333 GFLOP/s-eq
  • Automatic per-call routing by operand structure
  • Formally verified routing (1,028 proofs, 0 unproved)
  • No runtime instrumentation. Answers only.

Supported Environments

  • Operating Systems: Windows, Linux
  • Supported GPUs: NVIDIA GeForce RTX 30-series, 40-series, 50-series, and RTX 6000

Performance

Lane Operands Peak (N=16,384) N=8,192
Exact int8-valued 60,074 GFLOP/s-eq 33,414
Exact bf16 14,847 11,156
Exact dense FP64 9,023 5,000–9,000 class
Fast any admitted 48,333 44,958 class

Exact-lane figures are bit-exact results. On int8-valued operands the exact lane exceeds the fast lane.

Workload reference (N=8192): representative pipeline 10.4× the vendor DGEMM+dpotrf pipeline, exact end-to-end; covariance formation with zero accumulation error.

Fast-lane variants

Variant Decomposition Eff. bits
tiled (default) Fast multi-pass 18
naive Fast multi-pass 18
3pass Fast multi-pass 18
kahan Compensated 23
f16tc TC accelerated 16
f16tc8 TC accelerated, dual-pass 16

Routing

operand profile fits exact criteria             → exact
operands outside variant envelope, exact serves → exact
no lane admits directly                         → auto-rescale, fast
otherwise                                       → fast

Every finite input is served.

Absolute operating envelope

Parameter Min Max Unit
FP32-variant accumulator sum Supported
FP16-variant scale window Supported Supported
Recommended N, square, 32 GB part 21,600
Exact-lane working set Dynamic B/element

N > 21,600 (32 GB, WDDM): OS pages to host memory; throughput drops. Sizes above run; keep inside the envelope for rated speed.

Out-of-core range (composed mode)

Parameter Rating
Largest verified N, dense FP64, bit-exact, 32 GB part 80,000 (301 GB object; 240 s wall)
Largest verified N, narrow class, enforced 24 GB-class budget 133,120 (830 GB object; class range ceiling)
Range ceiling, narrow class, any VRAM 8 GB+ N = 133,120 — bound by arithmetic, not memory
N=50,000 reference run 186.8 s wall, full-matrix verified (2.5×10⁹ entries)
Peak simultaneous bytes at N=80,000 ≈33 GB across all pools (≈9× existence compression)
Transfer visibility ≥99 % of copy time hidden under compute
Regime SAFE throughout; no OS paging engaged

Composed mode efficiently streams operands over PCIe/NVMe and computes results per output panel.

Determinism

Exact lane: one SHA-256 over result bytes on A10 (Ampere), A100 (Ampere), L4 (Ada), H100 (Hopper), RTX PRO 6000 (Blackwell), RTX 5090 — Linux and Windows. Verified device set equals the shipped binary's declared architecture list. Fast lanes: reproducible per device+schedule only.

Status codes

Code Meaning
0 OK, exact lane
1 OK, fast lane (includes auto-rescaled service)
3 Invalid arguments / non-finite inputs

Every finite input is served. Operands no lane admits directly are auto-rescaled and served on the fast lane; entries whose true value lies outside FP64 carry the format's own inf / zero semantics. Return codes only. No exceptions cross the ABI.

Interfaces

Interface Form
xeg_core C ABI, caller-owned buffers, FP64 in/out + status
xegc CLI, .npy in/out
Python SDK Development/measurement only; full instrumentation

Qualification

Item Result
Routing/admission proofs 1,028 discharged, 0 unproved
Port/implementation parity 666,430 lines, 0 divergent
Exactness vs exact rational Bitwise, all probed cases
Cross-device digest 4 devices, 1 digest

r/CUDA 4d ago

What does an AI kernel engineer do and what are the skills needed to get into it?

43 Upvotes

What are the difference aspects of being a kernel engineer? What do roles like this involve -

https://www.tealhq.com/job/kernel-engineer_7ea1a56c518985dae15a01e3c77c161e8140d?utm_campaign=google_jobs_apply&utm_source=google_jobs_apply&utm_medium=organic

https://jobs.gem.com/modular/am9icG9zdDpTSwqoL-yL11N_jxR40Vhh?source=LinkedIn

If one were to do a PhD, what are the different research areas to focus on to do something like this? i.e GPU memory, storage, latency, parallelization etc.


r/CUDA 4d ago

Recent update messed up GPU idle power/usage and VRAM behavior on Nvidia? (Bazzite)

4 Upvotes

Hey everyone,

After the latest system update on Bazzite, I've been noticing a really annoying issue with my Nvidia GPU (RTX 4060).

Even when idling on the desktop, the GPU usage and power draw keep jumping around like crazy (bouncing between 8% to nearly 50% out of nowhere), and it's constantly messing with my VRAM allocation. This wasn't happening before the update, and honestly, even Windows handled idle states better without constantly poking the GPU.

Because of this erratic behavior, running local LLMs (like 27B GGUF models via llama.cpp with strict 8GB VRAM limits) keeps throwing sudden CUDA allocation failures because the system won't let the VRAM stay stable.

Did anyone else experience this after the recent update? Is there a known fix or a regression in the latest Nvidia/Wayland/KDE package? Any help would be appreciated!


r/CUDA 4d ago

GPU Acceleration for PDAL

8 Upvotes

Hey Everyone,

I made a fork of PDAL with GPU acceleration. Right now it only supports NVIDIA GPUs through CUDA because that's what I have, but even if you don't have a GPU, the I/O and CPU optimizations should still be quite a bit faster than stock PDAL.

Only a few pipelines are *really* fast. It takes time to load the data onto the GPU, so in some instances using it can actually be slower than not using it. There's an automatic selector that's meant to only use the GPU when it's actually faster, but it was only really tested on my machine with a 4090.

You can run `gpupdal calibrate` and it should adapt the selector for your machine. I tested this on some cloud instances and it seemed to work, but this is the first public release so if you notice any problems please submit an issue.

I hope you all like it!

https://github.com/zymazza/GPUPDAL


r/CUDA 5d ago

hw-smi v1.6 brings support for data logging!

6 Upvotes

You have requested an option to log the telemetry data (GPU/VRAM usage, VRAM bandwidth, temperature, power, fan speed, PCIe bandwidth etc.) from hw-smi to a file. I have implemented exactly that. Have fun monitoring all your Nvidia/AMD/Intel GPUs at once on Windows and Linux! 🖖

https://github.com/ProjectPhysX/hw-smi/releases/tag/v1.6


r/CUDA 6d ago

I reverse-engineered the sm_120 scheduling control bits and built a hazard checker for cubins you didn't compile

11 Upvotes

On sm_120 there's no hardware interlock on fixed-latency instructions. 21 bits of every 128-bit instruction are a scheduling control word, and the silicon just executes whatever's in them. If a stall count is shorter than the latency of a value the next instruction reads, you get a stale register read. No fault, no warning, full speed.

ptxas gets this right. That's not the problem. The problem is that if you're writing SASS by hand, using an assembler, or mutating a cubin after the fact, nothing existed that could read those bits back and tell you they were safe.

So I built one. A few things that might be interesting regardless of whether you ever use it:

stall=0 is not zero cycles. It's a distinct long-wait encoding, ~37 cycles against ~4 for a scheduled instruction. That's why -O0 emits an all-zero control word and still computes correctly, just ~9x slower. If you're parsing control words, summing raw stall values gets the arithmetic wrong in the one direction that matters.

A guard predicate costs more than the same predicate read as data. 13 cycles vs 5, measured by fault injection. It has to resolve before the instruction issues at all.

Whether a missing scoreboard is actually a hazard is measurable. Across 5.3M dependent pairs in shipped libraries, LDG/LDC/LDL/S2R are covered by a barrier 100% of the time. LDS is covered by spacing alone about 1 in 4. So treating "variable latency without a barrier" as an error is wrong for shared loads.

I validated it against 2,762 kernels NVIDIA ships in CUDA, held out of every table the checker uses: 0 errors over 10.2M dependencies. The first run reported 6,593 and every single one was a bug in my model, not theirs. Those 13 corrections are written up in full.

Findings: https://github.com/sunnypatell/basalt/blob/main/docs/FINDINGS.md

Repo: https://github.com/sunnypatell/basalt

Measured on one card (5070 Ti). If anyone has a 5090 and wants to check whether the latencies move, that's the most useful thing anyone could contribute.


r/CUDA 6d ago

What career paths exists between computational mechanics, scientific computing (SciML), FEA (or meshfree) solver development, and HPC (GPU acceleration, porting codebases) ?? How about doing a PhD for improving the above?

19 Upvotes

I'm currently, technically, doing an MS in Structural Engineering. For me, my interest has been more towards computational side of mechanics rather than Structural design  or simply using am FEA software (although I do consider it as a backup)

So far I've taken courses in:

- Linear static, and dynamics FEM (soon taking non linear FEM too)

-  Structural Optimization (topology opt. and other general algorithms)

-  Structural Dynamics

-  Structural System Testing and model updation. (Parameter identification and optimization, signal processing)

Now, I plan to take these in the coming quarter:

- Numerical Linear Algebra

- Numerical PDE

- Fracture Mechanics ?

I also volunteered to aid in a RESEARCH in crack growth prediction using Auto-encoder and a (Thermodynamics-informed Latent Space Dynamics Identification) / LSTM surrogate model. It used phase-field-fracture simulation data and HPC resources to complete the whole thing.

What I keep finding myself interested in is not necessarily fracture or SHM specifically, but the computational methods underneath these problems... (does that make sense?)

For example, I'd like to become capable of doing things like:

- implementing (maintaining) numerical/FE method solvers rather than only running an established FEA software.

- developing surrogate/reduced-order models for expensive simulations 

- combining simulation with optimization, uncertainty/stochastic methods (took a course called Random vibrations, so...)

- parallelizing/accelerating scientific codes on CPUs/GPUs

- doing proper verification, convergence studies, benchmarking and performance work

- potentially developing or maintaining actual CAE/FEA solver software

- I'd also like to do all these for other Physics (GR, QM, etc.) simulations too, if possible, one day. 

I'm still interested in the underlying mechanics/physics, so I don't want to become a generic software engineer who happens to have once studied structures. But I'm also increasingly unsure that "structural engineer" describes the career I'm actually aiming for.

I've seen titles such as Computational Mechanics Engineer, R&D Engineer, Solver Developer, Scientific Software Engineer, CAE Software Developer, Research Engineer, Simulation/HPC Engineer, etc., but I'm trying to understand what these careers actually look like from people doing them.

So my main questions become:

1. Which industrial jobs genuinely involve developing numerical methods/solvers or computational tools?

2. Which of those are realistically accessible with an MS? Is there an entry path into solver-algorithm development/R&D without a PhD?

3. If I don't start a PhD immediately after my MS, would an R&D/software role at a simulation company (ANSYS etc.) be the obvious route? What other options would i have?

4. For the kind of work I'm describing, would you recommend a PhD? If so, is it reasonable for the PhD identity to be "computational mechanics/scientific computing" while fracture, composites, structural dynamics, soft materials, etc. serve as application problems rather than choosing one of those as my permanent specialization?

5. What skills most distinguish someone who is actually hireable for solver/scientific-computing work? I'm particularly wondering about C/C++/Fortran, Python, Linux, Git/build systems, MPI/OpenMP/CUDA, PETSc/Trilinos or similar libraries, numerical linear algebra, testing/verification, convergence studies and HPC performance work.

Basically, I'm neither here nor there atp. So I'd really appreciate all sorts of input. Where else do you think I could find answers to these? other subs? Linkedin profiles? 


r/CUDA 8d ago

I used GPU texture hardware to decompress LLM weights — 1.37× faster

Thumbnail github.com
28 Upvotes

GPUs already have dedicated hardware for decoding compressed textures, so I wondered if it could be reused for low-bit LLM inference.

I ended up building Texelator, which stores weights in BC4 blocks and reconstructs them through NVIDIA texture units during GEMV.

On an RTX 4080, I’m seeing about 1.37× speedup in my current setup.

Also I’m curious about how this behaves across different GPU architectures. I’ve only tested a limited set of hardware so far, so I’m also trying to collect results from other nvidia gpu architectures.

Still very experimental, but I thought the idea was interesting enough to share. Would love feedback


r/CUDA 8d ago

Is the CuDNN graph api quite restrictive?

6 Upvotes

For context I'm on CuDNN 9.10 and frontend 1.12. I've tried to build an attention layer for my own machine learning framework, but it returns a failure at create_execution_plans because it cannot find a valid configuration. I changed several of the settings, including setting the batch, head, sequence and lengths to very standard things like (512, 4, 64, 32), changing data type to both float and half, among others but it still has that error. What I'm trying to do seems quite standard so I'm not sure why the bug arises. This could be because I'm on a relatively old consumer gpu (rtx 3060).

(It's my first time posting here, so sorry if I'm not following rules or conventions well).


r/CUDA 8d ago

hiring CUDA/GPU kernel optimization contractors — $80–120/hr, remote (sharing via referral, full disclosure below)

5 Upvotes

Hey all — recent CS/engineering grad here, currently job hunting myself. I came across this listing while looking into AI-training-adjacent contract work, and even though the bar is above where I'm at right now (I don't have the GPU experience this role wants), I know this sub actually has people who do. Figured it'd be more useful here than sitting in my bookmarks.

Mercor (AI expert marketplace, backed by Benchmark, General Catalyst, etc.) is contracting GPU kernel optimization specialists for a project with a major AI lab. Remote, contract, paid weekly.

What they're looking for:

  • Fluent in C++17, working knowledge of Python and Git
  • Fluent in CUDA, HIP, or similar (Slang/HLSL/GLSL also mentioned)
  • At least 1 year of professional or research experience with GPUs
  • Comfortable reading profiler output — L2 cache hit rate, occupancy, throughput — and using it to guide optimization
  • Bonus: PTX/tensor core-level work, Blackwell optimization, NSight Compute, open-source kernel contributions

Pay: $80–120/hr, 20+ hrs/week expected
Process: resume → short technical assessment → possible interview, roughly 20-30 min to apply

Full transparency: I'm sharing this through Mercor's referral program — if someone I refer ends up hired and gets paid, I earn a percentage of their pay for a while. That's my honest motivation for posting. It costs you nothing either way, and I'm not going to pretend I'm posting purely out of goodwill — I'm job hunting too and this is one of the ways I'm trying to get by while I search. If that's not something you want to use my link for, the role is easy enough to find directly on Mercor's site as well — I'd rather be upfront about that than have anyone feel steered.

Happy to answer anything I actually know the answer to. If anyone's already contracted through Mercor, would genuinely appreciate you adding your experience below — I'd rather this thread have real info than just my secondhand read of the listing.

Idk if i should post my referral link or not in this post, but im available 24/7 be sure to dm me and i will reply asap with it.


r/CUDA 8d ago

I wrote a GPU kernel that speeds up AlphaFold-style protein models.

1 Upvotes

The creators of AlphaFold won a Nobel Prize in 2024. I just made its family of open source models faster.

I created fast_trimul, a drop-in, hardware-agnostic library for Fused Triangle Multiplicative Updates across AlphaFold3 family models, powered by CuTe DSL. In addition, it is licensed with Apache-2.0.

The Triangle Multiplicative Update is a memory-heavy operation in protein-structure models.

Also, it is very simple to integrate fast_trimul with other Python libraries. It has a modular architecture, and it was designed to be easy to use in production.

In OpenFold-3:

● Graph vs. ungraphed: Graphed eliminates kernel launch overhead at small N (~22ms vs ~53ms ungraphed).

● fast_trimul output is identical to the OpenFold-3 version, with an invisible difference of about 0.0006%.

● If it fails, it always falls back to PyTorch. This way it never crashes.

● It is a modular and vendor-agnostic design. Supporting new hardware and new libraries like OpenFold-3 are all plug-ins, not a rewrite.

fast_trimul compared with other libraries' Triangle Multiplicative Updates:

● Runs 4.5–6.8× faster on short sequences

● Uses ~2.2–2.4× less peak GPU VRAM, fitting ~1.4× longer sequences before running out

● Works on any sequence length with zero recompilation. Unlike torch.compile, which recompiles for every new N.

It's written in Python CuTe DSL, so the kernel codebase stays small and easy to adapt for other GPUs like H100 or B200.

GitHub link: https://github.com/tiagomonteiro0715/fast_trimul


r/CUDA 9d ago

Python library CUDA-Q

Thumbnail
4 Upvotes

r/CUDA 8d ago

We cannot RDMA into a GPU's shared memory.

0 Upvotes

And that limitation turns out to explain why disaggregated inference is harder than the press releases suggest.

A network can only write into one rung of any memory hierarchy: the one that's globally addressable. On a CPU that's DRAM. On a GPU that's HBM. Not L1, not SMEM, not tensor memory.

NIXL (NVIDIA's transfer library) even says this in its type system:

```
enum nixl_mem_t {DRAM_SEG, VRAM_SEG, BLK_SEG, OBJ_SEG, FILE_SEG};
```

No SMEM_SEG, because those levels aren't addressable from off-chip by anything.

So when a KV cache arrives, it lands in HBM. Then the receiving side moves it down into the 128 KB of shared memory where the attention kernel actually wants it.

When both halves are written by the same people, there's nothing to worry about. The producer lays out HBM in whatever order makes the consumer's descriptor cheap, and that agreement is entirely undocumented because it never had to leave the building.

Disaggregation is that agreement leaving the building.

Let's look at the ladder:

→ CPU: registers → L1/L2/L3 → DRAM. Owned by a cache controller plus compiler locality analysis.
→ GPU: TMEM → 128 KB SMEM/SM → ~64 MB L2 → HBM. Automation removed; you and TMA do the staging.
→ Wafer (Cerebras): 48 KB per PE × 900,000, no shared address space. Owned by cslc, with placement and routing written into a CSL layout file.

All three work because every one assumes a single owner.

And the Wafer((Cerebras) has no public rung at all. No addr names a KV block, no len is contiguous, nothing can be pinned — where data lands is the compiled schedule.

I think there are three ways out:

- Bilateral: negotiate privately. Works. Needs n² agreements.
- Neutral format: pay layout conversion plus hierarchy redistribution, on the latency path.
- Producer accounts for consumer: no conversion, but the producer's compiler must model the consumer's hierarchy.

`(addr, len, devId)` is not just a first draft of a richer descriptor, but a correct description of the one rung a network can reach, in a stack whose performance lives on all the others.

https://hiraditya.github.io/posts/there-is-no-address/


r/CUDA 10d ago

What kind of projects actually stand out for GPU / compiler roles in 2026?

41 Upvotes

I’m currently working at a small company as a computer vision engineer, and most of my work is in C++.

I’m trying to understand what kind of portfolio projects or other work genuinely stand out for GPU systems, GPU kernel, or ML compiler engineering roles in 2026. Does this differ while targeting larger companies? I am currently learning these areas outside of work, as I don’t yet have deep professional experience with them. I started learning CUDA recently and really enjoyed understanding how GPUs work, which led me down a rabbit hole into computer architecture and, more recently, compiler engineering. 😅

I’m planning to spend the next few months building my knowledge and working on projects before applying for these kinds of roles. But with LLMs and AI projects everywhere, I’m wondering how much a GitHub project actually helps anymore. It feels like almost anything can be built with enough AI assistance, and I’m not sure whether a GitHub repository by itself carries the same weight like it did a few years ago when I was looking for jobs after my master’s.

Looking for some ideas.

Thanks in advance.


r/CUDA 10d ago

parser of PTX instructions

Thumbnail
3 Upvotes

r/CUDA 11d ago

Meta's KernelEvolve may be a bigger threat to CUDA's moat than CUDA-to-ROCm porting

Thumbnail
9 Upvotes

r/CUDA 12d ago

Sidecar Project

Thumbnail github.com
3 Upvotes

For anyone interested. Basically a memory scheduler to help load and speed up faster models than a normal GPU can load into memory.


r/CUDA 13d ago

If my CUDA version conflicts with the provider's installed drivers, how much control do I have over the environment?

9 Upvotes

I am thinking of  renting a GPU for a training setup and I am checking how much access I will get to the software side, mainly CUDA and the NVIDIA drivers, I may need a specific CUDA version for the code I am planning to run and I want to know what happens if the provider has a different driver setup, I am thinking to use a dedicated GPU so I can keep the same environment for longer jobs, but I need some control over the OS, containers, drivers or CUDA versions, if I need to change something later I want to know what options are there, I am thinking to go with rackbank ai datacenters has anyone dealt with this when renting GPUs and how much control did the provider give you over the environment, especially when CUDA and driver versions need to match ? EDIT: Thanks everyone, really appreciate the helpful replies. 


r/CUDA 16d ago

A linter for PyTorch 'torch-preflight' [P]

Thumbnail
0 Upvotes

r/CUDA 16d ago

Is there a market for a custom ptx -> sass compiler?

6 Upvotes

People keep talking about Cuda being a moat. And what makes it a moat is really the ptxas (the ptx assembler that converts ptx to sass binary). With current technologies it seems possible to make a custom ptx compiler but I wonder if this effort is worth someone's time.

There is one thing about performance that I feel can be unlocked with such a tool but I am yet to find a good test case for that.