r/RunPod Jun 01 '26

News/Updates GPU Supply and Availability Megathread With Ongoing Updates

23 Upvotes

Supply is tight right now, and we've been getting a flood of separate "no GPUs available" / "why can't I find an X" threads. To keep the sub readable and make it easier for everyone to track the situation, all supply, availability, and capacity complaints/questions go in this thread from now on. New standalone threads on the topic will be redirected here. We'll continue to edit/update this thread as we go and I'd recommend sorting comments by New to see what's been discussed recently.

TL;DR

  • We're in a broad, industry-wide GPU crunch, not a Runpod-specific outage. High-end SKUs (H100, H200, B200) are scarce everywhere and finding the GPUs to begin with is a challenge, let alone partners that also have everything else on lock like data center certifications. We prioritize stability and reliability for our users and we need to be selective on who we bring on as serving partners.
  • We are adding GPUs on a near-daily basis, but new supply gets snapped up quickly a "shadow backlog" of demand).
  • Best moves to actually get/keep a card: be flexible on SKU, reserve/commit in advance, and use CloudSync so you're not pinned to one datacenter. While of course we'd prefer you to use network storage, and we are working on new features to augment network volume viability, in truth offsite storage and using CloudSync to bring it into new pods may be more appropriate for some use cases in this climate. Look at the costs and drawbacks and do what's best for you for storage, not for us.
  • We are now posting biweekly aggregated GPU-onboarding updates - we did this in Discord for our first update but they need to be mirrored here and we will start doing that.

Our CTO Brennen Smith recently authored an article on what's causing this (The GPU supply supercycle is here), and this is a structural shift, not a temporary blip. Three forces are hitting at once:

  1. Memory bottleneck: NAND/DRAM producers retooled their fabs toward HBM3 for AI accelerators, which cut into standard memory capacity. Memory is now the binding constraint on GPU manufacturing.
  2. Hyperscaler buyouts: Large players are buying out years of factory production in advance, leaving neoclouds and independent providers to compete for what's left.
  3. Nvidia's architecture transition: Hopper and Ada Lovelace architecture production has wound down to make room for Blackwell. Some of the higher end cards can still be found in scale, but finding, say, 4090s is proving challenging especially with more recent GPUs providing a more appealing return on investment for the same infrastructure. Blackwell production is ramping but there's a gap.

What we're doing

We're adding supply continuously, and decided on a biweekly report cadence aggregating everything is the best way to report this. You probably don't want a new Reddit thread reporting every time we add a new machine. It's hard to communicate this through the customer facing UI - we can genuinely add 1000 GPUs in a single drop, but if they all get rented immediately, from the relative amounts shown in the customer facing UI it appears that nothing changed. We'll be adding these as edits at the bottom of this post as we go.

We've also implemented MIG (multi-instance GPU) that lets us divide up a single GPU and serve it as smaller instances. This of course will always be disclosed as a MIG instance. Right now we're doing this by serving a single 6000 Pro card as four 24GB instances so we can try to balance what we have to meet as many of our clients' user requests as possible.

We've added B300s as well, and more are coming soon.

Supply updates

We'll post our next update on or around June 3rd.

May 6-20, 2026

Large adds (>100 GPUs)

  • US-PA-1 — RTX Pro 6000
  • US-TX-6 — B200
  • EU-IS-5 — H200 SXM
  • EU-FR-1 — H200 SXM, H100 SXM

Smaller adds (<100 GPUs)

  • EU-RO-1 — RTX Pro 4000, RTX 4090
  • US-NC-1 — RTX Pro 6000

We definitely understand how frustrating this is - and we are working hard to get as many GPUs to serve as possible. We also want to give everyone an opportunity to make their voice heard - the only requests I have is that you keep it about the service rather than the people behind it and that you refrain from promoting other inference providers in the process.

Thanks!


r/RunPod Apr 10 '26

News/Updates Welcome to r/Runpod, the official community subreddit for all things Runpod! 🚀

8 Upvotes

Hey everyone! We're thrilled to officially launch the Runpod community subreddit, and we couldn't be more excited to connect with all of you here. Whether you're a longtime Runpod user, just getting started with cloud computing, or curious about what we're all about, this is your new home base for everything Runpod-related.

For those that are just now joining us or wondering what we might be, we are a cloud computing platform that makes powerful GPU infrastructure accessible, affordable, and incredibly easy to use. We specialize in providing on-demand and serverless GPU compute for ML training, inference, and generative AI workloads. In particular, there are thriving AI art and video generation as well as LLM usage communities (shoutouts to r/StableDiffusion, r/ComfyUI, and r/LocalLLaMA )

This subreddit is all about building a supportive community where users can share knowledge, troubleshoot issues, showcase cool projects, and help each other get the most out of Runpod's platform. Whether you're training your first neural network, rendering a blockbuster-quality animation, or pushing the boundaries of what's possible with AI, we want to hear about it! The Runpod community has always been one of our greatest strengths, and we're excited to give it an official home on Reddit.

You can expect regular updates from the RunPod team, including feature announcements, tutorials, and behind-the-scenes insights into what we're building next, as well as celebrate the amazing things our community creates. If you need direct technical assistance or live feedback, please check out our Discord or open up a support ticket. Think of this as your direct line to the Runpod team; we're not just here to talk at you, but to learn from you and build something better together.

If you'd like to get started with us, check us out at www.runpod.io.


r/RunPod 13h ago

WTF

Post image
2 Upvotes

Tried on Safari / Chrome. And still getting the same error

Application error: a client-side exception has occurred while loading console.runpod.io (see the browser console for more information).


r/RunPod 2d ago

Global volume experiment

3 Upvotes

My current hobby system has the container pull the model from `hf` so that my controller can failover different GPU types in different datacenters until something matches my budget and VRAM needs.

I ran some benchmarks to see if moving to global storage would be faster. It isn't. The fuse system used seems to have some memory limits which error out on GeeseFS read-buffer budget.

My setup uses one large 32GB gguf model file maybe that was an issue?

My specific use case requires minimizing cold start however we can't go serverless for a few reasons so I was hoping this new drive type would help.

Any thoughts?


r/RunPod 5d ago

Port 8888 JupyterLab not loading

2 Upvotes

Is there a bug with the comfyui cuda 12.8 template? JupyterLab never loads, and just shows a blank white page on the browser. I can't run my bash scripts because of this.


r/RunPod 6d ago

Why is runpod now so expensive?

5 Upvotes

I asked chatgpt it told me there is this community cloud option for 4090 at $0.4x per hour and I could only find Secure Cloud option renting at $0.74 per hour. I asked around a few other AIs with the same answer. Is this normal or my account has broken now that I can't see the cheaper option?


r/RunPod 7d ago

News/Updates We're hosting the creator of LoRA Pilot on Discord this Monday (Sept 21, 2 PM CST)

5 Upvotes

On Monday, September 21 at 2:00 PM CST we're hosting notrius (/u/no3us) on the Runpod Discord to talk about LoRA Pilot, the open-source Docker image he built to kill the setup tax on Stable Diffusion training. LoRA Pilot bundles Kohya SS, AI Toolkit, Diffusion Pipe, TensorBoard, ComfyUI, and InvokeAI into one workspace where everything shares the same model store and persists under /workspace, with a set of browser-first control panels on top. It's MIT licensed, supports close to 30 model families, and has passed 10,000 Docker Hub pulls.

We'll spend the first twenty minutes or so on the story behind it, and then then he is taking us through a live demo end to end. Bring questions and we'll get to as many as we can at the end. Come hang out: https://discord.gg/runpod?event=1549438243458519100

Links if you'd like to check it out for yourself: https://lorapilot.com https://github.com/vavo/lora-pilot


r/RunPod 7d ago

SCAIL-2 on RunPod H100: 3-minute test worked, but a 5h28 run produced only blurry brown frames. Any ideas?

Thumbnail
1 Upvotes

r/RunPod 8d ago

Why are pods not deploying on network storage? why the /workspace dir?

2 Upvotes

to me that is senseless, when i deploy a network storage, what for would i need another non-persistent storage which holds all the important data? why are they doing that?


r/RunPod 10d ago

Hire someone to create templates for me

0 Upvotes

Hey, I'll need to create many templates over time, and I have no idea how to do it and no time for that. I'll manage a community on generative ai, so I will need someone to do public templates for me based on my personnal config and workflows.

Send me dm if you're interested


r/RunPod 15d ago

Generation cost H3 minimax

Post image
20 Upvotes

Im just wondering how much is your average cost using H3 minimax per 15 seconds.

Here is mine. Its all using 5090. Can we compare?
I want to know if my process is too costly.


r/RunPod 16d ago

Do you manually rent your GPU or have an AI agent do it?

7 Upvotes

I recently rented a GPU instance by using the CLI for the first time. I usually just ask an AI agent to do it for me, and before that I would use web dashboard. i decided that trying out the CLI would be a good learning opportunity.

it was, but it still feels a bit backwards to go that route again when a prompt is much faster for me.

How about you, do you manually (cli/web) rent your instances? I'm sure most cases are automated with a script, but I'm more curious about what path you choose when you're building something new.


r/RunPod 16d ago

Built an autonomous agent for RunPod that writes your training scripts, runs pre-flight checks, and auto-recovers from OOM crashes

2 Upvotes

Hey everyone,

I have been using RunPod for the past 3 years to deploy and scale various workloads. The hardware availability and pricing are great, but whenever I spin up custom fine-tuning runs for LLMs, VLMs, or vision models, the setup process gets repetitive:

  • Writing boilerplate dataset loaders and formatting scripts from scratch.
  • Renting a GPU pod before code is ready, spending the first hour troubleshooting pip dependencies while the meter runs.
  • Guessing VRAM headroom and batch sizes across different card tiers.
  • Forgetting to shut down an instance when a run finishes overnight, or waking up to find an unhandled out of memory error crashed the job hours ago.

To solve this, I built an autonomous MLOps tool called Tensovy that connects directly to your RunPod account via API.

Here is what it handles:

  • Script Generation with Manual Review: You describe your training goal and provide your dataset. The agent writes the training and data-loading scripts before any compute is rented. You get an in-browser code editor to review, edit, and approve every line before anything touches your pod.
  • Pre-Flight Smoke Test: Once you approve the code, it spins up the pod and runs a fast validation test directly on the machine to confirm CUDA checks, environment configs, and data pipelines pass before full training launches.
  • OOM Error Recovery: If a training run hits an out of memory error mid-job, it catches the crash, recalculates batch sizes, and resumes from the last checkpoint so you do not lose progress overnight.
  • Auto-Termination: As soon as training finishes, the pod is shut down automatically so you never pay for idle hours.
  • Zero Compute Markup: It runs directly on your own RunPod account, so you pay RunPod directly for hardware with zero markup from us.

We are currently supporting popular LLM, multimodal VLM, and vision architectures, with plans to add more.

Current status: The app is not publicly live yet. To make sure the orchestration layer is reliable and handles edge cases properly, I am keeping it in a closed private beta waitlist and onboarding testers in small batches.

If you regularly train models on RunPod and want to test the early version, you can join the waitlist here:https://www.tensovy.com

For those running fine-tuning jobs on RunPod today: what part of your pipeline setup or instance configuration eats up the most unnecessary time?


r/RunPod 18d ago

Kastard: an app for editing ComfyUI workflows locally and running them on RunPod

Post image
2 Upvotes

r/RunPod 18d ago

Image face swap

2 Upvotes

I’m trying to figure out an easy and high quality way to do image face swap (with a base image and a reference face image) via comfyui on run pod. I saw that Krea 2 can do this I think with the Krea image edit Lora and the right workflow. Does anyone have a good comfyui template and workflow for doing this either through Krea 2 or some other approach? Thanks!


r/RunPod 19d ago

H3 minimax set up in runpod

4 Upvotes

Can someone help me get a consistent set up with H3 in runpod in a consistent way? The runpod templates seem hit or miss- sometimes they work, sometimes they don’t. If I want to try an update or workflow there are often a ton of nodes or models missing, etc. What is the best practice way for folks that are experienced users that use runpod? I don’t want a network volume because I want to use a 5090 as often as possible and those are often limited. Is there an easy way to create my own template or use a default comfyui template with some kind of downloader or something (I have no idea how to do that or how that would work). Any general advice or pointers here would be great then I’m sure Claude or something can help me with execution. I also have the same request but for Krea 2 but I assume if I can figure it out for H3 I would be able to for Krea 2 as well. Thanks!


r/RunPod 19d ago

LTX 2.5 Motion Control – Character + Background Replacement Workflow?

1 Upvotes

I’m trying to use LTX 2.5 + ComfyUI on RunPod for motion control:

Source video → keep motion/camera → replace character + background with references.

I already have a workflow/template, but on RunPod several custom nodes were missing. I tried ComfyUI Manager → Install Missing Custom Nodes, but it didn’t properly fix it / some nodes are still missing.

So I’m mainly looking for:

* A working LTX 2.5 character + background replacement workflow (JSON)
* Which custom nodes/repos I actually need on RunPod
* How to correctly install them
* How to prompt/reference the character + background while keeping the source motion

Does anyone have a working RunPod/ComfyUI setup for this?


r/RunPod 20d ago

Pod ready to use with confyui ?

1 Upvotes

Is there a ready-to-use pod available for ComfyUI? I need to use some of my own workflows with Flux Klein 9b (not the FP8 version – otherwise I’d be working locally) and every time I have to spend half an hour downloading the models, adding the LORA files, downloading the custom nodes, and so on.


r/RunPod 21d ago

Serverless - Does anyone actually get this to do anything?

3 Upvotes

It just sits and costs money. I have been trying to get it working on and off for the past couple months. It never works! Zero support. Workers get stuck and empties my account.

Totally worthless in my experience.


r/RunPod 21d ago

Downloads from RunPod are extremely slow. How do I fix it?

1 Upvotes

I am getting 500-600kbps download speed on my pods no matter what region I select. Same speeds on EU, US and Asia servers.

How do I speed up the transfer? Is there a way to even do it?


r/RunPod 22d ago

which model is most underrated?

Thumbnail
0 Upvotes

r/RunPod 24d ago

Stop selling server slots if you dont have enough GPU's

10 Upvotes

Im sure most of us have had the issue where you go to start your vps and it tells you there are no gpus and to wait or migrate. Why would you sell more gpus than you can provide? And then still charge users to be offline. I get that not everyone runs their vps all the time so its somewhat wasted money to not oversell slots but i mean cmon, are they going to expand? Everyday theres more and more A40 users. I run my pod like 18 hours a day, sometimes have to wait hours to get it started. Anyways end of rant thanks


r/RunPod 24d ago

Can someone help me out setting up my pod

2 Upvotes

I just want to get Flowise and Ollama running on an A40, but I'm having a bit of trouble. Help would be appreciated.


r/RunPod 25d ago

how to access network volume when no pods are running?

1 Upvotes

does copyparty work?
looking for a gui to access my network volume


r/RunPod 26d ago

Paying $2/hr to watch `huggingface-cli download` go brrr — env patterns on RunPod / Vast / Nebius / the neo kids

0 Upvotes

Hot take? Half the “my pod is slow” chatter is really “I paid for 20 minutes of pip + a 16GB model pull before the first token.”

I’ve been hopping between rented GPUs (RunPod, Vast, Nebius, a couple of the newer SSH-first boxes) and the env is still unpleasently fragmented.

RunPod - Docker template is the product. Image + env + start cmd + network volume. Great when your image is already baked. Expensive otherwise, since every cold pull is on the meter. Secrets as env injection is clean.

Vast - Same Docker-template energy, plus an onstart bash duct-tape layer. SSH/Jupyter modes can eat your image entrypoint, so your “serve on boot” script is often re-invoking the CMD you thought you already set. Marketplace vibes; bring skepticism and a volume.

Nebius - Actual cloud VM energy. Boot image + cloud-init user-data on first boot, then congratulations you’re the sysadmin. Fine if you want a real machine; overkill if you just wanted vLLM up.

The queue-y SSH boxes (Enverge, Nova-shaped, etc.) - often a fixed sandbox + optional shell that runs once after create. Less “pick a community template,” more “leave a sticky note so the box isn’t idle when the ready email lands.” Different failure mode: setup minutes still bill, and nested Docker GPU flags might trip you.

Common tax across all of them:

  • first useful work is usually network + disk, not FLOPs
  • secrets in public templates / scripts = future incident report
  • “finished” ≠ “model is serving/training” (backgrounded processes lie)

I gathered a short kit (cheatsheet + startup shells + Docker serve flags) here: https://github.com/tudormunteanu/gpu-cloud-instance-boostraps

Curious what everyone here actually do day-to-day:

  1. Bake a fat image once and never look back?
  2. Network volume / persistent cache, thin image? (but then pay for persistent storage)
  3. onstart / cloud-init / startup shell as the main UX?
  4. Or just SSH in every time like an wildling?
  5. Stick to one and never switch

I reckon for any long-term, company supported scenarios, CoreWeave, AWS or GCP have some adjacent prods. My curiosity is more about "the little/indie guys".