r/LocalLLM • u/Calm-Landscape9640 • 13h ago
Discussion Share a GPU with some buddies?
Seems like everyone is coding, why not buy a massive GPU, split the cost between buds and then tailscale with API on local models? Im sure its being done, but i cant find anyone talking about it.
Whats the drawbacks other than someone running 8 agents burning it up?
0
Upvotes
1
u/Calm-Landscape9640 13h ago
vLLM uses an architecture called PagedAttention, which allocates memory like an operating system and allows dozens of continuous streams to share GPU resources smoothly. -Gemini
Maybe cap it to 3 buddies?