Very large amounts of gaming gpus vs AI gpus

TheMightyCat@ani.social · 8 months ago

Very large amounts of gaming gpus vs AI gpus

brucethemoose@lemmy.world · edit-2 8 months ago

Be specific!

What models size (or model) are you looking to host?
At what context length?
What kind of speed (token/s) do you need?
Is it just for you, or many people? How many? In other words should the serving be parallel?

In other words, it depends, but the sweetpsot option for a self hosted rig, OP, is probably:

One 5090 or A6000 ADA GPU. Or maybe 2x 3090s/4090s, underclocked.
A cost-effective EPYC CPU/Mobo
At least 256 GB DDR5

Now run ik_llama.cpp, and you can serve Deepseek 671B faster than you can read without burning your house down with H200s: https://github.com/ikawrakow/ik_llama.cpp

It will also do for dots.llm, kimi, pretty much any of the mega MoEs de joure.

But there’s all sorts of niches. In a nutshell, don’t think “How much do I need for AI?” But “What is my target use case, what model is good for that, and what’s the best runtime for it?” Then build your rig around that.

TheMightyCat@ani.social · 8 months ago

My target model is Qwen/Qwen3-235B-A22B-FP8. Ideally its maxium context lenght of 131K but i’m willing to compromise. I find it hard to give an concrete t/s awnser, let’s put it around 50. At max load probably around 8 concurrent users, but these situations will be rare enough that oprimizing for single user is probably more worth it.

My current setup is already: Xeon w7-3465X 128gb DDR5 2x 4090

It gets nice enough peformance loading 32B models completely in vram, but i am skeptical that a simillar system can run a 671B at higher speeds then a snails space, i currently run vLLM because it has higher peformance with tensor parrelism then lama.cpp but i shall check out ik_lama.cpp.

GPU	VRAM	Price (€)	Bandwidth (TB/s)	TFLOP16	€/GB	€/TB/s	€/TFLOP16
NVIDIA H200 NVL	141GB	36284	4.89	1671	257	7423	21
NVIDIA RTX PRO 6000 Blackwell	96GB	8450	1.79	126.0	88	4720	67
NVIDIA RTX 5090	32GB	2299	1.79	104.8	71	1284	22
AMD RADEON 9070XT	16GB	665	0.6446	97.32	41	1031	7
AMD RADEON 9070	16GB	619	0.6446	72.25	38	960	8.5
AMD RADEON 9060XT	16GB	382	0.3223	51.28	23	1186	7.45