• Not an AI guy, but I do like using niche hardware wrong to get results cheap. Can anyone tell me what this would be like for gaming or general computing? My 1660 super was a budget pick when I got it back in '18.

  • 7 hours

    Yeah but the 0 point doing this unless you want to run AI models for some reason. These GPUs can’t do video game graphics so this isn’t a solution to the GPU shortage.

    This is a bit like me writing an article about NASCAR, now I can turn left whenever I want. But I haven’t magically acquired a functional vehicle for a fraction of its value. I’ve purchased a second hand specialist product that is usually useless outside of that environment.

  • Here I was expecting a graphics demo to blow our collective minds but instead I got a story about a local LLM for cheap. It is <current year>. I should have known better.

    Were I the author / tech cobbler here, I’d be concerned that too much time with an LLM, local or otherwise, might erode or dull my apparently fairly sharp reasoning and tech skills. (Clarification: Not my sharpness, theirs. I’m a potato.)

    Other thoughts: For a minute I thought this whole thing was a tribute to, or a troll in the manner of, that one Redditor that always spun their stories around to being about their dad beating them with jumper cables.

    Also, my old PC developed an issue like the warm reboot problem, except with the network interface. I couldn’t just restart, I had to power off and back on. I never did bother to find out whether it was early signs of hardware failure or whether it was an old hardware / newer kernel mismatch.

      • The change of the meaning of the G in GPU from “graphics” to “general” is even less well documented and used than the “V” of DVD changing from “video” to “versatile”.

        Indeed it only occurred to me what it must have changed to and to go looking to confirm after seeing your comment.

        And frankly they ought to have changed the name to something like “MPPU” if they wanted it to stick (massively parallel).

  • The article talks a lot trash about AMD and ROCm but vulkan works fine too. In fact from a datacenter GPU standpoint there is an AMD option called the V620 available on US EBay that I was able to haggle to $350, with 32GB VRAM, 512GB/s bandwidth, and runs the same Qwen-3.6-27b at about 20t/s. I would argue that’s even more cost effective.
    It requires a few of the same fan shenanigans this guy did but there is no need to pull specific past software versions to make it usable in Linux

    • Honestly even for the prices of around 500$ that I’m seeing it for it looks like a pretty good value to get 32gb of vram. I see it says 300w on AMD’s product page for it does it have any way of power limiting the card to get more efficiency/less heat?

      • I spent a lot of time researching and testing different methods for that, the only thing that worked was LACT in Linux. Using that I was able to undervolt 100mV and GPU power usage dropped about 10%. On my B450 ITX board with a Ryzen 2400GE CPU the entire system pulls 30w idle from the wall, and about 300w inferencing with VRAM filled. (330w before LACT).
        My fan solution ended up being to buy the 80mm 3d-printed shroud off ebay, the fan that came with it was super loud so I switched to an arctic p8 Max, and control it with the motherboard targeting a t-sensor header with the probe attached to the backplate.

  • 23 hours

    Sure it’s got a lot of VRAM, but the 4080 has five times the compute power.

      • Hmm… I have a 4080 from 2022. I wonder if I can get this thing too, for $200, and use them both? I hate only having 16gb VRAM.

        Would be nice if I could use them both at the same time. Like, make the 4080 recognize the other device as additional VRAM

        • 14 hours

          Yes, that’s basically what the article is about. They run the LLM across both GPUs.

          But that’s a feature of llama.cpp. SLI doesn’t really exist any more, and NVlink requires a specific setup, which the 4080 is not part of (the 3090 was the last consumer one, apparently). So you couldn’t pool the VRAM.

        • Not sure how well that without work even if you could bridge them properly to share their vram, as the latency from the other gpu will be pretty high compared to the local vram. Frame times won’t be that great is my prediction. It works ok for LLMs because they aren’t a realtime compute task like gaming graphics is. If an LLM takes an extra 0.12s to compute its result, you don’t notice, but if a gpu misses a frame deadline by 0.12s, that’s a stutter that represents less than 10 fps.

          I believe that’s why earlier attempts at dual gpu systems mostly fizzled out (plus cost concerns). You can get some pure acceleration if you can fit the entire working memory onto both GPUs’ vram (so your total vram is effectively min( GPUA_VRAM, GPUB_VRAM ) rather than GPUA_VRAM + GPUB_VRAM, though that also requires all pixels to be independent of anything calculated on the other GPU, other than maybe post processing effects that could be handled on whichever GPU is handling the display.

          That’s not to say that you can’t get extra performance out of multi-gpu setups that don’t just mirror their RAM, but it’s more complicated than “sum of the capabilities of each GPU”.

    • 20 hours

      That’s fine I just need to display pictures of your mom (they are very large) (/s)

  • 1 day

    Seems like an awful lot of trouble to save $100 not buying a 5060 Ti that also has 16GB.

    • 11 minutes

      The 5060 Ti does not support Nvlink though.

      • 16 hours

        Good point, that should theoretically be double the decode speed, though not sure how the prefill would differ since it’s mainly compute-bound.

    • I had no idea you could get a 16gb card this cheap! TIL.

      It’s pretty low performance, though… My card from 8 years ago nearly matches the passmark rating.

  • 1 day

    eBay has some rad Chinese mezzanine boards for these guys too. Nvlink works and everything lol 3566 file-QvfRnBhkoKQBxmqtmrGBLV

    • 14 hours

      Cool mezzanine boards, but can we talk about your dope af custom jig for offset mounting arbitrary boards?

      • 40 minutes

        I appreciate the kind words! Until recently, my day job was CAD monkey. I wanted to consolidate the hardware I was cobbling together, found a cheap 8 GPU mining rig and some extra 2020 extrusions to play with.

        The seller for the mezzanine said it followed the mATX mounting hole pattern, (it does but is ~75% in the width and doesn’t use all the points). I have a tendency to overthink designs and kind of stalled for a bit before finally taking the plunge and whipping up these struts. I used some brass heat set inserts to accept the standoffs and everything pretty much went right together lol

        3523

        • 17 minutes

          Damn, even cooler than I expected 🎸 Thank you for the extra details and picture!

      • 22 hours

        I am using PTM sheets and they idle at decent temps, though I did ziptie some high CFM fans behind them lol 3581

    • 19 hours

      They mostly don’t, but this is also not how the datacenters cool them.

      A datacenter will either have an open water loop, or an all in one taking heat to a more advantagous place for a radiator to be, or at the very least better managed airflow with bigger fans and more specific air baffles.

      This thing has no such luxury and has a small area and unknown broader thermal context, so screaming it is to make up for the limitations of the scenario.

    • If it doesn’t sound like a jet plane taking off, is it really a server at all?

    • No they really dont. Big ass fans running 24/7 to help the small fans running 24/7. It all blends into an easily ignored drone though just dont try to have a conversation in there.

  • The main thing that itches me with the V100 is the fact that given that pascal is about to be EOL, a 2017 card is probably soon next

  • This isn’t the smart way, though.

    What the homelabbers do (at least before the RAM crisis) is buy Xeon/TR/EPYC boards on the cheap, and then run gaming GPUs for hybrid inference.

    This is what I do. I run MiMo 2.5 at 8-10t/s on a 7800X3D/RTX 3090/128GB CPU RAM, more with Dflash. That’s a 300B model: it’s not even in the same class as Qwen 27B, which is what the dev in OP’s article is trying to run.

    And this is small-time: setups with 4-8 memory channels can run stuff like Kimi or Deepseek Pro, even faster. Or they can run smaller LLMs with quantization types that are very fast on CPUs, and get crazy speeds.

    …And besides, Qwen 27B can run fine on a 4080, with the right framework. It will fit in 16GB as an exl3.


    Not that this isn’t a cool hardware hacking project.

    …But it’s kind of the wrong approach. It’s about 2 years out of date, as MoEs are king in LLM land now. RAM is horrendously expensive, yes, but so are most used V100s, or used 3090s.

  • 1 day

    I never expected used DC GPUs would make it to eBay.

    • 19 hours

      V100s are ancient in a market obsessed with the very very latest.

      The datacenters are unlikely to bother directly with eBay, but they have asset recovery companies that will take the stuff off their hands and seek buyers, including over eBay.

      • 2 hours

        Hmm, there are some AMD Instict listings on ebay, after all.

    • 1 day

      Lots of used DC gear makes it to eBay. This is how homelabbers survive.

      • 1 day

        I know that and I use that. But recently market for switches and servers dried up. I thought they were shredding DC GPUs too. Most of them die after 3-5 years anyway.

    • 7 hours

      I thought parts felt like AI and parts felt like a British human writing it.

      AI(?):

      But here is the thing: this is a Volta GPU with 16GB of HBM2 memory, 5120 CUDA cores, and I picked it up for about £150 on eBay. The compute is still real. The VRAM is still real. And the memory bandwidth is where it gets genuinely surprising.

      Human (?):

      So I shoved some jumper wires into the connector and jammed the other ends into a spare fan header (turn your volume up)

      I suspect he’s used some AI polishing tool to turn his draft into a publishable post.

    • 15 hours

      It’s so clearly written with the aid of an LLM that I’m embarrassed for all the people downvoting this comment. Just because the basic outline of the information presented and the sequence of events comes from the real experience of a real person, it does not mean that the real person wrote the article themselves. Seriously, how many real people write like a bad 80s detective movie script? It’s literally a parody trope at this point.

      • Honestly, I just don’t care.

        I am a monkey, I live on a rock, in the middle of an infinite empty space, that none of the smartest monkeys on this rock can figure out.

        I am part of a species that has spent its entire history killing itself in the most horrifying and ghastly ways for the most benign reasons that the greatest war approaches us and everybody is so accepting.

        The planet I live on is slowly becoming inhabitable to point where the thought of kids is pointless. And banks have used to many money glitches that currency across the globe is going into hyperinflation.

        I really really cannot illustrate how much I do not give a fuck about the cadence of written words in that small space of relief I get from the inevitable destruction looming in a completely absurd and pointless existence.

    • 1 day

      Is nobody exasperatedly tired of comments calling anything remotely related to AI ‘slop??’

      • 15 hours

        It has nothing to do with the subject of the article. Are you blind? It’s the exact same bland, AI regurgitated writing style as every other AI written article.

        • 11 hours

          Truly, it does not read like it was written using an LLM to me. Maybe you’re expecting it to, so that’s what you see?

      • Yes.

        Lemmy is just by default super AI-hating. I wonder what they felt when Linus Torvalds said that LLM are actually useful tools.

        • I don’t really care what Linus Torvalds has to say, unless it’s about the Linux kernel, and I don’t think his opinions matter outside of that. Being a great programmer doesn’t mean he’s also an intelligent, socially aware person.

          He did an interview with Linus Tech Tips a few months ago, and when asked about artists being “upset about the large scale theft of work” coming from AI, he said “that’s reality, deal with it” and even went on to talk about photographers that are out of work because “you can fake pictures so much better now” while, within the same breath, saying programmers wouldn’t lose their jobs because he wants to have his cake and eat it too. (Obviously I’m paraphrasing a bit)

          I do agree AI can be a useful tool, and it’s a position I’ve held since OpenAI was first starting to become recognizable. But anything that’s generated by AI is slop, and I’ll continue to call it such even if nobody else will.

          The movie “Spider-Man: Into the Spider-Verse” used AI tools to place random blemishes and marks (that the artists created) on characters faces across multiple frames. I don’t really feel like finding a source for this one, but I’m sure you could fairly easily if you’re that curious. I think that’s an actual use for AI. Streamlining and speeding up the creation process to give creative people more time to do actual creative work. I don’t think asking an LLM to generate or “create” something for you is anything but slop.

          Slop isn’t about quality, it’s about the morality. The, again, large scale theft of work required to get to the point of being able to generate something of a passible quality level, plus the destruction of the environment in the process. I think it’s pointless to reject AI tools entirely due to how prominent and widespread they are, but “creating” something with AI will always be slop.

          Source for the LTT interview: https://youtu.be/mfv0V1SxbNA

          About 33:10

          • 18 hours

            What morality do you derive from creating competing works and throwing them out for free across from where someone is trying to make a living from similar product? -It’s what GPL is all about, and Linus and RMS both live high off the hog instead of living to their commie ideals -the same way communist leaders do.

            Sure, some developers find ways to monetize their work, but many don’t and shouldn’t have to. The ones that do are also catering to the competition while pretending to be competition (like Firefox). -As such, they end up playing politics instead of being run like a business.

        • 1 day

          We felt like he’s right about it being useful tools but the ethical aspects are hard to ignore. Yes, many ethical aspects.

          • 1 day

            You’re likely conflating the tools with the implementation or the implementers. An open source model on local hardware has next to zero ethical concerns.

            There are MANY ethical concerns about the actions of the big players in AI, their data centers, and to some extent, the original, bootstrapped training data sourcing.

            • 19 hours

              Note a key point of contention is how the training data is used and whether it is effectively discarding copyright. If you invested time making an open source project to do something people appreciate and you get attribution as a result, you may be unhappy that a model trained on your stuff can let a user prompt up an embedded implementation of what your project does without any attribution.

              This pretty much applies to all models. No one limited training data to explicitly public domain stuff.

              • 19 hours

                You might not appreciate it, but if it’s posted online then it’s no different from someone else learning to code from reading the project. It’s not making copies of the code, it’s just strengthening the weights on a neural network. Sure, if the code is so obscure that nothing else is like it then it’s possible to get the model to regurgitate some of it due to having so few relevant sources, but it’s very unlikely to be comprehensive enough that it’s violating any copyright. If a court finds that to be the case, for some fictitious example, then I’m certain they can find an agreeable resolution to the isolated case.

                However, none of that is justification for just writing off the technology entirely. Pandora’s Box has been opened. The genie isn’t going back in the bottle. You can’t close the barn door, all the cows already escaped. What do you think boycotting it will accomplish? What exactly is the goal by figuratively sticking your fingers in your ears and pretending the models don’t exist?

                • 18 hours

                  I have seen this argument before and it doesn’t make sense even in theory.

                  I used to work at a company that did open source work and also proprietary work with third party closed source code. The company didn’t let anyone who had seen proprietary code contribute to open source, because they felt once a person ‘learned’ from a proprietary codebase, then it’s too risky if similar looking code lands in a project.

                  Imagine if someone saw the source code for Excel. Then sometime later they notice that Calc didn’t have a feature that Excel did, and contributed an implementation of the feature. Even if they hadn’t been looking directly at the Excel source code in the moment of implementation, you think Microsoft would be so “understanding” when they see someone that once worked on Excel contributing what could be construed as infringing?

                  The AI companies also seem to acknowledge this, as they have offerings that promise not to use your proprietary code as training fodder. If it is not a risk of infringement, then why would it matter to promise that the proprietary code is kept out of training data? Though it was short lived, why would OpenAI have even made a deal to license Disney material if it’s all fair use anyway? If this sort of stuff is fair game, why do they get so pissy when other companies distill models?

                  Even as the AI company’s have roughly defended this scenario, their defense should be a cause for concern for users. Generally they say that anything they do with things they can read is ‘fair use’, and when exhibits of clearly infringing outputs are given, they respond with the model only did that because the user’s prompt directed it, and thus the responsibility for infringement should be with the AI user, not the engine that produced the infringing output. So the possibility of an unwitting infringement is possible as the AI companies explicitly say it’s the fault of the user even if it happens.

                  But we come to your last point, that essentially at this point, the whole thing is ‘too big to fail’ and thus the practical risk is low. Which is true. It’s just a bit disheartening that these companies are given free reign to interpret intellectual property law whichever way is convenient in the moment.

            • 21 hours

              You’re likely conflating the tools with the implementation or the implementers.

              I don’t think so, no?

              • 21 hours

                Okay. So, imagine I’m an indie game developer. I’ve got no artistic talent, no funds, and am just making a game that I want to play, not one I think will make any money.

                With me so far? I download a free, open source, open weights model on my laptop. I install the open source tools to allow the LLM to read my project files, and I give it a thorough description of the game I’m designing. I work on some aspect in the foreground, while in the background my GPU is utilized to achieve some task I’ve assigned the AI. When it’s done, I review the work, commit the changes, and assign it a new task.

                What are the ethical aspects you can’t ignore in my use case?

                • 18 hours

                  If you don’t intend to sell or distribute that game, the only aspects left that I can think of are:

                  • you’re not helping yourself by using AI, you’re just making yourself dumber, so if your goal is simply to make a game, fine, but if you want to maintain your coding skills, it’s a bad idea.
                  • the extra energy consumption, which happens regardless if you’re using your GPU or someone else’s GPU. Might even be worse if everyone uses their own GPU, because there’s a whole ass PC system built around that GPU that also consumes power. But I dunno if that’s the case, actually. I guess this is moot if you’re off the power grid with your own power source, such as full solar or something.
                  • how has the open source model been trained? On what data?
            • 24 hours

              Well, yes and no. Even open source models had to consume a lot of power for the training, including the use of non-owned data

              • 23 hours

                Many, many, many things “consume a lot of power” which is both common and benign. A whirlpool tub “consumes a lot of power” for example. So does an electric stove. And the power can be from any source, such as solar, hydro, or wind.

                That’s why I specified there’s minor issues that could be argued about the original data sources. However, that bird has flown the coop. There’s no putting that genie back in the bottle, and no way to undo it. You could argue the company’s responsible owe every single person on the planet some form of restitution, which I do, but I don’t think that qualifies as an ethical concern since it’s universal. No one person was harmed more than another, it was all publicly available data - so the public should benefit from the result. Refusing to use it only hampers yourself, at no detriment to those perceived as wrongdoers.

                • 18 hours

                  I rarely see people agonize over the ethics of playing a game on high settings for hours because of the frivolous use of electricity.

        • 19 hours

          I think broadly most people are less bothered by what the GenAI can do, but more bothered by what experiences are inflicted upon them by other people using the tech.

          All well and good if it can so code reviews and code gen at your behest and you can evaluate and get as much as you think you can out of it, but other people are enabled to be amazingly more obnoxious about it.

          “Slop” is mostly because people were inclined to do the slop level quality, but were self limited by the reality their poor ideas were roughly as hard to do as much better ideas. Now AI accelerates people that are not concerned with quality or specifics way more than it can accelerate people that care about specifics and quality.

        • Many people are incapable of nuance.

          I think it has some use cases, but it’s like religion: Don’t cram it down my throat, keep it to yourself, and I don’t mind it if it’s on my terms.

          • 20 hours

            Totally get what you mean and I can fully respect that stance. In recent times, it was more like a “let’s bash all the people who think slightly positive about AI”, so kind of the same indoctrination, just the other way around.

        • I’m conflicted about it too. But there is a useful software development tool in there. Its a beige mind good at regurgitating cookie-cutter work. Its pretty good at scaffold and cookie cutter unit tests.

          I hate how it makes my colleages behave, but I can totally see the legitimate $500M/yr industry worth of software in there, under all the BS.

          There’s a Jetbrains or Atlassian size companies worth of product here.

        • 1 day

          Ah, you got me. I copied the double negative of the original comment. You’re right, the replies here show there are indeed very many people annoyed by the ‘slop’ comments.

            • 21 hours

              🤣

              Oh wait, I know this one: “I know you are, but what am I?”

              Name-calling. The last bastion of the stubbornly illogical. I can picture you as a toddler, stamping your feet in a huff as you try to think up more names to call people who disagree with you.

              Am I meanie poopoo head, too?

              • 15 hours

                Name-calling. The last bastion of the stubbornly illogical.

                Just a factual observation. I mean, if it looks like a duck and quacks like a duck…

                • 11 hours

                  Yes, I lick my GPU’s boots… How did you even know it had boots!?

    • The guy is explaining how to hack around GPU fans which were never meant to be run at lower speed through jumper cables and you complain about slop?

      • 15 hours

        This “guy” isn’t explaining anything. The entire article is written by AI. It’s the same bland writing style everywhere now.

        • 13 hours

          I don’t know whether you read the article. It was quite interesting to me and gave me some good ideas. I not going to replicate what he did, but he has shown me a methodology I could apply to do other things.

        • Who gives a shit? You ever read a legal document? They’ve had the same writing style for 70 years.

    • Actually if you’re interested in running a local llm it’s quite a good article

      • 15 hours

        It has nothing to do with the subject of the article. Are you blind? It’s the exact same bland, AI regurgitated writing style as every other AI written article.

      • 15 hours

        It’s literally AI slop. Ignoring the subject of the article (which is whatever), the article is clearly entirely written by AI. The AI writing style is horrendous. “It’s not blah. It’s blah.”, “it’s genuinely <insert adjective of amazement>”

        • 12 hours

          Its not AI writing. It reads exactly like someone casually talking. Its consice and to the point. There is no “its not x, its y”

          He has been blogging dor dam near 6 years you can go read his older posts the writing is consistent.