• 0 Posts
  • 206 Comments
Joined 3 years ago
Cake day: July 5th, 2023

  • It hasn’t felt t like there’s been much significant performance increases or development in RAM in the last… decade?

    In memory? There’s been a ton of improvement, even if most of the coolest stuff isn’t making it into DIMMs that are installed in user laptops/desktops.

    Advanced packaging technology has allowed chip manufacturers to put different silicon dies together with increasingly high performance (high bandwidth, low latency) connections in the same package, including with some three dimensional stacking. That way they can mix and match different silicon dies for greater cost effectiveness, yield, performance, etc.

    This also means that in-package memory is now the standard in certain chips. Apple’s M-series silicon has its memory packaged right into the CPU/GPU package, as a system-in-a-package, so that the connection between the logic and memory is comparatively much higher performance, several times higher bandwidth than desktops or laptops that don’t follow that kind of architecture.

    Similarly, in data centers, the AI boom has caused all the memory manufacturers to switch their production lines to high bandwidth memory, where they vertically stack a bunch of DRAM chips on each other, with ultra-fast, high bandwidth connections, so that they can shove terabytes of memory into these data center servers. These recent generations have been improving speed and bandwidth in ways that make consumer level DDR5 RAM look like child’s play.

    So they’re improving things. Just not in ways that really show up in DIMM sticks.


  • Not strictly, there are usualy hurdles to overcome for home usage of datacentre tech, but it’s possible.

    The hurdles are basically insurmountable with the hardware released after 2024.

    The NVL72 for the Blackwell generation cost about $3 million and takes up a single server rack. The power consumption is about 130 kW, and most configurations require dedicated plumbing for the liquid cooling.

    To put things in perspective, a residential electrical hookup is usually 50A or 100A for a house, with recommendations that anyone who is going to be charging electric cars should have 100A service. 100A at 240V is 24 kW.

    So one server rack uses as much power as the maximum electrical capacity of 5 homes. You’ll never be able to pull that off in an actual residential environment.

    Oh, and the newest 2026 generation, the Rubin NVL72s, use something like 230 kW of electrical power, almost twice as much as the previous 2024 generation.


  • There’s always going to be a robust used market for phones that were purchased outright, to be resold on a different cycle than every 2 years (plenty of rich people changing phones every year, and plenty of people replacing on a 3, 4, 5, or 6 year cycle). You can expect the market to basically settle on a curve where it depreciates along a predictable rate.

    Leases don’t really change that, any more than leases changed the market for used cars, or even certified pre-owned by the same dealers and organized by the same manufacturers who sell new cars.

    There will be times that the predefined lease terms will unexpectedly prove to be either beneficial or detrimental to the consumer. Sometimes external factors will affect the entire used market, like currency issues, or component pricing issues (imagine if RAM prices dramatically swing again for new devices in a way that affects the value of the already-sold devices out in the world), where the predefined lease prices turn into a windfall for someone. Like in 2021 or so when expiring car leases allows the lessee to buy out the car at the end of the lease for much cheaper than the car itself was worth.

    It’s generally going to be a less than ideal financial decision to lease, but it also won’t collapse the used device market and it won’t be that far off the practice of selling your old phone when you buy a new one.


  • Law enforcement can legally trick you into giving up your password, too, and that’s full access right there. Having an unlocked phone but no password isn’t enough to get into certain parts of the core system/security settings, and trying to get into those will prompt a password anyway (and that generally gatekeeps the access to the phone through a physical connector plugged into the port).

    Neither pathway is perfect but I think for real world usage and real world adversaries (not just law enforcement, but also criminal thieves/scammers/hackers, and governmental adversaries that aren’t bound by legal limits, like foreign intelligence agencies), it’s better to have biometrics so that you are physically punching in your PIN/password much less frequently. Especially on modern systems that get spooked easily and require a password anyway when the phone has been idle too long or when the wrong face looks at it too many times.


  • The other underappreciated threat model is shoulder surfing, especially in an age of ubiquitous high resolution cameras. Punching in a numerical PIN within view of a camera potentially leaks that secret, and some high resolution cameras can even pick up letters and symbols from the on screen keyboards.

    Being compelled to give biometrics doesn’t do enough for an adversary (including government adversaries) to do everything with a phone, the way having the password or PIN does, and I would argue that governments would be better at tricking people into inadvertently giving up their PINs and passwords than they’d be at compelling biometrics within the time window that they still work (before the phones lockout biometrics as a valid unlocking method), or being able to do stuff to exploit extraction tools past the lock screen.

    So the threat model needs to be understood for what it is.




  • the growth itself is hella juiced because the GPUs are only relevant for about 3 years till the new ones are out and make more AI for less power. And they depreciate them over 7 years. More than twice as long as they can or should use the GPUs for.

    We don’t actually know this for sure, yet. I had expected the A100 generation (released in 2020) to no longer be profitable to run by now, but the backlog in new data centers being turned on and the high demand from Anthropic and OpenAI still leaves those chips useful for inference. You can rent those 2020 chips out today at some price above what they cost to continue running (300W, so electricity prices of USD $0.20 per kWh would translate into about 6 cents per hour. Prevailing spot prices appear to be about $2/hour right now.

    But just because I was wrong on 2020 chips, originally sold for about $15,000 in a low interest rate environment, doesn’t mean that I’m wrong about 2024 chips, the B100s that use 1000W and were sold for $35,000, requiring a ton more specialized cooling, power, and network infrastructure. Or the 2026 R100s that use 2000W, and whose prices I can’t seem to find published anywhere, but were set after the memory companies basically locked in their record breaking prices for their HBM. That’s an unsustainable path and at some point, data centers start struggling to find users willing to pay the bare minimum necessary to continue turning a profit on GPU usage.

    I doubt the 2024 chips stay in service to 2031. And I’m really, really skeptical that the 2026 chips stay in service to 2033, especially after NVIDIA switches to yearly release cycles next year.


  • But it isn’t encoding knowledge, it’s encoding word correlations.

    I’m saying that humans do this a lot, too. Qualitatively, it’s different, in that this particular batch of frontier LLMs will get things wrong in ways that most human brains wouldn’t, but as a category of error it’s not unique to LLMs.

    I know a ton of facts that I learned only through reading, and have no actual firsthand knowledge/experience or ability to test it: Jupiter is larger than Saturn, the atmosphere during the Carboniferous period was high in oxygen, cigarettes cause cancer, Thomas Jefferson owned slaves, the capital of Norway is Oslo. At best, I can cross reference other sources and see that things are consistent with each other. Is my belief in those facts “knowledge,” or is it merely recognizing from my training data that those particular words can validly be presented in that order?

    If you ask average people on the street whether FAT32 is a good filesystem for a 64GB removable drive, most of them won’t know, but there are a handful of bullshitters who might confidently parrot back things they can Google but not understand. That’s part of the human condition, too.

    I’m by no means an AI booster/enthusiast. I suspect LLMs/transformers are actually a dead end, and expect the upcoming crash to be economically and financially devastating to the tech and financial sectors. But I also have a pretty dim view of human intelligence, too, and see way too many parallels in LLMs as bullshit artists to humans as bullshit artists, too.


  • It modifies the prompt, aka the input, not the output. It is smuggling 3 bits of secret user/session data in a wrapper that doesn’t look like it contains that data. As the article explains:

    So the marker becomes part of the system context sent to the model.

    This is a normal timestamp on a prompt:

    Today's date is 2026-07-11.

    But if your system timezone is a Chinese mainland timezone, it looks like:

    Today's date is 2026/07/11.

    Then, if your base URL includes a keyword like “deepseek,” it silently replaces the apostrophe from a ' to a ʼ:

    Todayʼs date is 2026-07-11.

    Or if the base URL has one of the domains on the list, like any .cn domain, it replaces the apostrophe with another apostrophe character:

    Today’s date is 2026-07-11.

    And if it has both a URL and a keyword on the watchlist, the prompt context includes:

    Todayʹs date is 2026-07-11

    That’s 3 bits of information: does this system have a mainland Chinese time zone, does the base URL contain a known keyword (associated with Chinese AI competitors) or a known domain (associated with mainland China or its major tech companies). And it sneaks it on by without making it obvious.

    That’s steganography.



  • It’s that they are trying to use statistics to encode entire thought processes into hidden variables from conversation snippets. They want to use statistics to go from many individual interactions to a large model, and then use that model to predict individual interactions again.

    Has it been shown that the human brain doesn’t model the world in a similar way, though? A huge portion of human knowledge is both stored and transmitted in the form of language. Lots of human knowledge also follows the garbage in, garbage out theory, where you can have entire areas of knowledge that aren’t actually true but might be internally consistent, at least within certain scopes: conspiracy theories, belief in the supernatural, entire academic disciplines built on a religion or theology that not everyone believes, etc. Or even world building in fiction, the words on a page can be enough to convey ideas such that it “tricks” human brains into filling in the gaps so that they internally see a rich, fleshed out world that is entirely fictional and where specific details might not find strong direct support in the underlying text.

    it has no concept of correctness

    But statistical weight on what is more or less likely to be correct still makes a difference to objective quality of the outputs. If the model weights are trained on the reality that high quality university texts describe something and reflect some sort of underlying model of what is described using language, then can’t the model itself learn as much as a human could from those words on a page?

    All models are wrong, but some can be useful. And different models have different quality in different domains. So although I don’t believe LLMs will overtake the hump of getting ahead of human knowledge, I also don’t believe that any given LLM can be evaluated on quality, and that Facebook’s LLMs are significantly behind other LLMs we see.

    And that maybe a huge part of it is its internal process of preparing the model to evaluate the quality of its inputs, such that the output it produces can also score high on quality.





  • Basically they’d need about as much in radiator fin surface area as they would have in solar panel area. The ISS has 8 solar array wings, 35m x 12m, that can produce about 30 kW each, or 240 kW total, in sunlight (which is only half the time). The ISS has a complex cooling system, but relies on 4 radiators about 3.1 m x 13.6 m to reject up to 14 kW of heat each (56 kW total) for cooling the solar arrays themselves. The main cooling system uses 6 radiators, each 23.3 m x 3.4 m, to reject 70 kW of heat (from this report it sounds like each radiator may be capable of rejecting more than 1/6 of the heat but that the system as a whole needs to be kept under 70 kW of heat rejection).

    So that seems like about 650 square meters of radiators can provide about 120 kW of heat rejection.

    Today, a 72-GPU Blackwell server is 130 kW in a single server rack. The next generation rolling out now has 72 Rubin GPUs in a 230 kW server, in a single rack. And that’s not even a “data center.” That’s just a single (albeit very powerful) server. How many can you string together, with networking equipment beaming data connections back down to the ground, before the ratio of solar panels and radiators to the actual ship size becomes unworkable?

    That said, it’s technically possible, especially if you can radiate the heat at higher temperatures than the ISS does, as the Stefan-Boltzmann law shows that the hotter the radiator, the more heat it can reject. Just completely infeasible from an engineering and economical standpoint, for any data center that hopes to be relevant in an age of 100+ MW data centers.





  • I just pulled up the ChatGPT terms of use

    Who’s talking about ChatGPT or OpenAI?

    I just pulled up the Anthropic commercial API terms, since that’s the situation covered by the original article (big corporation using Anthropic’s paid API):

    Use Restrictions. Customer may not and must not attempt to (a) access the Services to build a competing product or service, including to train competing AI models except as expressly approved by Anthropic; (b) reverse engineer or duplicate the Services; or © support any third party’s attempt at any of the conduct restricted in this sentence.

    Ok, so it’s a contract that purports to prohibit pretty much this kind of model weight extraction, and I’m saying that Anthropic probably considers the model weights to be trade secrets.

    Are you under the impression that trade secret protection only happens when the contract says the words “trade secret”?

    Or, analogously, consider customer lists. Having a contract that says “don’t copy my customer lists even if I sometimes disclose a single customer at a time when we partner together on projects” is probably enough to adequately maintain trade secret protection over those customer lists, even if individual customers are sometimes disclosed under a contract.

    I’m just stating what I believe the law is, not what it should be, or even claiming that what the law is today is good. I’m just saying everyone should be aware that the law is quite protective of big corporations and their proprietary secrets. I still think this qualifies as a trade secret that they’ve protected with their own contracts.