DEV Community

Cover image for AMD's Move on Weight Storage: The Taalas Bet
Peremptory
Peremptory

Posted on • Originally published at peremptory.ai

AMD's Move on Weight Storage: The Taalas Bet

AMD's bet on weight storage as the real bottleneck in inference just got concrete. On Wednesday, the company announced it's acquiring Taalas, a Toronto startup that embeds AI model weights directly into custom silicon rather than storing them in high-bandwidth memory (HBM).

The move makes intuitive sense if you squint at the math. Weights are heavy. Moving them from storage into the compute units is expensive. If you bake them into the silicon itself, close to where you need them, you cut the most expensive part of the inference pipeline: the memory fetch.

Taalas has a working prototype. Their HC1 chip, built on TSMC's 6nm process, claims to serve Llama 3.1 8B at roughly 17,000 tokens per second. That's the performance claim. The comparison is messier: Taalas says this is 48x Nvidia's GPUs and 8.5x Cerebras. Both comparisons depend on setup, batch size, and whether we're talking about cost or raw throughput. The company is also shipping an HC2 with 20B parameters coming this summer.

What's interesting here is not whether these numbers hold up, they probably don't in the way Taalas markets them, but that AMD is buying into the architecture at all. This is a $34 billion company betting that the next generation of inference doesn't look like the last one. It's saying that VRAM bandwidth is no longer the constraint you throw more money at. You redesign the chip.

The deal closes in Q4 2026, subject to regulatory approval. That gives AMD time to figure out how to integrate Taalas's approach into its broader AI silicon roadmap. It also gives the market time to test whether baking weights into silicon actually solves the problem or just moves the bottleneck elsewhere. You can't update weights easily if they're fused into the wafer. You're committing to a model, a quantization level, a batch size. Flexibility trades for speed.

There's also a market question here about what actually matters to the companies buying inference hardware. Right now, the arms race is about absolute throughput and the cost per inference. Taalas's angle is latency per token and power efficiency for specific, fixed workloads. That's a different game than what Nvidia is playing. AMD is betting someone will care enough to pay for it.

The timing is interesting too. OpenAI's latest models are still training. Anthropic is shipping smaller, better versions of Claude. The industry is not optimizing for weight-serving performance yet. But AMD is building the hardware now, assuming it will be. That's either prescient or expensive.

Top comments (0)