← All Posts

AMD acquires AI chip startup Taalas to boost inference performance by etching models into silicon

The Hot Take: That sounds alot like an ASIC...

In AMD’s latest bid to upset Nvidia's dominance in AI hardware, the House of Zen has acquired AI chip company Taalas, which bakes model weights directly into silicon in a process that promises to boost inference performance by an order of magnitude or more. The deal, announced at market close on Thursday, appears to be framed in much the same context as Nvidia’s $20 billion licensing deal with Groq last December: make high-performance “premium” inference services prized for AI agents, like code assistants, faster and cheaper to run. AMD didn’t disclose the terms of the deal, but from what we understand, this is an actual acquisition rather than an acquihire. Founded in 2023 and based in Toronto, Taalas’ approach to inference is radically different from conventional GPUs or the dataflow architectures that underpin Groq LPUs or Cerebras' waferscale accelerators. A model-specific integrated circuit The startup’s chips don’t rely on HBM to store the model weights but rather etch them directly into the silicon. In a sense, Taalas’ chips are really model-specific integrated circuits or MSICs. Perhaps more importantly, Taalas’ tech isn’t just conceptual. In February, the startup revealed its first test chip fabbed on TSMC’s 6nm process tech, which it called the HC1. Initial benchmarks saw the chip serve Meta’s Llama 3.1 8B at a blistering 16,960 tokens a second — when announced last February, that was 48x faster than Nvidia's GPUs and 8.5x faster than Cerebras' accelerators. While Llama 3.1 is ancient by today’s standards, having made its debut all the way back in mid 2024, the reticle-sized chip was really intended to prove the concept. Taalas has been incredibly secretive about how its chips actually work, but we know its processors are comprised of two main regions: the mask-ROM recall fabric where model weights are etched, and the SRAM recall fabric where KV caches and fine-tuning adapters are stored. For its second-gen HC2 chip due out this summer, Taalas aims to boost parameter count to 20 billion parameters. That might not sound like much, but just like with GPUs for larger models, weights are simply distributed across multiple accelerators using pipeline parallelism. At 20 billion parameters per chip, you’d need just 50 accelerators to support a trillion-parameter model, and AMD just so happens to have a rack-scale compute platform and in-house system design team that can comfortably accommodate that. That’s quite a bit more space and power efficient than Nvidia’s recently unveiled LPX systems, which would need a few dozen GPUs and at least 2,000 Groq LPUs to serve the same model. From what we understand, AMD intends to pair its Instinct-based Helios racks with chips based on Taalas’ tech, which implies a disaggregated architecture where compute-heavy prompt processing is done on GPUs while token generation is offloaded to Taalas-based accelerators. It’s also possible that AMD could adopt a sort of tick-tock cadence in which customers initially deploy and validate models on Instinct accelerators and, once they’re satisfied with them, transition to Taalas accelerators. We can only speculate at this point, but here’s what AMD’s SVP of AI, Vamsi Boppana, had to say about it in a canned statement: “AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload." You better really love that model While the tech is blazing fast, if you hadn’t already figured it out, it comes with a pretty substantial downside. Once the chips are deployed you’re stuck with that model. Any change bigger than something like a LoRA adapter is going to require a re-spin of the chips, which is not only expensive but time-consuming. Nearly four years into the AI boom, new models are rolling out on a nearly monthly basis. In order to benefit from Taalas’ tech, AMD’s customers are going to have to be really sure about their choice of models, which will be easier for some than others. However, if the startup is to be believed, the situation isn’t quite as bad as it sounds. While new models will require a re-spin, it doesn’t require starting over from scratch. Instead, just two layers of metal need to be changed, which is a lot cheaper and less time-consuming. With that said, we strongly suspect this tech will largely be deployed by AI model devs, their infrastructure providers, and a handful of inference providers. In an interview with our sibling site The Next Platform in February, the company suggested that etching a model's weights into silicon is 100x less expensive than training a frontier model. AMD is certainly in a position to negotiate those deals. OpenAI, Anthropic, and Meta are all major Instinct customers. Given the close working relationship between the model houses and the chip designer, it wouldn't be surprising to see a GPT or Claude deployed on a combination of Taalas and instinct accelerators. The tech also has implications for model development. One of the ways developers have cut down on hallucinations is by trading time for accuracy. The technique, called test-time scaling, is quite simple in practice, and involves allowing a model to “think” for longer before responding. One drawback of test-time scaling is that it consumes substantially more tokens, which makes it expensive, and means users have to wait longer for the chatbot, code assistant, or agent to respond. If AMD’s Taalas buy can drive down the cost per token and boost output speeds by 10x or 20x, model devs may opt to extend the reasoning time even further. In any case, we may not have to wait long to see just how Taalas fits into AMD’s broader vision. Subject to regulatory approval, the deal is expected to close in the fourth quarter. ®

Read the full article

Intel grows as strongly as it has in more than 15 years: AI boom drives Xeon demand massively

The Hot Take: Good to hear, and hope it's true.

Intel is back with unusually strong quarterly figures. In the second quarter of 2026, the company generated revenue of around 16.1 billion US dollars, 25 percent more than in the same period last year. According to Intel, this was the strongest revenue growth in more than 15 years. The data-center processor business in particular benefited […]

Read the full article

AMD Is Reportedly Pricing Its Helios Rack 40% Above NVIDIA’s Second-Gen Rubin, Confident That Customers Will Pay The Premium

The Hot Take: Will have to take this with a bit rock of salt until we see the benches and real world performance. Just a price grab or actual performance?

AMD is apparently no longer satisfied with being perceived as a perennial underdog, and is reportedly choosing to make a bold statement by pricing its Helios rackscale solution at a hefty premium to NVIDIA's Rubin. What's more, AMD has reportedly managed to rope in Microsoft as the product's first confirmed client, demonstrating the utility that Helios offers even at its supposedly elevated price point. Futurum believes AMD will price Helios at $5 million to $5.5 million per rack, versus an expected retail price of $3.5 million to $4 million for the second-gen NVIDIA Rubin rack As we explained in a […]Read full article at https://wccftech.com/amd-is-reportedly-pricing-its-helios-rack-40-above-nvidias-second-gen-rubin-confident-that-customers-will-pay-the-premium/

Read the full article

Intel Posts Initial GCC Compiler Patches For AI Compute Extensions "ACE"

The Hot Take: I really wonder if this will deflate the Ai bubble even a little bit.

The x86 Ecosystem Advisory Group led by Intel and AMD recently firmed up the AI Compute Extensions (ACE) specification for optimizing x86 for AI computation tasks around matrix multiplication and the like for machine learning workloads. The cross-vendor ACE extension is ultimately a successor to Intel's Advanced Matrix Extensions (AMX). Posted to the GCC mailing list today by Intel engineers are the initial patches in preparing the compiler support for ACE...

Read the full article

Qualcomm Claims Single-Core Leadership for Its First Server CPU, the Dragonfly C1000, Delivering 250+ Cores & 5 GHz By 2028

The Hot Take: Interesting...

Qualcomm has introduced its first-ever CPU designed for Data Centers, the Dragonfly C1000, which leverages the Oryon architecture. Qualcomm Enters The Agentic AI CPU Race With Dragonfly C1000 Chip, Oryon-Based With Over 5 GHz Clocks, Over 250 Cores, & Aims To Achieve Single-Core Leadership One of the biggest announcements by Qualcomm today was its first release of a CPU for the data center segment, called the Dragonfly C1000. This is a chip purpose-built for Agentic AI & General-Purpose workloads, delivering best-in-class power efficiency and TCO. As per Qualcomm, the Dragonfly C1000 is based on a custom-designed Oryon core architecture that […]Read full article at https://wccftech.com/qualcomm-single-core-leadership-first-server-cpu-dragonfly-c1000-250-cores-5-ghz-2028/

Read the full article

Tensordyne's 3nm Napier AI Chip Promises 13x Higher Token Throughput Than Blackwell & Blazes Past Rubin With 1000 Tokens/s In Multi-Trillion Parameter Models

The Hot Take: I really hope something comes soon to alleviate all this nonsense AI is causing.

US-based AI company, Tensordyne, has announced the successful tape-out of its Napier chip, which it claims to demolish NVIDIA's Blackwell & Rubin chips with leading token throughput and efficiency. Tensordyne’s new Napier AI Chip arrives with one clear mission: to make NVIDIA’s Blackwell and Rubin chips look considerably less impressive The Napier chip will be the core component of the Tensordyne Napier TDN system, which is designed in collaboration with Broadcom and HPE Juniper Networks. The Napier platform has one goal: to unify AI through novel logarithmic AI math, a tightly integrated memory architecture, and a high-performance scale-up interconnect that […]Read full article at https://wccftech.com/tensordyne-3nm-napier-ai-chip-13x-higher-token-throughput-blackwell-blazes-past-rubin/

Read the full article

AWS Graviton5 Debuts with 192 Arm Cores and PCIe 6.0

The Hot Take: ARM seems to be breaking out from everywhere. Fujitsu, Nvidia, AWS and ARM. Qualcomm seems to be playing catch up in the server market from the looks of it.

AWS has provided a first look at its next-generation Graviton5 processor, a custom server CPU developed by Annapurna Labs for deployment across the company's cloud computing platform and AI inference infrastructure.

Read the full article

Microsoft is killing the Copilot+ PC advantage, brings Windows 11’s local AI to RTX 30+ PCs with 6GB vRAM

The Hot Take: Now we know why M$ is trying to squeeze out every ounce of performance in Windows 11.....

Microsoft says you’ll be able to run Windows 11’s local Language Model APIs on non-Copilot+ PCs as long as you meet the new hardware requirement: an RTX 30+ GPU with 6GB of VRAM. It’s a major change, as it means Copilot+ PCs’ advantages are getting “thin,” and I wouldn’t be surprised if Microsoft drops the NPU requirement entirely in the future. Copilot+ PCs officially debuted on June 18, 2024, and they’ve been driving sales for PC makers. However, it’s not because of the “Copilot” or “NPU” factor. It’s largely because newer PCs are now sold as “Copilot+ PCs,” so even a regular laptop purchase gets counted as proof that AI PCs are taking off. For a PC to meet the “Copilot+ PC” requirement, it would need to have 16GB of RAM, an SSD, and at least a 40 TOPS NPU. For those unaware, an NPU (Neural Processing Unit) is a chip designed to run AI models, specializing in efficiency rather than raw power. On the other hand, a GPU is a heavy-duty processor designed for massive parallel tasks. What is a “Copilot+ PC?” Microsoft sold you Copilot+ PCs as the only way to run local AI, but that was never…

Read the full article

Chinese military has been acquiring Nvidia chips, even post-Washington export controls, research claims — multiple institutions linked to the PLA asked for Nvidia AI chips, according to publicly available documents

The Hot Take: Tell me something I didn't know already. Why else would the GPU market go crazy prices wise?

A business-intelligence researcher said that the Chinese military has been actively acquiring Nvidia AI chips, even after the U.S. put export controls on them. Public documents show that some institutions ask for these chips either through the specifications they demand or by directly asking for Nvidia chips by name.

Read the full article