Itdaily - AMD Epyc Venice: Zen 6 to make AI agents salivate

AMD Epyc Venice: Zen 6 to make AI agents salivate

AMD Epyc Venice: Zen 6 to make AI agents salivate

With the sixth-generation Epyc 9006 server processors, codenamed Venice, AMD straightforwardly claims to be building the world’s best server CPU for cloud, enterprise, and AI. More important than the superlatives is the underlying vision: AI agents are changing the data center, and that requires a portfolio, not just a single CPU.

“Venice is the best server CPU in the world for cloud, enterprise, and AI. Period.” Ravi Kuppuswamy, senior VP and GM Server Product & Engineering at AMD, leaves no room for doubt during the Advancing AI 2026 pre-briefing in San Francisco.

He goes even further: “It doesn’t matter which competitor, which market segment, or which use case: for every benchmark you throw at Venice, you will see that Venice is better.” Bold words, but AMD has come a long way: from 0.2 percent market share in 2017 to 46 percent revenue share in servers today. According to the company, a Forbes Global 2000 company switches over every week.

From chatbot to AI agents

The common thread throughout the launch is the shift from simple chatbots to AI agents. Where a chatbot sends a single question to a GPU and receives an answer, an agent breaks its work down into a series of steps: gathering context, planning, reasoning on the GPU, calling tools, executing code, verifying results, and often starting over multiple times.

AMD analyzed that workflow in eight sub-steps and found that the majority of them run on the CPU, not the GPU.

This translates into three CPU roles in the modern data center. General-purpose servers run the databases, storage, and web services that agents rely on. Host nodes feed the GPUs in AI racks like AMD Helios. And then there is a new category: sandbox servers, where the agents’ control plane lives and where the disposable code generated on the fly by tools like Claude Code and ChatGPT runs.

According to AMD, the ratio of CPUs to GPUs in AI data centers has since evolved from one-to-four to one-to-one, and continues to rise.

What’s new in Venice?

AMD is building Venice on the new Zen 6 architecture, manufactured on TSMC’s two-nanometer process. The most densely packed Zen 6c variant scales up to 256 cores and 512 threads per socket, with a full gigabyte of L3 cache.

Memory is making a major leap forward: sixteen channels of MRDIMM up to 12,800 MT/s, providing 1.6 TB/s of bandwidth per socket—more than double its predecessor, Turin. PCIe Gen 6 doubles the bandwidth to GPUs and accelerators, and CXL 3.1 opens the door to memory expansion.

Also new are smart power features like Unified Power Provisioning, which dynamically distributes a single power budget between CPU and memory, along with security upgrades featuring post-quantum cryptography in the root of trust, FIPS 140-3 certification, and confidential I/O.

However, a single CPU profile is not enough, so Venice comes in four flavors. The current Epyc generation Turin is also available in multiple variants today.

The sandbox CPU: new silicon or clever positioning?

That sandbox category raised the question for us whether this truly requires separate silicon, or is rather a repositioning of existing cloud-native chips. The answer from Madhu Rangarajan, corporate VP of Compute and Enterprise AI Products, is nuanced.

“Customers would rather not roll out too many different SKUs. What we will see: you have a Venice with 256 cores at 600 watts, and for some racks, customers will simply cap it at, for example, 400 watts. For branchy code, where cores often wait for each other, that works perfectly fine. Those building dedicated sandbox racks will opt for a power-optimized SKU. It depends on the scale of your deployment.”

Interesting detail: it is precisely that ‘branchy’, stuttering Python code that agents generate en masse that makes simultaneous multithreading extra valuable. When one thread stalls, the second thread utilizes the freed-up computing power, accounting for 30 to 50 percent extra performance, according to AMD.

Host node for Helios

Venice is also the ‘under-the-hood’ CPU of AMD’s new Helios rack, where one Venice SP7 feeds four Instinct MI455X GPUs per compute tray. The CPU must never starve the GPU, and AMD is fully committed to that: high frequencies, more cores, and twice as fast CPU-to-GPU connections. Compared to the previous generation, it looks like this:

Host node5th Gen Epyc 9005 Turin6th Gen Epyc 9006 Venice
Frequency and cores5.0 GHz with 64 cores5.0 GHz with 96 cores
Memory12 channels DDR5, 614 GB/s16 channels DDR5/MRDIMM, 1.6 TB/s
CPU-to-GPUPCIe Gen 5PCIe Gen 6, double bandwidth

Furthermore, Verano is ready for the next generation of rack-scale systems—the LP variant with energy-efficient LPDDR5X memory that sits even closer to the GPUs as an AI host node. Over six generations, AMD now claims 18 times more throughput than the first Epyc Naples from 2017.

A story even without AI

The insight that sticks with us most from the briefing: what is currently in five-year-old data centers, and what you can replace it with. AMD calculates that a thousand outdated servers with Intel Xeon Gold 6258R processors can be replaced by 82 servers with the Epyc 9996—a consolidation of roughly twelve to one.

According to the company, this results in up to 77 percent less power consumption and up to 40 percent lower total costs over five years, plus freed-up space and power that can be invested in AI expansion. Such calculations show the best-case scenario, but the trend is clear: the business case for replacement lies as much in the energy bill as in computing power.

Hieronder vind je de volledige lijst met 9006 SP7-modellen. De volledige lijst met SP8-modellen en meer details vind je hier.

NameCPU CoresThreadsMax. Boost Clock Base ClockL3 CacheDefault CPU Power
AMD Epyc 9996256512Up to 4.1 GHz2.55 GHz1024 MB600W
AMD Epyc 9966192384Up to 4 GHz2.9 GHz768 MB600W
AMD Epyc 9846168336Up to 3.7 GHz2.85 GHz768 MB500W
AMD Epyc 9756128256Up to 4 GHz3.15 GHz512 MB500W
AMD Epyc 9G7696192Up to 4.8 GHz3.4 GHz384 MB500W
AMD Epyc 965696192Up to 3.7 GHz3.05 GHz512 MB400W
AMD Epyc 9686F96192Up to 5 GHz3.4 GHz384 MB500W
AMD Epyc 955664128Up to 4.3 GHz2.75 GHz384 MB300W
AMD Epyc 9586F64128Up to 5 GHz3.75 GHz384 MB500W

And what is Intel doing?

To put Venice into perspective, it’s worth looking at the competition. Earlier this year, Intel launched Xeon 6+, codenamed Clearwater Forest: the first server chip on its own 18A process (comparable to TSMC’s 2 nm), with up to 288 efficient E-cores. That chip aims for density and cloud-native work, but lacks the powerful P-cores for the high-end segment and thus competes at most with a portion of the Venice portfolio (think of the sandbox role).

The real answer is called Diamond Rapids, the Xeon 7 generation on 18A-P, which won’t appear until 2027. It promises up to 192 P-cores, sixteen memory channels, and PCIe 6.0: roughly the specifications that Venice is already delivering today. Notably, Intel is scrapping hyperthreading there, while AMD is highlighting simultaneous multithreading as an advantage for stuttering agent code. This gives AMD at least a year of free rein in the high-end segment, although Clearwater Forest does show that Intel’s foundry story is gradually getting back on track.

Nvidia is watching

Meanwhile, Nvidia is also directly entering the server CPU market. Its Vera processor, with 88 custom-designed Olympus cores, adjustable power consumption from 250 to 450 watts, and up to 1.2 TB/s of LPDDR5X bandwidth, is included by default in every Vera Rubin rack.

Nvidia itself is positioning the chip as a sandbox CPU for AI agents, claiming 1.8 times the performance of x86 according to its own marketing. Kuppuswamy could only chuckle at the fact that Nvidia recently published performance and consumption figures: “I thought we beat them by a smaller margin.”

Thanks to those details, we know: 2.2 times the throughput, and 20 percent more performance per core where we expected at least 10 percent. And we’re not even done tuning yet.” The real threat isn’t in those specifications, but in the bundle: anyone who buys a complete Nvidia rack gets the CPU de facto included.

Availability

Venice SP7 is in production today and is already being delivered to customers; servers from major manufacturers will follow in the fourth quarter of 2026. SP8 will appear in the first half of 2027, Venice-X and Verano in the second half of that year.

This completes the story around AI agents: Venice in the sandbox, Venice alongside the Instinct MI455X in Helios, and Venice in the regular server room. Furthermore, the timing could hardly be better: Intel won’t have a full response until 2027, and compared to Nvidia’s Vera, AMD offers more than double the throughput per socket.

“The best server CPU. Period” sounds like a bluff today, but as long as the competition cannot prove otherwise, there is no one to refute that point.