HyperAIHyperAI

Command Palette

Search for a command to run...

Arm Debuts Neoverse CSS N4 Platform With Up to 128 Cores

Arm has officially introduced the Neoverse CSS N4, codenamed Ranger, its next-generation semi-custom compute subsystem platform designed for high-efficiency cloud and data center architectures. Fabricated on TSMC’s N3P process, the platform supports configurations ranging from eight to 128 Neoverse N4 cores per die, with clocks reaching up to 3.8 GHz. Arm intends the architecture to scale beyond single-die limits through multi-chiplet and multi-socket designs enabled by UCIe interconnects. Each configuration features up to 256 MB of shared L3 cache, 2 MB of L2 cache per core, and integrated support for DDR5 or LPDDR6 memory, alongside 128 lanes of PCIe 6/7 and CXL 4.0 bandwidth. Compared to the preceding Neoverse CSS N2, which topped out at 64 cores and 64 MB of L3 cache, the N4 platform promises significant efficiency and throughput gains. Arm estimates the new architecture delivers double the socket performance and 1.75 times the memory bandwidth of the Neoverse N3, while improving performance per watt by 25 percent. As a semi-custom program, CSS allows silicon partners to tailor core counts, cache hierarchies, and I/O interfaces to specific workload requirements, a model already proven across hyperscale servers and dedicated data processing units. While the N4 cores, internally known as Dionysus, are optimized for performance-per-watt rather than peak throughput, Arm expects the platform to enter production alongside the upcoming Neoverse V4 family, though specific commercial integrations have not yet been disclosed. In parallel with the N4 launch, Arm expanded its lineup of third-party deployments for its proprietary AGI processor, a high-performance CPU built around Neoverse V3 cores. The chip, fabricated on a 3nm node, features a dual-die architecture housing up to 136 cores, 272 MB of L3 cache, and speeds up to 3.7 GHz. By integrating memory controllers and I/O directly onto the compute die, Arm reports memory latency below 100 nanoseconds and supports up to 6 terabytes of total memory capacity per chip at speeds reaching DDR5-8800. The company projects the AGI will deliver more than twice the rack-scale performance of contemporary x86 systems based on internal modeling. New partners confirmed for AGI deployment include Oracle and ByteDance, joining existing commitments from Meta, Lenovo, SAP, OpenAI, and Cloudflare. The announcement reinforces Arm’s broader strategy of supplying validated building blocks for custom silicon while establishing its own processor as a benchmark for advanced data center workloads.

Related Links