At its Arm Everywhere China 2026 event in Shanghai, Arm launched the CSS for Mobile 2, its second‑generation compute subsystem purpose‑built for agentic AI workloads. The platform integrates a new C2 CPU cluster with SME2 matrix extensions, and debuts Arm's first AI‑native GPU core, the Mali G2‑Ultra NX – embedding a neural accelerator directly into the GPU for native neuro‑graphics processing.

The shift to agentic AI – where apps perceive, reason, and act in continuous loops – demands lower latency, stronger parallel compute, and efficient task scheduling under strict power and thermal constraints. CSS for Mobile 2 is not just a CPU or GPU IP update; it's a complete mobile compute stack spanning CPU, GPU, system interconnect, physical implementation, and software toolchains.

C2 CPU Cluster: Combines C2‑Ultra (for single‑thread responsiveness) and C2‑Pro (for sustained performance). Single‑thread performance is up 15% higher , web browsing 15% faster , app launch 12% faster , and AI model execution up to 1.7x faster than the previous generation. The cluster supports up to 14 cores with independent frequency scaling, and SME2 delivers 6 TOPS of matrix compute – doubling the previous generation's matrix processing power.

Mali G2‑Ultra NX: The first GPU from Arm to integrate a neural accelerator (NX) into its shader cores, supporting INT8/INT16 inference at up to 2x core clock speeds. Three open‑source neuro‑graphics technologies – Neural Super Sampling (NSS), Neural Frame Rate Upscaling (NFRU), and NSSD (denoising + super‑resolution) – enable high‑quality graphics at low power (under 1W for NSS). Xiaomi's Xring O3 flagship SoC features a 16‑core Mali G2‑Ultra NX GPU.
Third‑gen Ray Tracing Unit (RTUv3): Adds geometry optimization (shared edges, repeated edge culling, geometry duplication elimination) to reduce DRAM traffic by up to 13% in ray‑traced scenes. Integrated OMM (Opacity Micro‑Mesh) technology cuts ray‑tracing load by 70% in specific game scenes, enabling console‑quality ray tracing on mobile.
SI L2 System Interconnect: Reduces CPU‑to‑DRAM load latency by 45% – critical for frequent memory‑access patterns in agentic workflows. Compatible with LPDDR5, LPDDR5X, and LPDDR6.
Software ecosystem: KleidiAI library (200+ optimized micro‑kernels) integrated into 14+ AI frameworks including ExecuTorch and Llama.cpp. Arm AI Portal provides pre‑optimized model libraries, quantization tools, and upcoming MCP server interfaces for agentic applications.
Strategic direction: Arm leaves NPU innovation to partners, focusing on CPU/GPU compute subsystems – while making standardized AI acceleration accessible through SME2 and Mali G2.
ICgoodFind Takeaway:
CSS for Mobile 2 signals a system‑level shift – from peak performance to sustained, low‑latency agentic AI execution. For SoC designers, this is the blueprint for next‑gen mobile compute. For buyers, it defines the hardware capabilities that will differentiate flagship devices in the AI era.