Arm has launched its second-generation Compute Subsystem for Mobile, CSS for Mobile 2, at its Arm Everywhere China event in Shanghai. This new platform is engineered for the evolving demands of intelligent agents and AI-native graphics on mobile devices, promising significant boosts in performance, energy efficiency, and data transfer.
Evolving Mobile Demands
The landscape of mobile computing is rapidly shifting. Arm highlights that intelligent agents are moving beyond simple Q&A to understanding user intent and autonomously executing tasks. This, coupled with increasingly complex AI-native graphics, places new demands on mobile processing power. Arm posits that future mobile performance will hinge on three key factors: the speed of execution across different task phases, the efficiency of data movement between various compute engines, and the system’s ability to orchestrate the entire task flow.
CSS for Mobile 2 addresses these challenges by optimizing performance, power, and data transfer efficiency at a system level. This optimization extends to the physical implementation phase, allowing partners to consider power, performance, and area (PPA) alongside architectural design even before SoC integration. The platform offers flexibility, enabling partners to select individual components or combine them with their own or third-party IP.
Arm C2 CPU Cluster with Dual SME2 Engines
The CPU component of CSS for Mobile 2 features the Arm C2 CPU cluster, which integrates C2-Ultra and C2-Pro CPUs with the second-generation Arm Scalable Matrix Extension (SME2) technology. Arm reports a performance uplift of up to 1.7 times on the latest AI models, alongside a 15% increase in single-thread performance, 15% faster web browsing, 12% quicker app launches, and a 12% boost in cluster-level multi-thread performance. This allows for parallel processing of tasks such as knowledge retrieval, model preparation, and application execution.
A key enhancement in the C2 CPU cluster is the inclusion of dual SME2 engines, effectively doubling SME2 capabilities compared to the previous generation. This translates to significant improvements in AI workflows:
- Reduced latency by 40% when running advanced voice-to-text models like Moonshine and Parakeet on C2-Ultra compared to C1-Ultra, leading to faster AI interactions.
- A 41% improvement in memory information retrieval and search performance during the ‘memory’ phase.
- A 25% increase in the average speed of generating structured commands by small language models, utilizing retrieved context in the ‘inference’ phase.
In end-to-end tests encompassing voice processing, memory retrieval, inference, application execution, and web browsing, the C2-Ultra configuration delivered a 24% overall performance gain over its predecessor. Arm also introduced the Orchestration Layer, which manages task states, consolidates context information, and determines which compute resource executes each workflow stage, coordinating permissions and processes across local apps, edge models, other compute resources, or cloud services.
Mali G2-Ultra NX GPU: AI Accelerator Integrated into GPU
The GPU component sees the introduction of the Mali G2-Ultra NX, Arm’s first AI-native Mali GPU. The major innovation is the direct embedding of dedicated neural network accelerators into the GPU shader cores. This allows neural graphics workloads to run in parallel with graphics and compute tasks, leveraging the GPU’s memory system, cache coherency, and control structures, thereby minimizing data transfers between compute units.
This GPU integrates the new Execution Engine and third-generation Ray Tracing Unit (RTUv3), providing an AI-native foundation for neural graphics. It supports three neural network acceleration technologies:
- Neural Super Sampling (NSS): Reconstructs higher-resolution images from lower-resolution renders.
- Neural Frame Rate Upscaling (NFRU): Enhances frame rates by generating intermediate frames.
- Neural Super Sampling and Denoiser (NSSD): Combines upscaling with denoising for complex ray tracing scenarios.
In a demonstration developed with Sumo Digital using Unreal Engine, NFRU and NSSD achieved up to a 4x performance efficiency improvement compared to traditional rendering, with external memory traffic reduced by up to 70%. This enables advanced techniques like Unreal Engine’s MegaLights on mobile devices. The GPU is designed for sustained high frame rates, supporting stable operation at up to 120 frames per second.
The Execution Engine represents the most significant instruction set architecture upgrade in Arm Mali’s last seven generations, featuring double the number of registers per thread group and a dynamic register allocation mechanism to reduce register spills. Compared to the previous generation, benchmark performance is up to 24% higher, with non-AI gaming performance improving by 14%. In ray tracing, the new architecture reduces DRAM traffic by up to 13% in mainstream benchmarks and offers hardware support for Opacity Micro-Maps (OMM), which can increase frame rates by 30% and reduce ray tracing workloads by up to 70%.
Software Ecosystem: KleidiAI and Arm AI Portal
On the software front, CSS for Mobile 2 leverages Arm’s global developer community of over 22 million. Through deep integration with major AI frameworks via KleidiAI, it provides optimized software execution paths for Arm CPUs, allowing developers to utilize underlying capabilities like SME2 without altering their existing workflows. The Arm AI Portal serves as a unified platform offering optimized and validated models, performance and accuracy data, code samples, and deployment resources. Arm MCP servers bring these capabilities into the intelligent agent AI development environment, accelerating the developer’s journey from evaluation to deployment.
CSS for Mobile 2 offers chip partners a flexible, configurable foundational platform, and its future integration into terminal chips and finished products will be interesting to follow.








