One Function, 92% of the CPU: Profiling an Azure Cobalt Workload with Arm Performix
Last Updated on September 25, 2026 by Editorial Team
Author(s): Dave R | Microsoft Azure & AI MVP ☁️
Originally published on Towards AI.
How targets, recipes, and an MCP server turn Arm hardware counters into a fix you can verify
Here is a result worth pausing on. When Arm Performix profiled a matrix-multiply workload on an Azure Cobalt virtual machine, a single function, matrix_multiply_impl, accounted for 92% of every CPU sample collected. The same investigation surfaced two more surprises: on a 96-core VM, roughly a third of the cores sat nearly idle under load, and the hot code already used Arm's Scalable Vector Extension, so "just add SIMD" was not the fix. None of this came from guesswork. It came from a repeatable workflow that starts with the whole system and narrows, step by step, to one line of code and a change you can measure.

The rest of the article explains why performance work often stalls on fragmented tools and manual stitching, then introduces Arm Performix as an end-to-end workflow built around five guided steps: qualify the instance, identify constraints, conduct analysis, apply insights, and repeat/iterate. It details why Azure Cobalt’s “one physical core per vCPU” makes system-level symptoms trustworthy, and shows how Performix’s targets, recipes, and stored runs (with compare/diff and baselines) turn hardware counter data into actionable code changes. Using the matrix-multiply case study, it connects system characterization (a cache “cliff”), system utilization (core heat maps and idle cores), and deeper code-level recipes (hotspots, instruction mix, and top-down microarchitecture) to a single root cause—data reuse failures in the innermost loop due to per-element dot products that thrash L1. Finally, it describes how an MCP server provides evidence to an AI coding agent to propose a specific SVE microkernel/tile-based rewrite, and stresses verification as the deciding step by rerunning against a baseline locally and even in CI so recommendations become measurable, trustworthy fixes rather than plausible guesses.
Read the full blog for free on Medium.
Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.