In the rapidly evolving landscape of cloud-native infrastructure, the metrics used to define success are undergoing a rigorous reassessment. Adrian Cockcroft, a veteran architect known for his pivotal roles at Sun Microsystems, Netflix, and Amazon, recently argued that our industry-wide fixation on the P99 metric is fundamentally insufficient for modern web services. Speaking ahead of the upcoming P99 CONF 2026, Cockcroft emphasized that in the era of complex, distributed systems, traditional percentiles often mask the true behavioral characteristics of an application.
Cockcroft’s career has been defined by his ability to demystify complex systems. During his tenure at Sun, he famously bypassed opaque manuals to read kernel source code directly, eventually documenting his findings in seminal texts like “Sun Performance and Tuning.” Today, while he acknowledges that modern observability tools offer far more visibility than the days of simple system stats, he remains critical of how that data is interpreted. His core contention is that percentiles provide a collapsed view that obscures the “why” behind latency spikes.
He posits that most service distributions are multimodal. A system might show a fast response peak driven by cache hits and a slower peak triggered by cache misses. As cache performance fluctuates, the heights of these peaks shift, yet their positions on a latency scale remain constant. Consequently, the mean, standard deviation, and P99 fluctuate in ways that appear erratic, even though the underlying mechanics of the system remain consistent. By relying on a single number to represent these complex distributions, engineers are often flying blind, misinterpreting natural caching cycles as system anomalies.
To move beyond these limitations, Cockcroft advocates for a move toward histogram-based peak analysis, which tracks the fluctuation of distinct response peaks over time. Historically, building the custom tooling required to visualize these distributions was a time-intensive endeavor. However, Cockcroft suggests that generative AI has fundamentally altered the economics of engineering efficiency. He describes current development as “vibe coding,” where LLMs allow him to generate specialized diagnostic tools in minutes rather than days. This speed, he notes, grants him an “infinite” relative speedup, enabling the creation of bespoke solutions for granular analysis that previously would have been discarded due to time constraints.
For practitioners currently managing high-performance environments, Cockcroft’s advice remains consistent with his engineering philosophy: adopt a macroscopic view to identify interesting telemetry patterns, then use that focus to narrow down to individual, slow-moving requests. By applying what he calls a “microscope” approach—moving from low-resolution monitoring to high-resolution, end-to-end inspection of specific anomalies—teams can move past the limitations of static dashboards.
As the industry pivots toward more autonomous agents and distributed AI workloads, the ability to build, iterate, and deploy custom performance tooling will be a primary competitive advantage. Cockcroft’s recent work serves as a reminder that while the tooling has evolved from kernel-level reading to AI-assisted code generation, the fundamental requirement remains the same: an engineer must understand the deep-seated mechanics of their environment before they can truly optimize it.
Source: The New Stack
