In the May 2026 edition: hunting CPU zombies in Kubernetes, when observability becomes the outage, and how JDK 27’s compact headers and G1GC shifts cut your pod memory footprint. The Optimal is the Akamas newsletter where DevRel Engineer Graziano Casto breaks down Kubernetes performance, reliability, and cost optimization.
Hello everyone! Graziano here.
If you’re reading this, we are officially wrapping up May 2026 and I am still catching my breath from an incredible run of conferences. From North America to Europe, it’s been fantastic connecting with so many of you to talk about the reality of performance engineering. Grab a coffee and let’s look at the logs for this month.
The Latest Replica
Hunting CPU Zombies in Kubernetes
A fascinating post from the Pinterest engineering team recently detailed a massive troubleshooting journey where their distributed machine learning jobs were mysteriously crashing due to network driver resets. After deep diving with perf and flamegraphs, they traced the root cause to CPU starvation triggered by “zombie” cgroups. It’s a brilliant real-world story that proves high-level metrics aren’t enough. When a single CPU core pegs at 100% for just a few seconds, it can choke network threads and bring down expensive workloads. It’s a stark reminder that robust Kubernetes scaling requires deep, continuous profiling at the OS and runtime levels, not just surface-level dashboarding.
When Observability Becomes the Outage
We all rely on telemetry to keep our systems running, but what happens when your monitoring is the very thing bringing your cluster down? A recent DZone article tackles this exact irony. Unoptimized observability, think excessive logging, high-cardinality metrics, and aggressive tracing, doesn’t just inflate your SaaS bills; it consumes vital CPU and memory resources directly on your nodes. It highlights a critical point we often emphasize: observability agents and sidecars need to be tuned just like any other application. If you aren’t optimizing your telemetry footprint, your monitoring strategy is actively contributing to cluster inefficiency and performance bottlenecks.
JDK 27: Compact Headers and G1GC Everywhere
The latest OpenJDK news roundup confirms some major shifts coming down the pipeline that will directly impact how we size our pods. First, G1GC is proposed to become the default garbage collector across all environments, not just server-class setups. Even more impactful, Compact Object Headers (JEP 534) are targeted to be enabled by default. This is huge news for memory footprint reduction. For those of us obsessed with Kubernetes pod right-sizing, smaller object headers mean less heap consumption and, ultimately, a lower risk of facing the dreaded K8s OOMKill.
Commits From The Lab
Upcoming Webinar: GenAI Optimization with Red Hat
Speaking of the intersection between AI and performance, we are thrilled to announce a joint webinar with Red Hat happening next week! We will be diving deep into “GenAI Optimization”, exploring the challenges of running LLM inference on Kubernetes. We’ll discuss how to properly size and tune your infrastructure to support these massive demands without letting your cloud bill spiral out of control. If you want to see how continuous optimization can tame GenAI infrastructure costs, you won’t want to miss this one.
Catch Us
Want to meet the Akamas team in person? We’re regularly at conferences and meetups across the Kubernetes and Java ecosystems. See where we’ll be next on our Events page, and follow me on LinkedIn for the talks and sessions I’ll be at.
Keep those CPUs cool and your latencies low.
Stay optimized,
Graziano Casto, DevRel @ Akamas

