Last week I attended KubeCon + CloudNativeCon North America 2023 in Chicago. I prioritize this conference every year, because it is the fastest way to learn what the cloud-native community is building and where it is heading.
Three themes dominated the event. Generative AI took over the keynotes, and Kubernetes turned out to be the infrastructure underneath most of it. Sustainability moved from discussion to working tooling, with projects that measure and then cut the carbon footprint of a cluster. Cost optimization filled a whole track, and nearly every session landed on the same two problems: pod sizing and autoscaler configuration.
Below are my key takeaways on the areas I find most crucial: cost efficiency, performance, and reliability in the K8s ecosystem. I hope they are useful for others as well.
Let’s dive in!

GenAI on Kubernetes: the next big thing in the K8s world
I have to admit, I was surprised by how much space Generative AI took during the keynotes. It underscores once again how AI is revolutionizing every sector of the computing industry.
But what does GenAI have to do with K8s? As Priyanka Sharma, CNCF Executive Director, put it: not many people know that Kubernetes is the infrastructure powering the GenAI rocket.
During her presentation, Priyanka showcased a chatbot powered by a Large Language Model (LLM) straight from her laptop. There was a minor hiccup when the database pod ran into trouble. I suspect the dreaded OOM kill or CPU throttling hit the KubeCon keynote.
The keynotes also went into the complexities of running AI workloads on Kubernetes. Representatives from NVIDIA discussed how hard it is to rightsize GPUs and manage AI workload performance. That job is markedly different from handling traditional web applications, the initial focus of K8s. A new Dynamic Resource Allocation model is in the works. It will allow better management of GPUs, including sharing one GPU among multiple pods.
Google’s Tim Hockin gave an insightful keynote on the future of Kubernetes. As a co-founder of K8s, he emphasized the challenge that AI workloads will pose for the platform. He predicted these workloads will demand substantial resources, potentially thousands of cores. Managing those resources efficiently is crucial, and not only for cost. It matters for sustainability too.
Efficiency and performance tuning are going to be critical for the future success of Kubernetes. We are going to see more configuration options to fine-tune resource allocation for AI workloads. That is great news. It also means K8s users get a few more “knobs” to figure out.
Interesting times ahead if you are interested in K8s optimization!
Kubernetes sustainability: measure the carbon footprint, then cut it
Sustainability really took the spotlight in the keynotes and the sessions. Why the focus? Data centers gobble up about 2% of global power, and we are nudging close to that critical 1.5 °C global warming limit. It is a wake-up call for the IT world to step up.
Before we can make a dent in this, we need the right metrics to pinpoint where the carbon goes. That is the tricky part of sustainability. A clear picture of the carbon footprint of software and infrastructure is not straightforward to get yet.
Four projects presented at KubeCon 2023 attack that visibility gap:
- Software Carbon Intensity specification, developed by the Green Software Foundation. It calculates the carbon intensity of software.
- Carbon Aware SDK. It helps identify the most sustainable energy sources.
- Kepler. An open-source project that monitors energy consumption through Prometheus and eBPF.
- Carbon-aware KEDA operator. It scales workloads based on the carbon emissions of different cloud regions.
Once you have visibility, you can start optimizing. Rightsizing your K8s pods can cut CO2 emissions by 45%, according to the figures shared in the kube-green session. Or you can turn the pods off entirely. That is the goal of kube-green, a project developed by our friend Davide Bianchi of Mia-Platform. kube-green shuts down Kubernetes resources during inactive periods.
Proper workload scheduling can also be effective, which is where the carbon-aware KEDA operator comes in.
Overall, it was encouraging to see sustainability as a focal point and to see the progress being made. In IT, energy waste is all too common, and we all need to play a part in reducing emissions. Interest in actual carbon reduction is still growing. The good news is that cost optimization is a far more common concern, and it often leads to carbon savings as well.

Kubernetes cost optimization roll call
This year, cost optimization emerged as a significant focus. Teams increasingly recognize how inefficient a Kubernetes cluster can be. With a strong optimization strategy, they can achieve substantial cost and carbon savings.
15,000 Minecraft players vs. one K8s cluster. Justin Head demonstrated that running K8s on bare metal led to 65% cost savings. His team used the K8s Cluster API and other open-source tools to get cloud-native benefits on a bare metal setup. One clever move was tuning K8s scheduler policies for more efficient pod placement, using NodeResourcesFit.MostAllocated. See the K8s scheduler docs for more info. Slide deck here.
Sustainable scaling of Kubernetes workloads. Vinay Kulkarni from eBay highlighted the challenge of defining pod resource requirements accurately. Efficiency is crucial, but under-resourced pods put performance and availability at risk. eBay is exploring in-place pod resizing to adjust settings without restarts, plus AI for automated pod sizing. The goal is to optimize resource use without missing Service Level Objectives (SLOs). Slide deck here.
Kubernetes on a budget: how to get pay-per-use right. Karim Lakhani from Intuit discussed cost-optimizing a high-scale API gateway across more than 30 clusters. His team fine-tuned Horizontal Pod Autoscaler (HPA) configurations and cluster autoscaler settings. They adjusted scaling metrics and pod resource limits manually. Slide deck here.
K8s pod autoscaling with application-aware AI. At Akamas, hosted by Dynatrace, I presented a new K8s optimization case study. We focused on configuring HPA and pod resources to minimize cost while keeping applications reliable. Our approach directly addresses the HPA tuning challenges the Intuit team described, and it makes scalable, cost-effective applications easier to reach. Slide deck here.

In summary, it is encouraging to see K8s teams increasingly aware of the vast optimization potential within the platform. K8s offers excellent capabilities, but its full efficiency and scalability benefits require proactive tuning. It is also exciting to observe the growing trend of using AI for K8s optimization. At Akamas, we have long advocated for AI-driven optimization, and we have built an enterprise platform that addresses both cost and reliability challenges in K8s.
FAQs
What were the main themes at KubeCon 2023?
Three themes dominated KubeCon + CloudNativeCon North America 2023 in Chicago. Generative AI took over the keynotes. Sustainability moved from discussion to working tooling. Cost optimization filled a track, and the sessions converged on two problems: pod sizing and autoscaler configuration.
How much CO2 can rightsizing Kubernetes pods save?
Rightsizing Kubernetes pods can cut CO2 emissions by 45%, according to figures shared in the kube-green session at KubeCon 2023. Shutting workloads down during inactive periods saves more. The kube-green project automates that shutdown.
Which tools for measuring Kubernetes carbon footprint were presented at KubeCon 2023?
Four projects stood out. The Software Carbon Intensity specification calculates the carbon intensity of software. The Carbon Aware SDK identifies cleaner energy sources. Kepler monitors energy consumption through Prometheus and eBPF. The carbon-aware KEDA operator scales workloads by cloud-region emissions.
Why are AI workloads harder to run on Kubernetes than web applications?
AI workloads need GPUs, and rightsizing a GPU is harder than rightsizing CPU and memory. NVIDIA speakers made this point at KubeCon 2023. Kubernetes was originally built for traditional web applications. A new Dynamic Resource Allocation model was in development to manage GPUs better, including sharing one GPU across pods.
What did teams share about Kubernetes cost optimization at KubeCon 2023?
Justin Head reported 65% cost savings from running K8s on bare metal. eBay described in-place pod resizing and AI-driven pod sizing to protect SLOs. Intuit fine-tuned HPA and cluster autoscaler settings by hand across more than 30 clusters. Akamas presented application-aware AI for pod autoscaling.

