This September, four Linux Foundation communities gathered in Shanghai.
KubeCon + CloudNativeCon, the OpenInfra Summit and the PyTorch Conference came together under one banner, and the message behind the co-location was hard to miss: the open-source and open infrastructure stack are becoming the default foundation for AI.
The Linux Foundation is clearly attempting to contain the geopolitical fragmentation of the opensource landscape, for example, via still organizing global conferences and striving for global software communities.
We would like to share the themes that stood out, a few sessions worth your time when the session recordings land on the CNCF YouTube channel, and some thoughts on what this means for those of us building and operating cloud infrastructure.
The strongest thread running through the keynotes was that AI has moved from an experiment on the side to the primary workload our infrastructure has to serve — and increasingly that workload is inference, not just training.
The interesting engineering problem is no longer "can we train the model" but "how do we schedule scarce accelerators between training and inference as demand shifts throughout the day."
That idea showed up concretely in a keynote on Tidal Auto-scaling for Training and Inference Based on Kubernetes + KEDA, where the pitch was exactly this: treat training and inference as tenants competing for the GPUs and let the control plane move capacity between them based on real demand. It is a very cloud-native answer to a very expensive problem.
A trend opposite to the cloud development of the past years is that AI workloads require a heterogeneous compute environment: the accelerator landscape is booming, and the infra must make it available.
Jonathan Bryce (OpenInfra / Linux Foundation) and speakers from the PyTorch Foundation, MiniMax and others reinforced the point from the model side: frontier intelligence is increasingly being built in the open, with open weights and open tooling, and it needs open infrastructure underneath it to run at scale.
Agents as the next unit of work (next step after model as a service)
Our favorite conceptual takeaway came from a keynote framing that we think will age well:
"Virtual machines gave us infrastructure automation. Containers gave us portable applications. Kubernetes gave us a control plane for distributed systems. Agents may become the next unit of work."
The speaker pointed out that an agent is not just an LLM call. It has identity, memory, goals, tools, permissions, execution state, and the ability to affect the outside world.
And crucially, it can fail in ways traditional software does not:
- by taking the wrong action
- using the wrong tool, leaking context
- exceeding its authority
- or simply behaving differently as its underlying models and environment change over time.
If agents really are the next unit of work, then the industry needs a place to *run* them safely. That is where the infrastructure sessions got interesting.
Running agents safely: sandboxes, Kata and confidential computing
Several talks converged on the same architecture for isolating untrusted, autonomous workloads. The pattern, roughly:
"Kubernetes Agent Sandbox manages its lifecycle, while Kata Containers provides a dedicated guest-kernel boundary. The rest of the chain handles container execution and efficient delivery of images and model artifacts.
In other words: Kubernetes orchestrates the agent's lifecycle; Kata Containers gives each agent a real kernel boundary instead of a shared one, and the ecosystem around it handles image and model-artifact delivery efficiently.
How an agent should be properly sandboxed is a topic of active experimentation and development - which technique will become standard is still an open question.
The natural next step, and one the community is clearly leaning into, is Kata Containers evolving toward confidential containers with GPU support, bringing hardware-backed isolation to accelerated AI workloads.
For anyone thinking about multi-tenant AI platforms — and for telco-grade environments in particular — this is a space worth watching closely.
Where Kubernetes fits
CNCF stresses the importance of Kubernetes certification, which Ericsson CCD already passes.
The k8s Gateway API is becoming increasingly popular with dozens of implementations - these seem to be mostly focused on not just IT but on Telco use cases also.
Increased interest in using Cluster API (CAPI) to declaratively manage the infrastructure clusters. Lennart Jern and Peppi-Lotta Kurjenhovi from EST already work on the maintenance of CAPI/CAPO and part of the release team.
Where OpenStack fits
For those of us with an OpenStack background, one line from the summit captured a shift that has been quietly happening for years:
"OpenStack and Kubernetes were once considered separate ecosystems, but today many operators run OpenStack on top of Kubernetes to get the benefits of both worlds."
A standout session here was Declarative Underlays: Scaling Purpose-Built Infrastructure (Clusters) for OpenStack with Cluster API by HanSol Park and Kangsub Song from Samsung SDS.
Seeing Cluster API (CAPI) and CAPO (Cluster API for OpenStack) used to declaratively manage the infrastructure clusters that OpenStack itself runs on shows the two worlds reinforcing each other rather than competing.
A couple of other OpenStack-flavored highlights from our week
- Redesigning OpenStack i18n for the AI era: Weblate & Zero-GPU AI
- A look at modernizing OpenStack's translation workflow by migrating from Zanata to Weblate. This is very much a live topic: it will need project-team coordination and is heading to the next PTG for discussion.
- An OpenClaw use case told from the perspective of an OpenStack developer — a nice reminder that our community keeps finding new, practical ways to apply the tools we build, and how to use these new tools without letting them destroy anything.
Conversations off the schedule
Some of the most valuable moments happened between sessions:
- A chat with a Cinder developer about backing up very large volumes — a scenario that maps directly onto the kind of workloads typical telco customers run, and something I want to keep an eye on.
- Discussions with OpenInfra representatives about pushing for more in-person design gatherings (PTG-style events) for project teams.
- There was broad agreement that face-to-face design time still delivers something that async collaboration cannot fully replace.
- We had a chance to meet Ildikó Váncsa an ex-Ericsson colleague, now working for OpenInfra foundation.
Takeaways
A few things we are carrying home:
1. AI inference is now a first-class infrastructure workload, and dynamic scheduling between training and inference is where a lot of the operational value is.
2. Agent isolation/sandboxing is an emerging infrastructure requirement. The Kubernetes + Kata + confidential-computing direction is worth tracking for any secure, multi-tenant platform.
3. OpenStack and Kubernetes are converging in practice, and Cluster API is one of the concrete bridges — relevant for anyone operating large private clouds.
4. The open-source community still values in-person collaboration, and stays active in that community (PTGs, upstream work, i18n modernization etc...) keeps us close to where the platforms are heading.
Sessions were recorded and soon should be appearing on the CNCF / OpenInfra Foundation / PyTorch YouTube channels.
Presentation slides are already available via the schedule:
https://kubecon-cloudnativecon-openinfra-pytorch-2026.sessionize.com/schedule/day/20260908