Speaker
Description
This session takes a practical look at what it actually means to build and run an AI Factory using open-source technologies. The goal is simple: show how to stand up a platform that can deliver GPU as a Service, Kubernetes as a Service, and LLM services in a way that fits alongside existing HPC environments—especially those built around Slurm.
We’ll dig into architectures that bring together cloud-native tools and traditional HPC systems. How do you run containerized workloads and batch jobs side by side without things getting messy? How can Kubernetes and GPU orchestration frameworks plug into what you already have instead of replacing it? Those are the kinds of questions we’ll tackle.
From there, we’ll break down the core pieces of an AI Factory, including resource abstraction, multi-tenancy, scheduling, and how services are exposed to users. A big part of the discussion will focus on running AI workloads—like LLM inference—efficiently on shared GPU infrastructure, without wasting resources or overcomplicating operations.
We’ll also spend time on the day-to-day reality: monitoring, automation, security, and access control. We will share real deployment patterns from research and enterprise environments and discuss what actually works. If you’re operating infrastructure or designing platforms, this session should give you a clear, workable path toward combining HPC and cloud-native AI.