Speaker
Description
Services hosted in federated research clouds are increasingly expected to stay available at all times. Yet the most common high-availability techniques, such as Kubernetes clusters, local HAProxy with Keepalived, or replicated virtual machines, all operate within a single cloud site. They cannot protect a service when the whole site becomes unavailable, and monitoring of the EGI Federated Cloud shows that such full-site outages happen regularly. Deploying instances of a service on several independent providers removes this dependence on one site. It also raises a new problem: users need a single, stable entry point that always leads them to a healthy instance, and that entry point must not itself be tied to one site.
We present the design and implementation of the FedCloud HA Load Balancer, a shared, open-source service that gives any HTTP/HTTPS service in the federation multi-site availability without requiring its owners to operate load-balancing infrastructure. The system separates a data plane from a replicated control plane. The data plane consists of identically configured HAProxy nodes on geographically distributed sites, which share rate-limiting state and check the health of backend instances. The control plane has four parts:
- a lightweight registry, integrated with EGI Check-in, that validates registrations, approves routine changes automatically, and publishes signed, versioned configuration snapshots with canary rollout;
- a quorum-based manager that probes nodes from several sites and publishes all healthy nodes through low-TTL DNS pool records;
- central certificate management;
- deduplicated monitoring and alerting.
A key design principle is that the data plane never depends on the control plane to keep serving traffic: nodes continue with their last valid configuration if any control-plane component fails. Service owners only deploy instances on different providers and submit a short declarative registration.