Speaker
Description
Dataverse is an open source data repository solution with increased adoption by research organizations and user communities for data sharing and preservation. Datasets stored in Dataverse are
cataloged, described with metadata, and can be easily shared and downloaded. After having dedicated a few years to the development and integration of a Dataverse based repository for research data, we have decided to improve the underlying hardware and high availability architecture.
In this presentation we will share some of the changes we did to the application, and the major hardware and software improvements that allows the service to be fault tolerant and improve its high availability.
With the usage of simple tools like haproxy, zookeeper, a distributed database with the patroni plugin and dns round robin, we were able to increase the availability in a wide range of fault situations.