Services for clusters

Services for clusters Newton, Lambda, Gipdeep

A system comprising services for cluster monitoring (Newton, Lambda, Gipdeep, Darwin). The system reads and analyzes data, and automatically manages the state in case of issues. It records the results in the relevant tables of the GPU Cluster system.

The number of nodes in the cluster is determined dynamically.

The main goals of the system are:

  • To ensure continuous access to and monitoring of clusters.
  • To provide state notifications in the form of graphs and cluster tables beyond the Technion network.

If this tool behaves incorrectly or shows inconsistent data, please report it to the DevOps team: devops@cs.technion.ac.il