Comparing cost of 3 different workload configurations on private vs public cloud

I compared three workload configurations on a private cloud and on a public cloud over five years, using heavy workloads.

The detailed report was supposed to be linked from this post. It never was, so the figures are not on this page.

What the comparison showed:

  • If the goal is cost over that five-year span, the private cloud came out ahead.
  • If the goal is ease of use, the public cloud came out ahead.

Archive note, December 2017. Newer writing here is on VMware Cloud Foundation, Avi, and NSX.

Alternatives to Kubernetes

Kubernetes manages containerized applications. These were other tools people reached for at the time.

Docker Swarm

Native clustering for Docker. It turns a pool of Docker hosts into one virtual Docker host. Because Swarm serves the standard Docker API, a tool that already talks to a Docker daemon can use Swarm to scale across hosts. Tools that fit that pattern included the Docker client, Dokku, Docker Compose, Docker Machine, and Jenkins. Docker Swarm overview

Apache Mesos

A platform for fine-grained resource sharing in the data center. Introduction

Nomad

Manages a cluster of machines and runs applications on them. You declare what you want to run. Nomad decides where and how. Nomad introduction

Marathon

A framework for orchestrating containers. Marathon introduction

Written in December 2017, and left as a record of that moment. Newer writing here is on VMware Cloud Foundation, Avi, and NSX.

High-level architecture of Kubernetes

A functioning cluster needed six components:

  1. API server
  2. Scheduler
  3. Controller manager
  4. kubelet
  5. kube-proxy
  6. etcd

Each one could run as a normal Linux process, or inside a container.

Control plane

The master ran the API server, the scheduler, and the controller manager. It could be set up as multi-master. The scheduler and the controller manager elect a leader. The API servers can sit behind a load balancer.

The API server exposes a REST interface to Kubernetes resources and is configurable. The scheduler places containers on nodes using policies, metrics, and resource requests, and it can be tuned with command-line flags. The controller manager reconciles the cluster’s actual state with the desired state from the API. It is a control loop, and it is configurable too.

Workers

Every worker ran kubelet, kube-proxy, and the Docker engine. kubelet talked to Docker and kept the containers that should be running actually running. kube-proxy handled network connectivity to those containers.

The post also pointed at rkt as an alternative to the Docker engine: rkt getting started guide.

The diagram was from a Kubernetes training slide by @wattsteve.

Written in December 2017. Newer writing here is on VMware Cloud Foundation, Avi, and NSX.

What is a multicore processor and a manycore processor?

Multicore

One processor with two or more independent cores. The cores read and execute ordinary CPU instructions (add, move, branch) at the same time, which helps programs that can run in parallel.

Examples from the original note: AMD A-Series, Athlon II, FX-Series, Phenom, and EPYC; IBM POWER4, POWER5, and POWER6; Intel Core 2 Duo, Core i3, Core i5, and Core i7.

Manycore

A specialist multicore design for a high degree of parallel work, with a large number of simpler cores (tens, hundreds, or thousands). Used in embedded systems and high-performance computing.

Examples from the original note: the Sunway TaihuLight supercomputer, GPUs, Intel Xeon Phi, Tilera, and the Teraflops Research Chip.

Archive note, October 2017. Newer writing here is on VMware Cloud Foundation, Avi, and NSX.

Difference between a CPU and a GPU

A CPU is a general-purpose processor. It can in principle do any computation, but not always in the best way for that computation. You can do graphics on a CPU. A GPU built for the job will usually finish sooner.

A GPU is a special-purpose processor, tuned for the calculations computer graphics repeat, especially SIMD work.

MIMD, the shape associated here with the CPU

Many actions happen at once on many pieces of data. Adding and multiplying at the same time to solve a problem with separate parts is the example used in the original note. The work may or may not be synchronized. MIMD is used when an algorithm splits into independent parts and each part goes to a different processor.

SIMD, the shape associated here with the GPU

One identical action runs at the same time on many pieces of data: retrieve, calculate, or store. Fetching multiple files at once is the example used in the original note. Processors with their own local data execute the same instruction in step, and they communicate to shift work around. SIMD is used when a lot of computation is the same operation done in parallel.

Archive note, October 2017. Newer writing here is on VMware Cloud Foundation, Avi, and NSX.