Shared Server

Serves everything at once, and slows under load. Input and output.

How it works#

The Shared Server takes in every entity it is given and serves them all together. It stands for an external system that answers all its callers at once and gets slower the more requests it holds: a database, a downstream API, a shared cache, a disk.

Unlike a server, it is never busy with just one entity, and nothing waits in front of it. It accepts every entity the moment it arrives: it never refuses one and never makes a sender wait. The cost of a crowd shows up inside instead. Every entity it holds progresses at the same speed, and that speed depends on how many it holds, N:

speed = 1 / (1 + contention × (N − 1) + coherency × N × (N − 1))

With one entity inside the speed is 1. It changes the moment an entity enters or leaves, for every entity inside. Each entity draws its amount of work once, on entry, from the Distribution, and leaves when that work is done — so the Service Rate describes an entity served alone, and the time an entity really takes is that work stretched by the company it kept.

With Contention and Coherency both 0, the defaults, entities do not slow each other at all and a shared server behaves like a Delay with a random latency. Contention alone makes each entity slower in proportion to the others inside, and the work completed per second levels off. Coherency does more damage: past some load the shared server completes less work the more it holds, entities arrive faster than they leave, and they pile up without limit. The curve is the shape of the Universal Scalability Law.

A shared server is passive on its input and active on its output: it is fed like a queue, and it pushes onward like a generator. So it can follow a generator, a server, a delay, a router or another shared server, and hand on to a queue, a delay, a router or another shared server. A queue cannot feed it directly — put a server between them — and a server cannot pull from it: to serve what comes out of a shared server, put a queue between them. Several components can connect into its one input.

A shared server must have an incoming connection. Its outgoing side is optional: with nothing connected, entities count as completed when their work is done and leave the model. A shared server has no Wait Timeout and does not wait at a full queue: an entity the queue has no room for is dropped at once and counted against the shared server.

Settings#

Service Rateentities per second — default 3
The mean rate μ at which one entity is served when it is alone in the shared server. Alone, the mean time is 1/μ, whichever distribution is chosen: a third of a second at the default.
Distributiondefault Exponential
How the amounts of work entities bring are spread around their mean: Exponential, Deterministic, Uniform, Erlang or Lognormal. It describes an entity served alone; company stretches every time by the same factor.
Spreadshare of the mean, 0 to 1 — default 0.5
Shown when the distribution is Uniform. How far either side of the mean an amount of work can fall.
Stagesexponential stages, 1 to 100 — default 2
Shown when the distribution is Erlang. More stages, less variation.
Variabilitystandard deviation over the mean, 0.1 to 10 — default 1
Shown when the distribution is Lognormal. The coefficient of variation of the amounts of work.
Contention0 to 1, in steps of 0.01 — default 0
The slowdown each other entity inside adds: entities sharing one resource, such as a lock, a disk or a connection. At 0.2, a second entity makes both take 1.2 times as long.
Coherency0 to 1, in steps of 0.001 — default 0
The slowdown each pair of entities inside adds: entities that have to coordinate with each other. It grows with the square of the number inside, so a small value matters once the shared server is busy.

Service Rate, Contention and Coherency can be changed while the simulation runs, and a new value applies at once to the entities already inside. So can Spread, Stages and Variability, but those apply only to entities that enter afterwards, because an entity's work is drawn on entry. The distribution itself is fixed for the length of a run. Distributions describes each one and what its shape setting does.

Statistics#

  • Requests inside — how many entities the shared server holds, over time. A line that keeps climbing means it has passed the load it can sustain.
  • Entity rate — completed and dropped entities per second over time. Drops here mean the queue after the shared server was full.
  • Time per request — the mean time an entity spent inside, in seconds, over the chart window. It is measured from the moment an entity enters to the moment its work is done.
  • After a batch run, the mean requests inside, the completed and dropped rates, and the time per request, each as a mean per run with a confidence interval. Time per request is the mean in Basic mode and the mean, median, 95th and 99th percentile in Advanced mode; an entity that entered during the warm-up is left out.

Example: a server pool#

The Server pool model in the library is ready to run. Entities arrive at 4 a second into a queue, three servers at 2 a second each take them from it, and all three hand over to a shared server named Database: 3 a second alone, Contention 0.2, Coherency 0.02. At these settings the Database holds about 2.3 entities on average and takes about 0.56 seconds per request, against 0.33 alone.

Raise Contention to 0.3 and Coherency to 0.05 while it runs. The arrival rate has not changed, but the Database can no longer keep up with it once it gets crowded, and the entities inside pile up without limit.

One limit to keep in mind: a server hands its entity over and is free at once. It does not wait for the shared server's answer, so the model is a pipeline, not a blocking call, and the number of entities inside is not capped by the size of the pool.