Serial and Parallel Computing
An ordinary computer can perform a calculation one step after another. This is called serial computing. HPC makes greater use of parallel computing: when a problem can be divided into parts, multiple processors can work on those parts at the same time.
Imagine a problem with many calculations. In a serial approach, the work is done in sequence:
flowchart LR
subgraph Serial[Serial]
direction LR
S1[Task 1] --> S2[Task 2] --> S3[Task 3] --> S4[Task 4]
end
style Serial fill:#f8fafc,stroke:#94a3b8,color:#334155,stroke-width:2px
classDef task fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px
class S1,S2,S3,S4 task
linkStyle default stroke:#64748b,stroke-width:2px
If the problem can be divided, some of those calculations can happen at the same time:
flowchart TB
subgraph Parallel[Parallel]
direction LR
T1[Task 1] --> C1[Core 1]
T2[Task 2] --> C2[Core 2]
T3[Task 3] --> C3[Core 3]
T4[Task 4] --> C4[Core 4]
end
style Parallel fill:#f0fdfa,stroke:#5eead4,color:#134e4a,stroke-width:2px
classDef task fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px
classDef core fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px
class T1,T2,T3,T4 task
class C1,C2,C3,C4 core
linkStyle default stroke:#64748b,stroke-width:2px
This can reduce the time needed for a computation. But not every problem can be perfectly parallelized. Some calculations depend on earlier results, so the amount of useful parallelism depends on the problem.
In a simple way:
flowchart LR
P[Large problem] --> D[Smaller tasks] --> R[Parallel processing]
classDef problem fill:#ffedd5,stroke:#c2410c,color:#7c2d12,stroke-width:2px
classDef tasks fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px
classDef parallel fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px
class P problem
class D tasks
class R parallel
linkStyle default stroke:#64748b,stroke-width:2px
High Performance Computing (HPC) uses computing resources to solve large or demanding problems. Parallel processing is an important part of HPC, but a complete HPC system also needs memory, storage, networking, and software to manage the work [1, 2].
Why Do We Need HPC?
We use HPC when a computation is too large or time-consuming for a single computer to handle efficiently. Scientific simulations, engineering calculations, and large-scale data analysis are some examples [2, 4].
What Is an HPC Cluster?
An HPC cluster is a group of connected computers that work together as a computing system. The computers are called nodes. Each node has its own processors, memory, and other resources, and the nodes communicate over a network [1].
flowchart TB
Scheduler[Scheduler] --> Network[Cluster network]
Network --> N1[Compute node 1]
Network --> N2[Compute node 2]
Network --> N3[Compute node 3]
classDef coordinator fill:#ffedd5,stroke:#c2410c,color:#7c2d12,stroke-width:2px
classDef network fill:#e0e7ff,stroke:#4f46e5,color:#312e81,stroke-width:2px
classDef node fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px
class Scheduler coordinator
class Network network
class N1,N2,N3 node
linkStyle default stroke:#64748b,stroke-width:2px
So an HPC cluster is not just one extremely powerful computer. It is a group of computers that can share a workload when the software and problem allow it.
What Is a Node?
An individual computer in a cluster is called a node. A node may have:
- CPU cores
- RAM
- storage
- network interfaces
- GPUs, on some systems
Different workloads benefit from different resources. For example, GPUs can accelerate some workloads with many parallel mathematical operations, but they are not required for every HPC job [2].
flowchart TB
Node[Compute node]
Node --> CPU[CPU cores]
Node --> RAM[Memory]
Node --> Storage[Storage]
Node --> Network[Network interface]
Node --> GPU[Optional GPU]
classDef node fill:#e0e7ff,stroke:#4f46e5,color:#312e81,stroke-width:2px
classDef compute fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px
classDef memory fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px
classDef io fill:#ffedd5,stroke:#c2410c,color:#7c2d12,stroke-width:2px
class Node node
class CPU,GPU compute
class RAM memory
class Storage,Network io
linkStyle default stroke:#64748b,stroke-width:2px
Scale-up and Scale-out
There are two common ways to increase the resources available to a computation [1].
Scale-up
With scale-up, a single computer uses more of its own resources, such as additional CPU cores or memory.
flowchart LR
subgraph ScaleUp[Scale-up: one computer]
U1[CPU] --- U2[More cores and memory]
end
classDef hardware fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px
class U1,U2 hardware
linkStyle default stroke:#64748b,stroke-width:2px
Scale-out
With scale-out, work is distributed across multiple computers connected through a network. This is one way to use an HPC cluster.
flowchart LR
subgraph ScaleOut[Scale-out: multiple computers]
O1[Node 1] --- Net[Network]
Net --- O2[Node 2]
Net --- O3[Node 3]
end
classDef node fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px
classDef network fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px
class O1,O2,O3 node
class Net network
linkStyle default stroke:#64748b,stroke-width:2px
How Does an HPC Job Work?
The exact workflow depends on the software and the problem, but a typical HPC job follows these steps:
- Prepare the problem. The program and input data are prepared for computation.
- Divide the work. If possible, the problem is split into tasks or data regions that can be processed in parallel.
- Request resources. A job scheduler manages the cluster and assigns resources such as nodes, CPU cores, memory, and sometimes GPUs. Slurm is one example of a workload manager [3].
- Compute and communicate. The assigned resources perform their parts of the calculation. They may exchange data or intermediate results as the program runs.
- Collect results. The program saves or combines its output for analysis.
In this series I will explain SLURM, the workload manager used on many HPC systems, including YZ HPC, which I am using. We will look at how to request resources and run jobs in the next articles.
flowchart LR
Input[Program and data] --> Submit[Submit job]
Submit --> Schedule[Scheduler allocates resources]
Schedule --> Run[Parallel computation]
Run --> Output[Results]
classDef prepare fill:#ffedd5,stroke:#c2410c,color:#7c2d12,stroke-width:2px
classDef schedule fill:#e0e7ff,stroke:#4f46e5,color:#312e81,stroke-width:2px
classDef compute fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px
classDef result fill:#dcfce7,stroke:#15803d,color:#14532d,stroke-width:2px
class Input,Submit prepare
class Schedule schedule
class Run compute
class Output result
linkStyle default stroke:#64748b,stroke-width:2px
In Short
An HPC cluster is a collection of interconnected computers that work together to solve large computational problems using parallel computing.
References
- Intel, What is HPC?
- IBM, What is high-performance computing (HPC)?
- SchedMD, Slurm Workload Manager overview
- NVIDIA, High-performance computing