What Is Inside a CPU?
In the previous article, we looked at storage, RAM, cache, and the CPU. Now we can look one level deeper. A CPU is not a single calculator: it contains registers, control and instruction-handling logic, and execution units that cooperate to run instructions.
The diagram is a conceptual model, not a complete design. Modern processors contain many additional structures, and their exact organization varies.
%%{init: {"themeVariables": {"fontSize": "18px"}, "flowchart": {"nodeSpacing": 36, "rankSpacing": 48}}}%%
flowchart TB
PC["Program counter<br/>instruction address"] --> Fetch[Instruction fetch]
Fetch --> Decode[Instruction decode]
Decode --> Control[Control logic]
Decode --> Registers[Register file]
Registers --> ALU[ALU<br/>integer and logic]
Registers --> FPU[Optional FPU<br/>floating point]
Registers --> LSU[Load/store unit]
Control --> ALU
Control --> FPU
Control --> LSU
ALU --> Registers
FPU --> Registers
LSU <--> Cache[Cache and memory interface]
PC --> Control
classDef control fill:#e0e7ff,stroke:#4f46e5,color:#312e81,stroke-width:2px
classDef registers fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px
classDef execute fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px
classDef memory fill:#ffedd5,stroke:#c2410c,color:#7c2d12,stroke-width:2px
class PC,Fetch,Decode,Control control
class Registers registers
class ALU,FPU,LSU execute
class Cache memory
linkStyle default stroke:#64748b,stroke-width:2px
The Main Parts
| Part | What it does |
|---|---|
| Registers | Hold the small values an instruction is currently using, along with processor state. |
| Program counter (PC) | Tracks instruction location. In RISC-V, the architectural pc holds the address of the current instruction [1]. |
| ALU | Performs integer arithmetic and logical operations such as addition, comparison, and bitwise operations. |
| FPU | Performs floating-point operations when the processor includes suitable floating-point hardware. In RISC-V, floating-point support is an ISA extension [3]. |
| Control logic and decoder | Interpret an instruction and coordinate the required register reads, execution, and result write. |
| Load/store unit | Transfers data between registers and memory. In RISC-V’s base integer ISA, arithmetic operates on registers; load and store instructions access memory [1]. |
For example, ADD x3, x1, x2 means: read values from registers x1 and x2, add them, and write the result to x3 [1]. The register names and instruction details depend on the processor’s instruction set architecture.
Fetch, Decode, Execute
A useful first model of instruction execution is:
%%{init: {"themeVariables": {"fontSize": "18px"}, "flowchart": {"nodeSpacing": 34, "rankSpacing": 44}}}%%
flowchart LR
Fetch[Fetch instruction] --> Decode[Decode operation]
Decode --> Read[Read operands]
Read --> Execute[Execute operation]
Execute --> Write[Write result]
Write --> Fetch
classDef fetch fill:#e0e7ff,stroke:#4f46e5,color:#312e81,stroke-width:2px
classDef operands fill:#ccfbf1,stroke:#0f766e,color:#134e4a,stroke-width:2px
classDef execute fill:#dbeafe,stroke:#2563eb,color:#172554,stroke-width:2px
class Fetch,Decode fetch
class Read,Write operands
class Execute execute
linkStyle default stroke:#64748b,stroke-width:2px
For ADD x3, x1, x2, the processor fetches and decodes the instruction, reads x1 and x2, performs the addition, then writes the result to x3. The PC advances to the next instruction unless a branch or jump changes the flow [1].
This cycle is a teaching model, not a claim that every modern CPU completes one instruction at a time. Pipelining lets different instructions occupy different stages at once; multiple execution units and out-of-order execution can add further overlap.
Clock Speed and Performance
A clock frequency of 4 GHz means about four billion clock cycles per second. It does not mean four billion instructions per second: an instruction may take multiple cycles, and processors may execute parts of several instructions concurrently. Performance also depends on the program, execution units, memory behavior, and parallelism.
ISA and Microarchitecture
An instruction set architecture (ISA) defines the instructions and processor state software can use. The microarchitecture is how a particular CPU implements that interface. Different processors can implement the same ISA with different pipelines, caches, and execution units [2].
Why This Matters for HPC
Scientific simulations repeat operations over many values and time steps. Their performance can therefore depend on floating-point capability, registers, cache and memory access, execution units, and how well work is parallelized. The CPU turns each program into instructions, but the data path often determines how quickly those instructions can make progress.
In Short
The CPU fetches and decodes instructions, uses registers to supply operands, performs work in execution units, and writes results back. Modern processors overlap this work, so fetch-decode-execute is a useful model, not a literal one-step-at-a-time schedule.
References
- RISC-V International, RV32I Base Integer Instruction Set
- RISC-V International, Introduction to the RISC-V ISA
- RISC-V International, RV64I Base Integer Instruction Set