Parallel computing uses multiple hardware execution resources to make progress on more than one operation at a time. Those resources may exist inside one CPU core, across several cores or processor sockets, on a GPU, or across multiple networked machines. The useful performance gain depends not only on how much work can run in parallel, but also on memory bandwidth, communication cost, synchronization, and how well the workload matches the hardware.
Hardware and software expose parallelism at several levels:
These forms of parallelism are often combined. A cluster node, for example, may contain several CPU sockets, each CPU may contain many cores with SIMD units, and the node may also include one or more GPUs.
A single-core CPU has one physical processing core, but that does not mean it performs only one operation at a time. Modern cores exploit substantial parallelism internally through pipelining, multiple execution units, out-of-order execution, and vector instructions. The operating system can also time-slice several software threads on one core, giving them concurrent progress even though only one hardware context may be issuing instructions at a given instant.
A simplified instruction path is:
Several instructions can occupy different pipeline stages at the same time. A superscalar core may also begin multiple independent instructions in one cycle.
The amount of useful ILP depends on the instruction stream.
Clock frequency matters, but performance cannot be predicted from clock speed alone. Instructions per cycle (IPC), cache behavior, branch prediction, vector width, memory latency, and the workload itself are equally important.
A multi-core CPU contains several physical cores on one processor package. Different cores can execute different software threads at the same time, so applications with independent work can achieve true parallel execution.
Simultaneous multithreading (SMT) allows one physical core to maintain multiple hardware thread contexts and issue instructions from more than one thread. Intel commonly calls its implementation Hyper-Threading. SMT can improve utilization when one thread is stalled, but the hardware threads share execution units, caches, and bandwidth, so two SMT threads do not provide the same throughput as two independent physical cores.
Caches reduce the latency of accessing frequently used data, but multiple cores introduce a consistency problem: several cores may cache copies of the same memory location.
A cache-coherence protocol coordinates those copies so cores agree on the value of each cache line. Common protocols are based on states such as Modified, Exclusive, Shared, and Invalid (MESI and related protocols).
Coherence has a performance cost:
Cache coherence is not the same as a programming-language memory model. Correct concurrent programs still need appropriate synchronization such as atomics, locks, barriers, or message passing.
Adding cores increases potential compute throughput, but all cores still depend on finite memory and interconnect bandwidth. A program can therefore stop scaling long before all cores are computationally saturated.
Important hardware limits include:
For memory-bound workloads, adding more cores may provide little improvement once memory bandwidth is saturated.
GPUs are throughput-oriented processors designed to execute large numbers of similar operations across many data elements. They are especially effective when a workload exposes substantial data parallelism and relatively regular control flow.
I. A GPU contains many execution lanes grouped into larger compute units. Rather than optimizing a small number of threads for minimum latency, the design keeps many threads in flight so that ready work can execute while other threads wait for memory.
II. General-purpose GPU computing commonly uses programming models such as CUDA, HIP, OpenCL, and SYCL. Graphics APIs such as OpenGL, Vulkan, and Direct3D are primarily designed for rendering, although modern graphics pipelines can also expose programmable compute stages.
III. GPUs are widely used for graphics, scientific simulation, numerical linear algebra, machine learning, image and signal processing, and other workloads with high parallelism.
IV. GPUs also require very high memory bandwidth. Graphics workloads continuously read geometry, textures, depth information, intermediate render targets, and final image data. Modern GPUs therefore use high-bandwidth memory systems such as GDDR or HBM and rely heavily on caches and locality.
A framebuffer is memory that stores image data associated with rendering or display. Depending on the rendering pipeline, related buffers can include:
The exact memory layout is implementation-dependent; modern GPUs do not simply move every rendered value directly between compute units and DRAM for each operation. Caches, compression, tiling, and on-chip storage reduce unnecessary traffic.
V. Floating-point operations per frame
The computational cost of rendering a frame can be estimated in part by the number of arithmetic operations required for geometry processing, shading, lighting, image effects, and other stages. However, floating-point operation count alone does not determine performance. A frame can also be limited by:
GPUs are therefore designed to balance compute throughput with memory bandwidth and specialized graphics or matrix-processing hardware.
I. CPU Architecture Diagram
+---------+---------+---------+
| Control | ALU | ALU |
| CPU +---------+---------+
| | ALU | Vector |
+---------+---------+---------+
| Cache Hierarchy |
+-----------------------------+
| DRAM |
+-----------------------------+
II. GPU Architecture Diagram
+------------------------------------------------+
| Many Compute Units / Execution Lanes |
| [CU] [CU] [CU] [CU] [CU] [CU] [CU] [CU] |
+------------------------------------------------+
| Registers / Shared Memory / GPU Caches |
+------------------------------------------------+
| L2 Cache |
+------------------------------------------------+
| High-Bandwidth Device Memory |
+------------------------------------------------+
The diagrams are intentionally simplified. Both CPUs and GPUs contain multiple cache levels, schedulers, execution units, and specialized hardware.
| Aspect | CPU (Central Processing Unit) | GPU (Graphics Processing Unit) |
| Core architecture | Fewer, complex cores optimized for strong single-thread performance and general-purpose control flow. | Many execution lanes organized for high-throughput parallel work. |
| Control logic | Aggressive branch prediction, speculation, and out-of-order execution are common. | More throughput-oriented scheduling; groups of threads often execute in SIMD/SIMT fashion. |
| Cache system | Large cache hierarchy designed to reduce latency for diverse workloads. | Cache and explicitly managed on-chip memories are tuned for high bandwidth and data reuse. |
| Execution model | Strong at latency-sensitive, branch-heavy, irregular, and sequential work. | Strong at regular, massively parallel workloads with many independent elements. |
| Memory system | Optimized for low latency and general-purpose access patterns. | Optimized for high bandwidth; discrete GPUs often have separate device memory. |
| Parallelism | ILP, SIMD/vector operations, SMT, and multiple CPU cores. | Large-scale thread/data parallelism, usually executed in SIMD/SIMT groups. |
| Programming model | General-purpose languages plus threading, vectorization, and process libraries. | CUDA, HIP, OpenCL, SYCL, shader languages, and higher-level accelerator libraries. |
| Energy use | Varies from low-power mobile CPUs to high-power server processors. | Varies widely; high-end accelerators can consume substantial power but offer high throughput per watt for suitable workloads. |
| Best fit | Operating systems, control-heavy applications, databases, compilers, interactive workloads, and mixed computation. | Graphics, machine learning, simulation, dense numerical work, image processing, and other highly parallel workloads. |
A GPU is not automatically faster than a CPU. Data-transfer cost, problem size, branch behavior, synchronization, arithmetic intensity, and available parallelism determine whether acceleration is worthwhile.
CPUs and GPUs occupy different points on this design spectrum. CPUs generally invest more hardware in reducing the latency of a small number of instruction streams; GPUs generally invest more hardware in sustaining many concurrent operations.
A throughput-oriented design tries to maximize completed work across many threads rather than minimize the execution time of one individual thread.
CPUs also benefit from parallel threads, but their cores devote more area to features that improve the latency of general-purpose code. GPU threads are typically much lighter-weight and are most useful in large numbers.
More execution units do not guarantee proportional speedup. Common limits include:
This is why parallel speedup usually becomes sublinear as more hardware is added.
SISD (Single Instruction, Single Data) A single instruction stream operates on one data stream. A simple scalar processor is the classic example, although modern scalar CPUs may still exploit internal ILP.
SIMD (Single Instruction, Multiple Data) One instruction stream operates on multiple data elements. Vector processors and CPU SIMD/vector instructions are common examples. GPUs also use SIMD-like hardware execution internally, although their programming model is usually described as SIMT.
MISD (Multiple Instructions, Single Data) Multiple instruction streams operate on the same data stream. Pure MISD machines are rare, and real systems are seldom classified this way. Fault-tolerant redundant computation and some specialized pipelines are sometimes used as illustrative examples, but the category has no widely used general-purpose counterpart.
MIMD (Multiple Instructions, Multiple Data) Multiple instruction streams operate on multiple data streams. Multicore CPUs, multiprocessor servers, and distributed-memory clusters are common MIMD systems.
Parallel Computer Architectures
|
--------------------------------------------------------------------
| | | |
SISD SIMD MISD MIMD
| | | |
scalar processors vector / array rare / specialized ----------------
processors | |
shared-memory distributed-memory
systems systems
/ \ / \
UMA NUMA clusters MPP
Flynn’s taxonomy is useful as a high-level classification, but modern systems often combine categories. A multicore CPU is MIMD across cores while each core may also execute SIMD instructions. A GPU runs many thread groups independently, while each group is executed on SIMD-like hardware.
Keeping the distinction clear helps explain why source-level concurrency does not automatically imply hardware parallelism, and why hardware can exploit some parallelism even when the source code appears sequential.
Parallelism is often described in terms of how work is divided.
I. Data Parallelism
In data parallelism, the same or similar operation is applied to many data elements.
Data parallelism and SIMD are related but not identical: data parallelism is a software/workload property, while SIMD is one possible hardware execution mechanism.
II. Task Parallelism
In task parallelism, different tasks or functions execute concurrently, often on different data.
Task and data parallelism are often combined. A scientific application may run different simulation tasks across nodes while using SIMD instructions or GPUs inside each task.
Shared-memory systems give multiple processors or cores access to one shared address space. Communication can therefore occur through loads and stores rather than explicit network messages.
A shared address space does not mean that every write becomes instantaneously visible everywhere. Modern processors use private caches, store buffers, speculative execution, and relaxed memory ordering. Hardware cache coherence keeps cached copies of memory consistent, while the programming model still requires synchronization to establish safe ordering between threads.
There are two common organizations:
I. Uniform Memory Access (UMA)
II. Non-Uniform Memory Access (NUMA)
Modern multi-socket x86 servers are common examples of CC-NUMA systems.
These terms describe different dimensions of a system, so they should not be treated as mutually exclusive categories.
| Concept | CPUs | GPUs |
| SIMD / vector execution | CPU cores use vector instruction sets such as SSE, AVX, SVE, or NEON. | Execution lanes typically operate in SIMD/SIMT groups such as warps or wavefronts. |
| MIMD | Multiple CPU cores can execute independent instruction streams. | Different compute units or thread groups can make progress independently, even though execution within a warp/wave is SIMD-like. |
| Shared address space | Threads in a process normally share virtual memory; cache coherence coordinates copies across cores. | Threads can access device-global memory, while faster explicitly shared on-chip memory often has block/workgroup scope. Integrated or unified-memory systems may provide broader shared-address-space mechanisms. |
| UMA | Common as an abstraction on smaller shared-memory systems. | Not a defining GPU property. Some integrated systems share physical memory, but access cost can still vary by location, cache state, or processor. |
| NUMA | Common in multi-socket servers and some many-core systems. | Multi-GPU and chiplet-based accelerator systems can have non-uniform access to local versus remote memory, even when software exposes a unified address space. |
A single platform can therefore be MIMD across processors, SIMD within each processor, and NUMA in its memory organization at the same time.
Distributed and cluster computing both use multiple computers, but they emphasize different operating assumptions. A cluster is a type of distributed system whose machines are usually managed together and connected by a relatively fast network.
Distributed computing uses independent computers that coordinate over a network to provide a service or solve a problem.
I. Architecture
II. Characteristics
III. Examples
IV. Challenges
Cluster computing uses a group of machines that are usually colocated, administratively coordinated, and connected by a fast network. HPC clusters are designed to run tightly coupled parallel jobs, while data-processing clusters may emphasize storage and throughput.
I. Architecture
II. Characteristics
III. Examples
IV. Challenges
| Aspect | Distributed Computing | Cluster Computing |
| Scope | Broad category covering independent networked machines that cooperate. | A more tightly managed form of distributed computing, usually within one organization or facility. |
| Network | May span datacenters or the public internet; latency can be high or variable. | Usually uses a local, relatively high-speed network. |
| Management | Can be centralized, decentralized, or hierarchical. | Usually coordinated by shared management and scheduling infrastructure. |
| Hardware | Often heterogeneous and geographically dispersed. | Often more uniform, though heterogeneous CPU/GPU clusters are common. |
| Failure model | Partial failures and network partitions are normal design concerns. | Node failures still occur, but networking and administration are usually more controlled. |
| Typical use cases | Cloud services, distributed databases, global systems, volunteer computing. | HPC simulations, batch analytics, AI training, and tightly coupled parallel jobs. |