GPU Architecture MCQs (Multiple-Choice Questions)

Practice GPU Architecture MCQs to test your knowledge of GPU design, streaming multiprocessors, execution units, warps, memory hierarchy, caches, GPU scheduling, and parallel execution. These questions cover important architectural concepts used in modern GPUs for graphics, scientific computing, artificial intelligence, and high-performance computing. They are useful for students, developers, system architects, and professionals preparing for technical interviews. The set includes both fundamental concepts and architecture-focused scenarios that help connect hardware organization with application performance.

GPU Architecture MCQs

These GPU Architecture multiple-choice questions cover the internal organization and execution model of modern graphics processing units. The questions include topics such as SMs, SIMT execution, warp scheduling, registers, shared memory, cache hierarchy, memory bandwidth, occupancy, instruction throughput, and specialized accelerators. They also examine how architectural characteristics influence kernel performance and parallel workloads. This set is suitable for learning GPU fundamentals as well as reviewing concepts used in CUDA-based and accelerator-oriented systems.

List of GPU Architecture MCQs

Below is a collection of technically focused GPU Architecture MCQs with answers and explanations.

1. What is the primary purpose of a GPU's massively parallel architecture?

  1. To execute a small number of sequential tasks with very low latency
  2. To execute many parallel operations with high throughput
  3. To replace system storage
  4. To eliminate the need for CPU processing

Answer: B) To execute many parallel operations with high throughput

Explanation:

GPUs are designed to provide high throughput by executing large numbers of parallel operations. This architecture is particularly effective for workloads containing substantial data-level and thread-level parallelism.

2. In a modern NVIDIA GPU, what is an SM?

  1. System Memory
  2. Streaming Multiprocessor
  3. Storage Manager
  4. Shared Module

Answer: B) Streaming Multiprocessor

Explanation:

A Streaming Multiprocessor (SM) is a major execution unit in NVIDIA GPU architecture. It contains resources such as registers, shared memory, caches, scheduling logic, and functional units used to execute GPU workloads.

3. What execution model groups GPU threads into warps on NVIDIA GPUs?

  1. SIMD
  2. SIMT
  3. MIMD only
  4. VLIW

Answer: B) SIMT

Explanation:

SIMT stands for Single Instruction, Multiple Threads. NVIDIA GPUs organize threads into groups called warps and execute their instructions using the SIMT model.

4. How many threads are traditionally grouped into a warp on NVIDIA GPUs?

  1. 8
  2. 16
  3. 32
  4. 64

Answer: C) 32

Explanation:

NVIDIA CUDA groups threads into warps of 32 threads. Warp-level execution is an important architectural consideration for control-flow divergence and memory-access behavior.

5. Which component stores thread-local values that are allocated by the compiler on an NVIDIA SM?

  1. Register file
  2. Global memory
  3. Texture memory
  4. Host memory

Answer: A) Register file

Explanation:

The register file provides fast on-chip storage for values used by individual threads. Register allocation is generally determined by the compiler and contributes to the resources required by a resident thread block.

6. What is the main architectural advantage of on-chip memory compared with off-chip DRAM?

  1. It generally provides lower access latency
  2. It always has unlimited capacity
  3. It is physically located in system RAM
  4. It eliminates all synchronization requirements

Answer: A) It generally provides lower access latency

Explanation:

On-chip memories such as registers and shared memory are physically close to the execution units and generally provide much lower access latency than off-chip DRAM. Their capacity is limited, so they must be used carefully.

7. Which GPU memory is typically shared by threads within a thread block on NVIDIA CUDA architectures?

  1. Registers
  2. Shared memory
  3. Constant memory only
  4. CPU cache

Answer: B) Shared memory

Explanation:

Shared memory is an on-chip programmable memory space accessible by threads in the same thread block. It is commonly used for cooperation and data reuse among those threads.

8. What is the primary role of a GPU's global memory?

  1. Store data accessible by GPU threads across the device
  2. Store only CPU registers
  3. Store only instruction scheduling metadata
  4. Replace the GPU register file

Answer: A) Store data accessible by GPU threads across the device

Explanation:

GPU global memory is typically backed by device DRAM and is accessible by GPU threads. It provides large capacity but has substantially higher latency than on-chip storage.

9. Which memory hierarchy level is generally closest to an executing GPU thread?

  1. Device DRAM
  2. Registers
  3. PCIe memory
  4. Remote storage

Answer: B) Registers

Explanation:

Registers are located within the SM and are directly associated with executing threads. They provide very fast access but have limited capacity per SM.

10. What is the main purpose of an L2 cache in a GPU?

  1. Provide a shared cache accessible across GPU execution units
  2. Replace all registers
  3. Store only CPU instructions
  4. Generate GPU clocks

Answer: A) Provide a shared cache accessible across GPU execution units

Explanation:

In NVIDIA GPU architectures, the L2 cache is shared across SMs and can reduce the number of accesses that need to reach device DRAM. It is part of the GPU memory hierarchy between the SM-side caches and device memory.

11. What does memory bandwidth measure?

  1. The number of instructions executed per branch
  2. The rate at which data can be transferred through a memory interface
  3. The number of registers per thread
  4. The number of SMs per GPU

Answer: B) The rate at which data can be transferred through a memory interface

Explanation:

Memory bandwidth represents the amount of data that can be transferred per unit of time, commonly expressed in GB/s or TB/s. High bandwidth is particularly important for memory-bound workloads.

12. What is memory latency?

  1. The time required for a memory operation to begin producing its result
  2. The total number of memory channels
  3. The number of bytes stored in a register
  4. The number of active warps

Answer: A) The time required for a memory operation to begin producing its result

Explanation:

Memory latency is the delay associated with accessing data. GPUs use massive parallelism and scheduling of independent warps to help hide memory latency.

13. What technique allows a GPU to hide memory latency by executing other ready warps?

  1. Hardware multithreading
  2. Disk caching
  3. Instruction serialization
  4. CPU polling

Answer: A) Hardware multithreading

Explanation:

GPU SMs maintain execution state for many resident warps. When one warp is waiting on a long-latency operation, the scheduler can issue instructions from another ready warp.

14. What is warp divergence?

  1. When threads in a warp follow different control-flow paths
  2. When two GPUs share memory
  3. When all threads execute identical instructions
  4. When a kernel uses shared memory

Answer: A) When threads in a warp follow different control-flow paths

Explanation:

Warp divergence occurs when threads in the same warp take different branches. The GPU may need to execute the different paths separately with inactive lanes masked for each path, reducing execution efficiency.

15. Which workload is generally most suitable for a GPU?

  1. A highly sequential algorithm with unavoidable dependencies
  2. A workload containing thousands of independent operations
  3. A workload that performs one operation once per hour
  4. A workload that cannot be parallelized at all

Answer: B) A workload containing thousands of independent operations

Explanation:

GPUs are designed to exploit large-scale parallelism. Workloads such as matrix operations, image processing, simulations, and many machine-learning operations can contain substantial independent work.

16. What is a GPU execution unit responsible for?

  1. Performing computational operations such as arithmetic
  2. Managing only external storage
  3. Replacing system firmware
  4. Managing Ethernet packets exclusively

Answer: A) Performing computational operations such as arithmetic

Explanation:

GPU execution or functional units perform operations such as integer arithmetic, floating-point calculations, and specialized operations. Different GPU architectures provide different types and quantities of these units.

17. What is a floating-point execution unit primarily designed to process?

  1. Floating-point arithmetic
  2. Only disk requests
  3. Only network packets
  4. Only source-code parsing

Answer: A) Floating-point arithmetic

Explanation:

Floating-point execution units perform arithmetic operations using floating-point data types such as FP32 or FP64, depending on the architecture and workload.

18. What is the primary purpose of Tensor Cores in supported NVIDIA GPUs?

  1. Accelerate matrix and tensor operations
  2. Provide additional system RAM
  3. Replace the GPU scheduler
  4. Control display resolution

Answer: A) Accelerate matrix and tensor operations

Explanation:

Tensor Cores are specialized hardware designed to accelerate matrix and tensor operations used heavily in workloads such as deep learning. Their supported data types and capabilities vary by GPU generation.

19. Why can lower-precision arithmetic improve AI workload performance on suitable GPU hardware?

  1. It can enable higher arithmetic throughput and reduce data movement
  2. It always produces mathematically identical results
  3. It eliminates memory access completely
  4. It disables GPU parallelism

Answer: A) It can enable higher arithmetic throughput and reduce data movement

Explanation:

Lower-precision formats can require fewer bits per value, reducing storage and data movement. Specialized hardware may also provide substantially higher throughput for supported lower-precision operations.

20. What does occupancy generally describe on a GPU?

  1. The ratio of active warps to the maximum supported resident warps
  2. The percentage of disk space used
  3. The GPU's physical temperature only
  4. The number of installed drivers

Answer: A) The ratio of active warps to the maximum supported resident warps

Explanation:

Occupancy describes how many warps are active on an SM relative to the hardware maximum. Registers, shared memory, block limits, and architectural constraints can limit occupancy.

21. Which resource can directly limit the number of resident blocks on an SM?

  1. Per-block shared-memory usage
  2. Monitor resolution
  3. Keyboard buffer size
  4. Operating-system font size

Answer: A) Per-block shared-memory usage

Explanation:

Each SM has a finite amount of shared memory. If a kernel consumes a large amount per block, fewer blocks may be able to reside concurrently on the SM.

22. How can high register usage per thread affect GPU occupancy?

  1. It can reduce the number of resident warps or blocks
  2. It always increases occupancy
  3. It has no relationship to resource usage
  4. It disables the GPU cache

Answer: A) It can reduce the number of resident warps or blocks

Explanation:

Registers are a finite per-SM resource. A kernel requiring many registers per thread may consume enough register capacity to limit how many blocks or warps can remain resident.

23. What is coalesced memory access intended to improve?

  1. Efficiency of global memory transactions generated by threads
  2. Number of CPU cores
  3. GPU display resolution
  4. Disk compression ratio

Answer: A) Efficiency of global memory transactions generated by threads

Explanation:

When threads in a warp access memory in a suitable pattern, their requests can be combined into efficient memory transactions. This improves effective global-memory bandwidth.

24. Which access pattern is generally preferable for a warp processing consecutive array elements?

  1. Consecutive threads accessing consecutive elements
  2. Every thread accessing a random distant address
  3. Every thread accessing a different memory page
  4. Only one thread accessing all elements serially

Answer: A) Consecutive threads accessing consecutive elements

Explanation:

When neighboring threads access neighboring elements, the memory requests can often be serviced more efficiently. This is a common pattern for achieving good global-memory throughput.

25. What is a shared-memory bank conflict?

  1. A situation where multiple threads access the same shared-memory bank in a conflicting pattern
  2. A failure of GPU power delivery
  3. A conflict between two CUDA kernels on different GPUs
  4. A cache replacement algorithm

Answer: A) A situation where multiple threads access the same shared-memory bank in a conflicting pattern

Explanation:

Shared memory is divided into banks. Certain access patterns can cause multiple threads to contend for the same bank, potentially serializing accesses and reducing performance.

26. Why is shared memory useful for tiled matrix multiplication?

  1. It allows threads to reuse data from fast on-chip storage
  2. It increases CPU clock speed
  3. It eliminates arithmetic operations
  4. It stores unlimited matrices

Answer: A) It allows threads to reuse data from fast on-chip storage

Explanation:

Tiled matrix multiplication can load portions of matrices into shared memory and reuse those values across multiple calculations. This can reduce repeated accesses to slower global memory.

27. What is the main purpose of a GPU warp scheduler?

  1. Select a ready warp for instruction issue
  2. Allocate disk partitions
  3. Compile C++ source code
  4. Control the CPU operating system

Answer: A) Select a ready warp for instruction issue

Explanation:

A warp scheduler selects eligible warps and issues their instructions to the available execution resources. Scheduling helps keep GPU functional units busy while other warps wait on dependencies or memory operations.

28. Why can multiple resident warps improve GPU utilization?

  1. They provide independent work that can be scheduled when another warp is waiting
  2. They increase the capacity of GPU DRAM automatically
  3. They remove all data dependencies
  4. They eliminate instruction execution

Answer: A) They provide independent work that can be scheduled when another warp is waiting

Explanation:

Having multiple resident warps gives the scheduler more choices. This can help hide latency from memory operations and other dependencies.

29. What is instruction-level parallelism (ILP)?

  1. The ability to execute independent instructions from a computation concurrently or in an overlapped manner
  2. The ability to install multiple GPU drivers
  3. The number of GPU memory channels
  4. The number of graphics displays

Answer: A) The ability to execute independent instructions from a computation concurrently or in an overlapped manner

Explanation:

Instruction-level parallelism exists when instructions do not depend on one another and can be overlapped by the processor's execution pipeline. GPUs can exploit ILP in addition to thread-level parallelism.

30. What is thread-level parallelism in GPU architecture?

  1. Executing many threads concurrently to expose independent work
  2. Running one thread repeatedly on a CPU
  3. Increasing hard-drive capacity
  4. Executing only one instruction in the entire GPU

Answer: A) Executing many threads concurrently to expose independent work

Explanation:

Thread-level parallelism allows a GPU to maintain many threads in flight. The architecture can use this large amount of independent work to keep execution resources busy.

31. What does a GPU's compute capability identify in NVIDIA's architecture model?

  1. Supported architectural features and hardware parameters
  2. Internet bandwidth
  3. Operating-system version
  4. Monitor refresh rate

Answer: A) Supported architectural features and hardware parameters

Explanation:

NVIDIA compute capability is represented as a major and minor version and identifies supported GPU features and various hardware characteristics.

32. Which component provides large-capacity storage for data processed by GPU kernels?

  1. Device DRAM
  2. Register file
  3. Warp scheduler
  4. Instruction decoder only

Answer: A) Device DRAM

Explanation:

Device DRAM provides much larger storage capacity than on-chip memories such as registers and shared memory. It is commonly exposed to CUDA kernels as global memory.

33. Which statement about registers and shared memory is correct?

  1. Registers are generally thread-local, while shared memory is shared by threads in a block
  2. Both are stored exclusively in host RAM
  3. Shared memory is private to one thread
  4. Registers are shared by every GPU thread

Answer: A) Registers are generally thread-local, while shared memory is shared by threads in a block

Explanation:

Registers are allocated to individual threads, whereas shared memory is an explicitly managed on-chip memory space available to threads within a thread block.

34. What happens when a CUDA thread block is scheduled for execution on an SM?

  1. Its threads are resident on that SM while the block executes
  2. Its threads are permanently distributed across all GPUs
  3. Its threads execute only on the CPU
  4. Its threads are stored exclusively on disk

Answer: A) Its threads are resident on that SM while the block executes

Explanation:

In the CUDA programming model, threads belonging to a thread block execute on the same SM. Multiple blocks can be resident simultaneously when sufficient SM resources are available.

35. Why are GPU thread blocks designed to be independently schedulable?

  1. To allow grids to scale across different numbers of SMs
  2. To force all blocks onto one SM
  3. To prevent parallel execution
  4. To make CPU cache larger

Answer: A) To allow grids to scale across different numbers of SMs

Explanation:

CUDA allows blocks to be distributed among available SMs without requiring a fixed execution order between blocks. This makes the same kernel launch scalable across GPUs with different numbers of SMs.

36. What is a GPU cache designed to do?

  1. Reduce the need to repeatedly fetch data from slower memory levels
  2. Increase the number of CUDA threads directly
  3. Replace arithmetic execution units
  4. Store source code permanently

Answer: A) Reduce the need to repeatedly fetch data from slower memory levels

Explanation:

Caches exploit temporal and spatial locality by retaining recently accessed or nearby data. A cache hit can avoid a more expensive access to a lower memory level.

37. Which characteristic most directly indicates how many arithmetic operations a GPU can perform per second at a specified precision?

  1. Compute throughput
  2. Storage capacity
  3. Display resolution
  4. Memory address width only

Answer: A) Compute throughput

Explanation:

Compute throughput expresses the rate at which a processor can perform operations. GPU throughput specifications are usually dependent on operation type, precision, and architectural resources.

38. A GPU kernel performs very little arithmetic but reads a large amount of data from DRAM. What is it most likely to be?

  1. Memory-bound
  2. Compute-bound
  3. Branch-bound by definition
  4. Instruction-cache independent by definition

Answer: A) Memory-bound

Explanation:

A memory-bound kernel spends much of its execution time constrained by data movement rather than arithmetic throughput. Improving memory access patterns or data reuse can therefore have a significant impact.

39. A kernel performs a very large number of arithmetic operations while using relatively little memory bandwidth. What is it likely to be?

  1. Compute-bound
  2. Storage-bound
  3. Network-bound
  4. Display-bound

Answer: A) Compute-bound

Explanation:

A compute-bound workload is primarily limited by the available arithmetic or specialized compute throughput. Optimizing instruction count, arithmetic intensity, or execution-unit utilization may be important.

40. What does arithmetic intensity describe?

  1. Arithmetic operations performed per unit of data moved
  2. The number of registers in an SM
  3. The number of warps in a GPU
  4. The GPU clock frequency only

Answer: A) Arithmetic operations performed per unit of data moved

Explanation:

Arithmetic intensity relates computation to memory traffic, often expressed as operations per byte transferred. It is useful when analyzing whether a workload is likely to be constrained by computation or memory bandwidth.

41. What is the purpose of a GPU interconnect such as NVLink in a multi-GPU system?

  1. Provide high-speed communication between GPUs or other supported components
  2. Replace GPU execution units
  3. Convert GPU instructions into source code
  4. Increase monitor resolution

Answer: A) Provide high-speed communication between GPUs or other supported components

Explanation:

High-speed GPU interconnects can provide substantially faster communication paths than conventional system interconnects in supported systems. They are important for workloads that exchange significant amounts of data between GPUs.

42. What is PCIe commonly used for in a discrete GPU system?

  1. Connecting the GPU to the host system
  2. Connecting registers inside an SM
  3. Replacing shared memory
  4. Executing GPU warps

Answer: A) Connecting the GPU to the host system

Explanation:

PCI Express is commonly used as the host-to-device interconnect for discrete GPUs. Data transfers between CPU-side memory and GPU memory may travel over this interconnect depending on the system architecture.

43. Why can GPU clock frequency alone not determine application performance?

  1. Performance also depends on architecture, parallelism, memory behavior, instruction mix, and utilization
  2. Clock frequency has no relationship to hardware
  3. All GPUs execute exactly the same number of instructions per cycle
  4. Memory bandwidth is always irrelevant

Answer: A) Performance also depends on architecture, parallelism, memory behavior, instruction mix, and utilization

Explanation:

Two GPUs with similar clock frequencies can have very different numbers of execution units, memory bandwidth, cache structures, and supported instructions. Actual application performance therefore depends on the complete architecture and workload behavior.

44. What is the main purpose of GPU instruction pipelines?

  1. Allow different stages of instruction processing to overlap
  2. Store all GPU data permanently
  3. Eliminate all memory latency
  4. Increase disk capacity

Answer: A) Allow different stages of instruction processing to overlap

Explanation:

Pipelining divides instruction processing into stages so multiple instructions can be in different stages at the same time. This helps increase instruction throughput.

45. Which factor can cause a GPU's theoretical compute throughput to exceed its achievable application performance?

  1. Insufficient parallelism or memory bandwidth
  2. Using more than one thread
  3. Having an SM
  4. Using floating-point arithmetic

Answer: A) Insufficient parallelism or memory bandwidth

Explanation:

Theoretical peak throughput assumes favorable conditions and sufficient work. Real applications may be limited by memory bandwidth, latency, instruction dependencies, synchronization, divergence, or insufficient utilization.

46. Which architectural property is particularly important when a GPU workload contains many independent threads?

  1. Massive thread-level parallelism
  2. Large sequential instruction chains only
  3. Single-thread optimization only
  4. Disk seek time

Answer: A) Massive thread-level parallelism

Explanation:

GPU architectures are designed to maintain and execute many threads concurrently. Large amounts of independent work allow the architecture to utilize its parallel execution resources effectively.

47. What is the primary architectural consequence of using a very large amount of shared memory per thread block?

  1. Fewer blocks may be able to reside simultaneously on an SM
  2. The GPU automatically gains more DRAM
  3. The number of GPU cores doubles
  4. Warp size increases automatically

Answer: A) Fewer blocks may be able to reside simultaneously on an SM

Explanation:

Shared memory is a finite SM resource. High per-block shared-memory usage can reduce the number of blocks that can be resident simultaneously, which can affect occupancy and latency hiding.

48. A kernel has high occupancy but performs poorly because every warp repeatedly accesses uncached global memory. What is the most relevant architectural issue?

  1. Memory access behavior and insufficient data locality
  2. Insufficient number of CPU cores
  3. GPU display resolution
  4. Excessive filesystem capacity

Answer: A) Memory access behavior and insufficient data locality

Explanation:

High occupancy does not guarantee high performance. If the kernel generates inefficient global-memory accesses and has little data reuse, memory latency and bandwidth can still dominate execution time.

49. A CUDA kernel uses 96 registers per thread and a large amount of shared memory per block. What should an architect or performance engineer investigate first?

  1. Whether register and shared-memory usage are limiting resident warps or blocks
  2. Whether the monitor supports a higher refresh rate
  3. Whether the CPU has enough USB ports
  4. Whether the source code uses enough comments

Answer: A) Whether register and shared-memory usage are limiting resident warps or blocks

Explanation:

Registers and shared memory are finite per-SM resources. High usage of both can restrict the number of resident blocks and warps, potentially reducing the GPU's ability to hide latency. Resource usage and occupancy should therefore be analyzed together with actual performance metrics.

50. A matrix-multiplication kernel has excellent global-memory coalescing but spends most of its execution time repeatedly loading the same matrix tiles from DRAM. Which architectural optimization is most directly relevant?

  1. Stage reusable tiles in shared memory to increase on-chip data reuse
  2. Increase the number of CPU threads
  3. Disable all GPU caches
  4. Serialize all GPU threads

Answer: A) Stage reusable tiles in shared memory to increase on-chip data reuse

Explanation:

Even with coalesced accesses, repeatedly fetching the same data from DRAM can create substantial memory traffic. Tiling allows threads in a block to load data into shared memory and reuse it for multiple matrix operations, reducing repeated global-memory accesses and improving arithmetic intensity.

Comments and Discussions!

Load comments ↻



Copyright © 2026 www.includehelp.com. All rights reserved.