NVIDIA CUDA MCQs (Multiple-Choice Questions)

Practice NVIDIA CUDA MCQs to test your knowledge of CUDA programming, GPU kernels, thread hierarchy, memory management, CUDA streams, synchronization, unified memory, and GPU performance optimization.

NVIDIA CUDA MCQs

These NVIDIA CUDA multiple-choice questions cover fundamental and advanced concepts used to develop and optimize applications for NVIDIA GPUs.

List of NVIDIA CUDA MCQs

Below is the list of NVIDIA CUDA MCQs with answers and explanations.

1. What is NVIDIA CUDA?

  1. A parallel computing platform and programming model for NVIDIA GPUs
  2. A relational database system
  3. An operating system
  4. A web development framework

Answer: A) A parallel computing platform and programming model for NVIDIA GPUs

Explanation:

CUDA provides a programming model, APIs, libraries, and tools that allow developers to execute general-purpose computations on NVIDIA GPUs.

2. In CUDA terminology, what is the CPU generally called?

  1. Device
  2. Host
  3. Kernel
  4. Warp

Answer: B) Host

Explanation:

In the CUDA programming model, the CPU and its associated memory are referred to as the host, while the GPU and its memory are referred to as the device.

3. What is CUDA device code?

  1. Code executed on the GPU
  2. Code executed only by the operating system
  3. Code stored exclusively in system memory
  4. Code used only for network communication

Answer: A) Code executed on the GPU

Explanation:

CUDA applications can contain host code running on the CPU and device code running on the GPU.

4. What is a CUDA kernel?

  1. A function launched for execution on the GPU
  2. A CPU cache
  3. A database table
  4. A GPU driver file

Answer: A) A function launched for execution on the GPU

Explanation:

A CUDA kernel is a function executed by many GPU threads in parallel after it is launched by host code.

5. Which CUDA function qualifier identifies a function that can be launched as a GPU kernel?

  1. __global__
  2. __device__
  3. __host__
  4. __shared__

Answer: A) __global__

Explanation:

The __global__ qualifier identifies a CUDA kernel function that is launched from host code and executed on the device.

6. What is the basic execution hierarchy of CUDA threads?

  1. Threads → Blocks → Grid
  2. Threads → Grid → Blocks
  3. Blocks → Threads → Grid
  4. Warps → Grids → CPUs

Answer: A) Threads → Blocks → Grid

Explanation:

CUDA organizes threads into thread blocks, and thread blocks into a grid for a kernel launch.

7. Which built-in CUDA variable identifies the index of the current thread within its block?

  1. gridIdx
  2. threadIdx
  3. blockSize
  4. kernelIdx

Answer: B) threadIdx

Explanation:

threadIdx provides the position of the current thread within its thread block.

8. Which CUDA variable identifies the index of the current thread block within a grid?

  1. threadIdx
  2. blockIdx
  3. gridIdx
  4. warpIdx

Answer: B) blockIdx

Explanation:

blockIdx identifies the current thread block's position within the grid.

9. What does blockDim represent?

  1. The dimensions of the current thread block
  2. The dimensions of GPU memory
  3. The number of SMs in the GPU
  4. The dimensions of the CPU cache

Answer: A) The dimensions of the current thread block

Explanation:

blockDim provides the number of threads in each dimension of the current thread block.

10. What does gridDim represent?

  1. The dimensions of the grid launched for a kernel
  2. The dimensions of shared memory
  3. The number of CPU cores
  4. The size of the GPU register file

Answer: A) The dimensions of the grid launched for a kernel

Explanation:

gridDim contains the dimensions of the grid of thread blocks used for the kernel launch.

11. How many dimensions can a CUDA thread block have?

  1. Only one
  2. Only two
  3. One, two, or three
  4. Exactly four

Answer: C) One, two, or three

Explanation:

CUDA thread blocks can be organized in one, two, or three dimensions.

12. How many dimensions can a CUDA grid have?

  1. Only one
  2. One, two, or three
  3. Exactly four
  4. Only two

Answer: B) One, two, or three

Explanation:

CUDA grids can be one-, two-, or three-dimensional, allowing thread organization to match different data structures.

13. What is a warp in CUDA?

  1. A group of 32 threads
  2. A group of 8 blocks
  3. A group of 64 SMs
  4. A group of 16 kernels

Answer: A) A group of 32 threads

Explanation:

CUDA organizes threads within a block into warps of 32 threads. The warp is a fundamental execution unit in the CUDA SIMT model.

14. What does SIMT stand for?

  1. Single Instruction, Multiple Threads
  2. Shared Instruction, Multiple Tasks
  3. Single Interface, Multiple Transfers
  4. System Instruction, Memory Threads

Answer: A) Single Instruction, Multiple Threads

Explanation:

SIMT is the execution model used by CUDA GPUs in which threads are grouped into warps and execute instructions together while retaining individual thread state.

15. What is a Streaming Multiprocessor (SM)?

  1. A GPU hardware unit that executes CUDA thread blocks
  2. A CPU-only memory controller
  3. A host-side compiler
  4. A storage device

Answer: A) A GPU hardware unit that executes CUDA thread blocks

Explanation:

Streaming Multiprocessors contain processing resources, registers, and on-chip memory resources used to execute CUDA workloads.

16. Where are all threads belonging to a CUDA thread block executed?

  1. On a single SM
  2. Only on the CPU
  3. Across unrelated GPUs
  4. Only in system RAM

Answer: A) On a single SM

Explanation:

CUDA schedules all threads of a thread block on a single SM, enabling them to communicate through shared memory and synchronization mechanisms.

17. Which memory space is accessible by all threads in a CUDA grid?

  1. Global memory
  2. Register memory
  3. Local memory only
  4. Thread-private memory

Answer: A) Global memory

Explanation:

Global memory is accessible by threads across the grid and is commonly used for input and output arrays.

18. Which CUDA memory space is shared by all threads in a thread block?

  1. Shared memory
  2. Register memory
  3. Constant memory only
  4. Local memory

Answer: A) Shared memory

Explanation:

Shared memory is accessible to all threads within a thread block and provides relatively low-latency on-chip storage for cooperative computations.

19. What is the scope of CUDA registers?

  1. Thread
  2. Grid
  3. Block
  4. Entire application only

Answer: A) Thread

Explanation:

Registers are private to individual CUDA threads and are located on the SM.

20. What is local memory in CUDA?

  1. Memory associated with an individual thread's local variables when they cannot reside in registers
  2. Memory shared by every block
  3. Host RAM only
  4. Constant memory shared by the entire grid

Answer: A) Memory associated with an individual thread's local variables when they cannot reside in registers

Explanation:

Local memory has thread scope. Despite its name, it resides in device memory rather than being a small dedicated on-chip memory area for each thread.

21. Which CUDA API is commonly used to allocate device global memory?

  1. cudaMalloc()
  2. cudaLaunch()
  3. cudaThread()
  4. cudaGrid()

Answer: A) cudaMalloc()

Explanation:

cudaMalloc() allocates memory in the device's global memory space.

22. Which CUDA API releases memory allocated with cudaMalloc()?

  1. cudaFree()
  2. cudaRelease()
  3. cudaDelete()
  4. cudaDestroyMemory()

Answer: A) cudaFree()

Explanation:

cudaFree() releases memory previously allocated through compatible CUDA memory allocation APIs.

23. What is the purpose of cudaMemcpy()?

  1. Copying data between CUDA-accessible memory locations
  2. Launching a kernel
  3. Creating a thread block
  4. Compiling CUDA source code

Answer: A) Copying data between CUDA-accessible memory locations

Explanation:

cudaMemcpy() is used to copy data between host and device memory or between compatible device memory locations.

24. What is Unified Memory in CUDA?

  1. A memory model that allows CPU and GPU code to access managed allocations
  2. A GPU-only register file
  3. A compiler optimization only
  4. A type of CUDA kernel

Answer: A) A memory model that allows CPU and GPU code to access managed allocations

Explanation:

CUDA Unified Memory allows managed allocations to be accessed from CPU or GPU code, with the CUDA runtime and underlying system handling memory placement and migration as appropriate.

25. Which CUDA API is commonly used to allocate managed memory?

  1. cudaMallocManaged()
  2. cudaManagedAlloc()
  3. cudaUnifiedAlloc()
  4. cudaSharedMalloc()

Answer: A) cudaMallocManaged()

Explanation:

cudaMallocManaged() allocates managed memory that can be accessed by CPU and GPU code.

26. What is a CUDA stream?

  1. A sequence of operations that execute in issue order on the GPU
  2. A type of GPU memory
  3. A collection of thread blocks stored permanently on an SM
  4. A CPU register

Answer: A) A sequence of operations that execute in issue order on the GPU

Explanation:

CUDA streams provide an execution context for ordering asynchronous operations such as kernel launches and memory operations.

27. Why are multiple CUDA streams useful?

  1. They can enable independent operations to overlap when hardware and dependencies permit
  2. They increase the number of CPU cores
  3. They eliminate GPU memory
  4. They force all operations to execute sequentially

Answer: A) They can enable independent operations to overlap when hardware and dependencies permit

Explanation:

Using multiple streams can allow independent kernels or memory operations to overlap, potentially improving hardware utilization.

28. What is a CUDA event commonly used for?

  1. Recording and measuring progress or timing of operations on a stream
  2. Allocating global memory
  3. Defining thread dimensions
  4. Compiling kernels

Answer: A) Recording and measuring progress or timing of operations on a stream

Explanation:

CUDA events can be recorded in streams and used for synchronization and GPU-side timing measurements.

29. What does cudaDeviceSynchronize() do?

  1. Waits until previously issued device work has completed
  2. Allocates GPU memory
  3. Starts a new kernel
  4. Changes the number of threads in a block

Answer: A) Waits until previously issued device work has completed

Explanation:

cudaDeviceSynchronize() blocks the calling host thread until previously issued device work has completed.

30. What is synchronization used for in CUDA?

  1. Coordinating execution or memory visibility among threads or operations
  2. Increasing GPU memory capacity
  3. Changing a CPU into a GPU
  4. Compiling C++ source code

Answer: A) Coordinating execution or memory visibility among threads or operations

Explanation:

Synchronization mechanisms help coordinate threads and operations when one computation depends on the completion or visibility of another.

31. What does __syncthreads() provide?

  1. A barrier for threads within a thread block
  2. A barrier across all GPUs in a cluster
  3. A memory allocation mechanism
  4. A kernel compilation command

Answer: A) A barrier for threads within a thread block

Explanation:

__syncthreads() synchronizes threads within the same thread block. It should be used carefully so that participating threads reach the barrier correctly.

32. What is coalesced global memory access?

  1. A memory access pattern where threads in a warp access memory in a way that can be serviced efficiently
  2. A method of allocating shared memory
  3. A method of launching multiple grids
  4. A CPU-only caching technique

Answer: A) A memory access pattern where threads in a warp access memory in a way that can be serviced efficiently

Explanation:

Coalesced accesses allow memory requests from threads in a warp to be combined efficiently, improving global memory throughput.

33. What is warp divergence?

  1. When threads in the same warp follow different control-flow paths
  2. When GPU memory becomes completely full
  3. When multiple GPUs share the same PCIe slot
  4. When a kernel fails to compile

Answer: A) When threads in the same warp follow different control-flow paths

Explanation:

When threads in a warp take different branches, the hardware may execute the divergent paths separately, reducing parallel efficiency. NVIDIA recommends minimizing unnecessary warp divergence for performance.

34. What is occupancy in CUDA?

  1. The ratio of active warps on an SM to the maximum supported active warps
  2. The percentage of GPU memory that is permanently allocated
  3. The number of kernels stored on disk
  4. The number of CUDA streams in an application

Answer: A) The ratio of active warps on an SM to the maximum supported active warps

Explanation:

Occupancy describes how many active warps are resident on an SM relative to its hardware limit. Higher occupancy can help hide latency, although maximum occupancy does not automatically mean maximum performance.

35. Which resource can limit the number of active blocks on an SM?

  1. Registers
  2. Shared memory
  3. Maximum resident threads
  4. All of the above

Answer: D) All of the above

Explanation:

SM resources such as registers, shared memory, and limits on resident threads can restrict how many blocks can be active simultaneously.

36. Why can excessive register usage reduce kernel occupancy?

  1. Each thread consumes registers, limiting how many threads can be resident on an SM
  2. Registers are stored in system RAM only
  3. Registers disable shared memory
  4. Registers increase the number of GPUs automatically

Answer: A) Each thread consumes registers, limiting how many threads can be resident on an SM

Explanation:

Registers are a finite SM resource. A kernel requiring many registers per thread may reduce the number of simultaneously resident threads or blocks.

37. What is the purpose of shared memory in a CUDA kernel?

  1. To provide fast, programmer-managed storage shared by threads in a block
  2. To permanently store application files
  3. To replace all global memory
  4. To store CPU operating-system files

Answer: A) To provide fast, programmer-managed storage shared by threads in a block

Explanation:

Shared memory is useful when threads in a block need to cooperate on data or repeatedly reuse values with lower latency than global memory.

38. What is a shared-memory bank conflict?

  1. A situation where multiple threads access different addresses that map to the same memory bank in a conflicting pattern
  2. A failure of GPU power delivery
  3. A conflict between CUDA versions
  4. A conflict between CPU cores

Answer: A) A situation where multiple threads access different addresses that map to the same memory bank in a conflicting pattern

Explanation:

Bank conflicts can serialize shared-memory accesses and reduce performance. Appropriate data layout and access patterns can help avoid them.

39. What is constant memory in CUDA?

  1. A read-only memory space intended for data that remains constant during kernel execution
  2. A memory space that can only store thread-local variables
  3. A type of system RAM
  4. A memory space used exclusively for kernel output

Answer: A) A read-only memory space intended for data that remains constant during kernel execution

Explanation:

Constant memory is a device memory space intended for values that are constant for the lifetime of a kernel and can benefit from caching.

40. What is pinned host memory?

  1. Page-locked host memory that can support efficient host-device data transfers
  2. GPU shared memory
  3. Constant device memory
  4. Thread-local register memory

Answer: A) Page-locked host memory that can support efficient host-device data transfers

Explanation:

Pinned or page-locked host memory is not pageable by the operating system and can be useful for asynchronous host-device transfers.

41. Which API is commonly used to allocate pinned host memory?

  1. cudaMallocHost()
  2. cudaPinnedDevice()
  3. cudaHostGPU()
  4. cudaPageLock()

Answer: A) cudaMallocHost()

Explanation:

cudaMallocHost() allocates page-locked host memory that can be used with CUDA data-transfer operations.

42. What is the main purpose of cudaMemcpyAsync()?

  1. To initiate an asynchronous memory copy associated with a CUDA stream
  2. To compile a kernel asynchronously
  3. To allocate shared memory
  4. To create a new GPU device

Answer: A) To initiate an asynchronous memory copy associated with a CUDA stream

Explanation:

cudaMemcpyAsync() can initiate asynchronous memory transfers, allowing suitable operations to overlap when the hardware, memory type, stream, and dependencies permit.

43. What is CUDA Graphs designed to optimize?

  1. Repeated execution of a sequence of GPU operations by representing them as a graph
  2. GPU physical cooling
  3. Database indexing
  4. CUDA source-code formatting

Answer: A) Repeated execution of a sequence of GPU operations by representing them as a graph

Explanation:

CUDA Graphs allow a sequence of operations to be captured or constructed as a graph and launched with reduced CPU-side launch overhead in suitable workloads.

44. What is the purpose of CUDA error handling APIs?

  1. To detect and report errors from CUDA API calls and asynchronous GPU execution
  2. To increase GPU memory capacity
  3. To change thread-block dimensions automatically
  4. To train neural networks

Answer: A) To detect and report errors from CUDA API calls and asynchronous GPU execution

Explanation:

CUDA provides error-reporting mechanisms such as cudaGetLastError() and cudaPeekAtLastError() for diagnosing runtime and kernel-launch errors.

45. What does the CUDA compiler driver nvcc primarily do?

  1. Compiles CUDA source code and coordinates host and device compilation
  2. Allocates GPU memory at runtime
  3. Schedules CUDA threads directly
  4. Measures GPU temperature only

Answer: A) Compiles CUDA source code and coordinates host and device compilation

Explanation:

nvcc is NVIDIA's CUDA compiler driver and handles compilation of CUDA code for host and device components.

46. A CUDA vector-add kernel uses threadIdx.x + blockIdx.x * blockDim.x. What does this expression typically calculate?

  1. The global one-dimensional thread index
  2. The GPU device ID
  3. The number of SMs
  4. The shared-memory bank number

Answer: A) The global one-dimensional thread index

Explanation:

The expression combines the thread's local index with its block index and block size to calculate a unique one-dimensional index across the grid.

47. A CUDA kernel processes 10,000 array elements using blocks of 256 threads. What should the kernel normally do for threads whose calculated index is greater than or equal to 10,000?

  1. Skip the out-of-range element using a bounds check
  2. Access a random array element
  3. Write beyond the allocated array
  4. Terminate the entire GPU

Answer: A) Skip the out-of-range element using a bounds check

Explanation:

When the grid contains more threads than the data size, kernels commonly use a condition such as if (idx < n) before accessing the array.

48. A CUDA kernel is slow because each thread repeatedly reads the same data needed by neighboring threads. Which CUDA feature can often reduce redundant global-memory accesses?

  1. Shared memory
  2. Host-only memory
  3. Constant CPU registers
  4. Kernel comments

Answer: A) Shared memory

Explanation:

Threads within a block can cooperatively load reusable data into shared memory, allowing neighboring threads to reuse values rather than repeatedly accessing global memory.

49. A CUDA application launches many very small kernels repeatedly, and GPU computation is fast but CPU-side launch overhead becomes significant. Which CUDA feature can help reduce repeated launch overhead?

  1. CUDA Graphs
  2. Constant memory
  3. Thread-local memory
  4. Increasing the number of source-code comments

Answer: A) CUDA Graphs

Explanation:

CUDA Graphs can represent repeated sequences of GPU operations and replay them with lower CPU launch overhead than issuing every operation independently in suitable workloads.

50. A CUDA matrix-processing application has poor performance. Profiling shows uncoalesced global-memory accesses, significant warp divergence, excessive register usage, and low occupancy. Which optimization strategy is most appropriate?

  1. Optimize memory-access patterns, reduce unnecessary divergence, control register usage, and tune the kernel configuration
  2. Increase the number of CPU threads without changing the CUDA kernel
  3. Move every array to host memory
  4. Increase kernel complexity to use more registers

Answer: A) Optimize memory-access patterns, reduce unnecessary divergence, control register usage, and tune the kernel configuration

Explanation:

CUDA performance depends on several interacting factors. Coalesced memory access improves global-memory efficiency, reducing warp divergence can improve execution efficiency, and controlling register usage can increase the number of active warps or blocks that fit on an SM. Kernel configuration should be tuned based on profiling rather than relying on a single optimization metric.

Comments and Discussions!

Load comments ↻



Copyright © 2026 www.includehelp.com. All rights reserved.