Home »
Trending Technologies MCQs
NVIDIA CUDA MCQs (Multiple-Choice Questions)
Practice NVIDIA CUDA MCQs to test your knowledge of CUDA programming, GPU kernels, thread hierarchy, memory management, CUDA streams, synchronization, unified memory, and GPU performance optimization.
NVIDIA CUDA MCQs
These NVIDIA CUDA multiple-choice questions cover fundamental and advanced concepts used to develop and optimize applications for NVIDIA GPUs.
List of NVIDIA CUDA MCQs
Below is the list of NVIDIA CUDA MCQs with answers and explanations.
1. What is NVIDIA CUDA?
- A parallel computing platform and programming model for NVIDIA GPUs
- A relational database system
- An operating system
- A web development framework
Answer: A) A parallel computing platform and programming model for NVIDIA GPUs
Explanation:
CUDA provides a programming model, APIs, libraries, and tools that allow developers to execute general-purpose computations on NVIDIA GPUs.
2. In CUDA terminology, what is the CPU generally called?
- Device
- Host
- Kernel
- Warp
Answer: B) Host
Explanation:
In the CUDA programming model, the CPU and its associated memory are referred to as the host, while the GPU and its memory are referred to as the device.
3. What is CUDA device code?
- Code executed on the GPU
- Code executed only by the operating system
- Code stored exclusively in system memory
- Code used only for network communication
Answer: A) Code executed on the GPU
Explanation:
CUDA applications can contain host code running on the CPU and device code running on the GPU.
4. What is a CUDA kernel?
- A function launched for execution on the GPU
- A CPU cache
- A database table
- A GPU driver file
Answer: A) A function launched for execution on the GPU
Explanation:
A CUDA kernel is a function executed by many GPU threads in parallel after it is launched by host code.
5. Which CUDA function qualifier identifies a function that can be launched as a GPU kernel?
__global__
__device__
__host__
__shared__
Answer: A) __global__
Explanation:
The __global__ qualifier identifies a CUDA kernel function that is launched from host code and executed on the device.
6. What is the basic execution hierarchy of CUDA threads?
- Threads → Blocks → Grid
- Threads → Grid → Blocks
- Blocks → Threads → Grid
- Warps → Grids → CPUs
Answer: A) Threads → Blocks → Grid
Explanation:
CUDA organizes threads into thread blocks, and thread blocks into a grid for a kernel launch.
7. Which built-in CUDA variable identifies the index of the current thread within its block?
gridIdx
threadIdx
blockSize
kernelIdx
Answer: B) threadIdx
Explanation:
threadIdx provides the position of the current thread within its thread block.
8. Which CUDA variable identifies the index of the current thread block within a grid?
threadIdx
blockIdx
gridIdx
warpIdx
Answer: B) blockIdx
Explanation:
blockIdx identifies the current thread block's position within the grid.
9. What does blockDim represent?
- The dimensions of the current thread block
- The dimensions of GPU memory
- The number of SMs in the GPU
- The dimensions of the CPU cache
Answer: A) The dimensions of the current thread block
Explanation:
blockDim provides the number of threads in each dimension of the current thread block.
10. What does gridDim represent?
- The dimensions of the grid launched for a kernel
- The dimensions of shared memory
- The number of CPU cores
- The size of the GPU register file
Answer: A) The dimensions of the grid launched for a kernel
Explanation:
gridDim contains the dimensions of the grid of thread blocks used for the kernel launch.
11. How many dimensions can a CUDA thread block have?
- Only one
- Only two
- One, two, or three
- Exactly four
Answer: C) One, two, or three
Explanation:
CUDA thread blocks can be organized in one, two, or three dimensions.
12. How many dimensions can a CUDA grid have?
- Only one
- One, two, or three
- Exactly four
- Only two
Answer: B) One, two, or three
Explanation:
CUDA grids can be one-, two-, or three-dimensional, allowing thread organization to match different data structures.
13. What is a warp in CUDA?
- A group of 32 threads
- A group of 8 blocks
- A group of 64 SMs
- A group of 16 kernels
Answer: A) A group of 32 threads
Explanation:
CUDA organizes threads within a block into warps of 32 threads. The warp is a fundamental execution unit in the CUDA SIMT model.
14. What does SIMT stand for?
- Single Instruction, Multiple Threads
- Shared Instruction, Multiple Tasks
- Single Interface, Multiple Transfers
- System Instruction, Memory Threads
Answer: A) Single Instruction, Multiple Threads
Explanation:
SIMT is the execution model used by CUDA GPUs in which threads are grouped into warps and execute instructions together while retaining individual thread state.
15. What is a Streaming Multiprocessor (SM)?
- A GPU hardware unit that executes CUDA thread blocks
- A CPU-only memory controller
- A host-side compiler
- A storage device
Answer: A) A GPU hardware unit that executes CUDA thread blocks
Explanation:
Streaming Multiprocessors contain processing resources, registers, and on-chip memory resources used to execute CUDA workloads.
16. Where are all threads belonging to a CUDA thread block executed?
- On a single SM
- Only on the CPU
- Across unrelated GPUs
- Only in system RAM
Answer: A) On a single SM
Explanation:
CUDA schedules all threads of a thread block on a single SM, enabling them to communicate through shared memory and synchronization mechanisms.
17. Which memory space is accessible by all threads in a CUDA grid?
- Global memory
- Register memory
- Local memory only
- Thread-private memory
Answer: A) Global memory
Explanation:
Global memory is accessible by threads across the grid and is commonly used for input and output arrays.
18. Which CUDA memory space is shared by all threads in a thread block?
- Shared memory
- Register memory
- Constant memory only
- Local memory
Answer: A) Shared memory
Explanation:
Shared memory is accessible to all threads within a thread block and provides relatively low-latency on-chip storage for cooperative computations.
19. What is the scope of CUDA registers?
- Thread
- Grid
- Block
- Entire application only
Answer: A) Thread
Explanation:
Registers are private to individual CUDA threads and are located on the SM.
20. What is local memory in CUDA?
- Memory associated with an individual thread's local variables when they cannot reside in registers
- Memory shared by every block
- Host RAM only
- Constant memory shared by the entire grid
Answer: A) Memory associated with an individual thread's local variables when they cannot reside in registers
Explanation:
Local memory has thread scope. Despite its name, it resides in device memory rather than being a small dedicated on-chip memory area for each thread.
21. Which CUDA API is commonly used to allocate device global memory?
cudaMalloc()
cudaLaunch()
cudaThread()
cudaGrid()
Answer: A) cudaMalloc()
Explanation:
cudaMalloc() allocates memory in the device's global memory space.
22. Which CUDA API releases memory allocated with cudaMalloc()?
cudaFree()
cudaRelease()
cudaDelete()
cudaDestroyMemory()
Answer: A) cudaFree()
Explanation:
cudaFree() releases memory previously allocated through compatible CUDA memory allocation APIs.
23. What is the purpose of cudaMemcpy()?
- Copying data between CUDA-accessible memory locations
- Launching a kernel
- Creating a thread block
- Compiling CUDA source code
Answer: A) Copying data between CUDA-accessible memory locations
Explanation:
cudaMemcpy() is used to copy data between host and device memory or between compatible device memory locations.
24. What is Unified Memory in CUDA?
- A memory model that allows CPU and GPU code to access managed allocations
- A GPU-only register file
- A compiler optimization only
- A type of CUDA kernel
Answer: A) A memory model that allows CPU and GPU code to access managed allocations
Explanation:
CUDA Unified Memory allows managed allocations to be accessed from CPU or GPU code, with the CUDA runtime and underlying system handling memory placement and migration as appropriate.
25. Which CUDA API is commonly used to allocate managed memory?
cudaMallocManaged()
cudaManagedAlloc()
cudaUnifiedAlloc()
cudaSharedMalloc()
Answer: A) cudaMallocManaged()
Explanation:
cudaMallocManaged() allocates managed memory that can be accessed by CPU and GPU code.
26. What is a CUDA stream?
- A sequence of operations that execute in issue order on the GPU
- A type of GPU memory
- A collection of thread blocks stored permanently on an SM
- A CPU register
Answer: A) A sequence of operations that execute in issue order on the GPU
Explanation:
CUDA streams provide an execution context for ordering asynchronous operations such as kernel launches and memory operations.
27. Why are multiple CUDA streams useful?
- They can enable independent operations to overlap when hardware and dependencies permit
- They increase the number of CPU cores
- They eliminate GPU memory
- They force all operations to execute sequentially
Answer: A) They can enable independent operations to overlap when hardware and dependencies permit
Explanation:
Using multiple streams can allow independent kernels or memory operations to overlap, potentially improving hardware utilization.
28. What is a CUDA event commonly used for?
- Recording and measuring progress or timing of operations on a stream
- Allocating global memory
- Defining thread dimensions
- Compiling kernels
Answer: A) Recording and measuring progress or timing of operations on a stream
Explanation:
CUDA events can be recorded in streams and used for synchronization and GPU-side timing measurements.
29. What does cudaDeviceSynchronize() do?
- Waits until previously issued device work has completed
- Allocates GPU memory
- Starts a new kernel
- Changes the number of threads in a block
Answer: A) Waits until previously issued device work has completed
Explanation:
cudaDeviceSynchronize() blocks the calling host thread until previously issued device work has completed.
30. What is synchronization used for in CUDA?
- Coordinating execution or memory visibility among threads or operations
- Increasing GPU memory capacity
- Changing a CPU into a GPU
- Compiling C++ source code
Answer: A) Coordinating execution or memory visibility among threads or operations
Explanation:
Synchronization mechanisms help coordinate threads and operations when one computation depends on the completion or visibility of another.
31. What does __syncthreads() provide?
- A barrier for threads within a thread block
- A barrier across all GPUs in a cluster
- A memory allocation mechanism
- A kernel compilation command
Answer: A) A barrier for threads within a thread block
Explanation:
__syncthreads() synchronizes threads within the same thread block. It should be used carefully so that participating threads reach the barrier correctly.
32. What is coalesced global memory access?
- A memory access pattern where threads in a warp access memory in a way that can be serviced efficiently
- A method of allocating shared memory
- A method of launching multiple grids
- A CPU-only caching technique
Answer: A) A memory access pattern where threads in a warp access memory in a way that can be serviced efficiently
Explanation:
Coalesced accesses allow memory requests from threads in a warp to be combined efficiently, improving global memory throughput.
33. What is warp divergence?
- When threads in the same warp follow different control-flow paths
- When GPU memory becomes completely full
- When multiple GPUs share the same PCIe slot
- When a kernel fails to compile
Answer: A) When threads in the same warp follow different control-flow paths
Explanation:
When threads in a warp take different branches, the hardware may execute the divergent paths separately, reducing parallel efficiency. NVIDIA recommends minimizing unnecessary warp divergence for performance.
34. What is occupancy in CUDA?
- The ratio of active warps on an SM to the maximum supported active warps
- The percentage of GPU memory that is permanently allocated
- The number of kernels stored on disk
- The number of CUDA streams in an application
Answer: A) The ratio of active warps on an SM to the maximum supported active warps
Explanation:
Occupancy describes how many active warps are resident on an SM relative to its hardware limit. Higher occupancy can help hide latency, although maximum occupancy does not automatically mean maximum performance.
35. Which resource can limit the number of active blocks on an SM?
- Registers
- Shared memory
- Maximum resident threads
- All of the above
Answer: D) All of the above
Explanation:
SM resources such as registers, shared memory, and limits on resident threads can restrict how many blocks can be active simultaneously.
36. Why can excessive register usage reduce kernel occupancy?
- Each thread consumes registers, limiting how many threads can be resident on an SM
- Registers are stored in system RAM only
- Registers disable shared memory
- Registers increase the number of GPUs automatically
Answer: A) Each thread consumes registers, limiting how many threads can be resident on an SM
Explanation:
Registers are a finite SM resource. A kernel requiring many registers per thread may reduce the number of simultaneously resident threads or blocks.
37. What is the purpose of shared memory in a CUDA kernel?
- To provide fast, programmer-managed storage shared by threads in a block
- To permanently store application files
- To replace all global memory
- To store CPU operating-system files
Answer: A) To provide fast, programmer-managed storage shared by threads in a block
Explanation:
Shared memory is useful when threads in a block need to cooperate on data or repeatedly reuse values with lower latency than global memory.
38. What is a shared-memory bank conflict?
- A situation where multiple threads access different addresses that map to the same memory bank in a conflicting pattern
- A failure of GPU power delivery
- A conflict between CUDA versions
- A conflict between CPU cores
Answer: A) A situation where multiple threads access different addresses that map to the same memory bank in a conflicting pattern
Explanation:
Bank conflicts can serialize shared-memory accesses and reduce performance. Appropriate data layout and access patterns can help avoid them.
39. What is constant memory in CUDA?
- A read-only memory space intended for data that remains constant during kernel execution
- A memory space that can only store thread-local variables
- A type of system RAM
- A memory space used exclusively for kernel output
Answer: A) A read-only memory space intended for data that remains constant during kernel execution
Explanation:
Constant memory is a device memory space intended for values that are constant for the lifetime of a kernel and can benefit from caching.
40. What is pinned host memory?
- Page-locked host memory that can support efficient host-device data transfers
- GPU shared memory
- Constant device memory
- Thread-local register memory
Answer: A) Page-locked host memory that can support efficient host-device data transfers
Explanation:
Pinned or page-locked host memory is not pageable by the operating system and can be useful for asynchronous host-device transfers.
41. Which API is commonly used to allocate pinned host memory?
cudaMallocHost()
cudaPinnedDevice()
cudaHostGPU()
cudaPageLock()
Answer: A) cudaMallocHost()
Explanation:
cudaMallocHost() allocates page-locked host memory that can be used with CUDA data-transfer operations.
42. What is the main purpose of cudaMemcpyAsync()?
- To initiate an asynchronous memory copy associated with a CUDA stream
- To compile a kernel asynchronously
- To allocate shared memory
- To create a new GPU device
Answer: A) To initiate an asynchronous memory copy associated with a CUDA stream
Explanation:
cudaMemcpyAsync() can initiate asynchronous memory transfers, allowing suitable operations to overlap when the hardware, memory type, stream, and dependencies permit.
43. What is CUDA Graphs designed to optimize?
- Repeated execution of a sequence of GPU operations by representing them as a graph
- GPU physical cooling
- Database indexing
- CUDA source-code formatting
Answer: A) Repeated execution of a sequence of GPU operations by representing them as a graph
Explanation:
CUDA Graphs allow a sequence of operations to be captured or constructed as a graph and launched with reduced CPU-side launch overhead in suitable workloads.
44. What is the purpose of CUDA error handling APIs?
- To detect and report errors from CUDA API calls and asynchronous GPU execution
- To increase GPU memory capacity
- To change thread-block dimensions automatically
- To train neural networks
Answer: A) To detect and report errors from CUDA API calls and asynchronous GPU execution
Explanation:
CUDA provides error-reporting mechanisms such as cudaGetLastError() and cudaPeekAtLastError() for diagnosing runtime and kernel-launch errors.
45. What does the CUDA compiler driver nvcc primarily do?
- Compiles CUDA source code and coordinates host and device compilation
- Allocates GPU memory at runtime
- Schedules CUDA threads directly
- Measures GPU temperature only
Answer: A) Compiles CUDA source code and coordinates host and device compilation
Explanation:
nvcc is NVIDIA's CUDA compiler driver and handles compilation of CUDA code for host and device components.
46. A CUDA vector-add kernel uses threadIdx.x + blockIdx.x * blockDim.x. What does this expression typically calculate?
- The global one-dimensional thread index
- The GPU device ID
- The number of SMs
- The shared-memory bank number
Answer: A) The global one-dimensional thread index
Explanation:
The expression combines the thread's local index with its block index and block size to calculate a unique one-dimensional index across the grid.
47. A CUDA kernel processes 10,000 array elements using blocks of 256 threads. What should the kernel normally do for threads whose calculated index is greater than or equal to 10,000?
- Skip the out-of-range element using a bounds check
- Access a random array element
- Write beyond the allocated array
- Terminate the entire GPU
Answer: A) Skip the out-of-range element using a bounds check
Explanation:
When the grid contains more threads than the data size, kernels commonly use a condition such as if (idx < n) before accessing the array.
48. A CUDA kernel is slow because each thread repeatedly reads the same data needed by neighboring threads. Which CUDA feature can often reduce redundant global-memory accesses?
- Shared memory
- Host-only memory
- Constant CPU registers
- Kernel comments
Answer: A) Shared memory
Explanation:
Threads within a block can cooperatively load reusable data into shared memory, allowing neighboring threads to reuse values rather than repeatedly accessing global memory.
49. A CUDA application launches many very small kernels repeatedly, and GPU computation is fast but CPU-side launch overhead becomes significant. Which CUDA feature can help reduce repeated launch overhead?
- CUDA Graphs
- Constant memory
- Thread-local memory
- Increasing the number of source-code comments
Answer: A) CUDA Graphs
Explanation:
CUDA Graphs can represent repeated sequences of GPU operations and replay them with lower CPU launch overhead than issuing every operation independently in suitable workloads.
50. A CUDA matrix-processing application has poor performance. Profiling shows uncoalesced global-memory accesses, significant warp divergence, excessive register usage, and low occupancy. Which optimization strategy is most appropriate?
- Optimize memory-access patterns, reduce unnecessary divergence, control register usage, and tune the kernel configuration
- Increase the number of CPU threads without changing the CUDA kernel
- Move every array to host memory
- Increase kernel complexity to use more registers
Answer: A) Optimize memory-access patterns, reduce unnecessary divergence, control register usage, and tune the kernel configuration
Explanation:
CUDA performance depends on several interacting factors. Coalesced memory access improves global-memory efficiency, reducing warp divergence can improve execution efficiency, and controlling register usage can increase the number of active warps or blocks that fit on an SM. Kernel configuration should be tuned based on profiling rather than relying on a single optimization metric.