AI Infrastructure MCQs (Multiple-Choice Questions)

Practice AI Infrastructure MCQs to test your knowledge of GPUs, AI servers, networking, storage, distributed computing, Kubernetes, GPU scheduling, model serving, monitoring, and production AI infrastructure.

AI Infrastructure MCQs

These AI Infrastructure multiple-choice questions cover the hardware and software systems required to train, deploy, scale, and operate modern artificial intelligence workloads.

List of AI Infrastructure MCQs

Below is the list of AI Infrastructure MCQs with answers and explanations.

1. What is AI infrastructure?

  1. A collection of hardware and software resources used to build and operate AI workloads
  2. A programming language for neural networks
  3. A database containing only training labels
  4. A single machine learning algorithm

Answer: A) A collection of hardware and software resources used to build and operate AI workloads

Explanation:

AI infrastructure includes compute, accelerators, memory, storage, networking, orchestration, software runtimes, monitoring, and other systems required for AI workloads.

2. Why are GPUs widely used in AI infrastructure?

  1. They are designed primarily for text editing
  2. They can execute large numbers of parallel mathematical operations
  3. They eliminate the need for memory
  4. They are exclusively designed for database storage

Answer: B) They can execute large numbers of parallel mathematical operations

Explanation:

Neural networks contain many operations that can be executed in parallel, making GPUs well suited to AI training and inference.

3. What is an AI accelerator?

  1. Specialized hardware designed to accelerate AI computations
  2. A storage protocol
  3. A database backup system
  4. A network authentication service

Answer: A) Specialized hardware designed to accelerate AI computations

Explanation:

AI accelerators include GPUs, NPUs, TPUs, and other specialized processors designed to execute machine learning operations efficiently.

4. What does GPU memory commonly store during AI workloads?

  1. Model weights, activations, tensors, and other workload data
  2. Only operating system passwords
  3. Only network configuration files
  4. Only source-code comments

Answer: A) Model weights, activations, tensors, and other workload data

Explanation:

GPU memory, commonly called VRAM or device memory, stores data required by GPU computations, including model parameters and intermediate activations.

5. Why is GPU memory capacity important when deploying large language models?

  1. The model weights and runtime state must fit within available memory or be distributed across resources
  2. GPU memory only controls network speed
  3. Large models never use GPU memory
  4. GPU memory determines the programming language

Answer: A) The model weights and runtime state must fit within available memory or be distributed across resources

Explanation:

Large models can require substantial memory for weights, activations, KV cache, and other runtime data. When one GPU is insufficient, model or workload parallelism can distribute the requirements.

6. What is GPU memory bandwidth?

  1. The rate at which data can be transferred between GPU memory and the GPU processing system
  2. The number of GPUs in a cluster
  3. The physical size of a server rack
  4. The number of Kubernetes namespaces

Answer: A) The rate at which data can be transferred between GPU memory and the GPU processing system

Explanation:

High memory bandwidth can be important for workloads that repeatedly move large amounts of data between memory and compute units.

7. What is multi-GPU computing?

  1. Using multiple GPUs to execute an AI workload
  2. Running multiple operating systems on one CPU
  3. Using multiple databases without networking
  4. Using several keyboards for one server

Answer: A) Using multiple GPUs to execute an AI workload

Explanation:

Multi-GPU systems can distribute training or inference workloads across several GPUs to increase available compute and memory capacity.

8. What is multi-node distributed training?

  1. Training a model across GPUs located on multiple servers
  2. Training only on a laptop CPU
  3. Training without network communication
  4. Training multiple unrelated models without synchronization

Answer: A) Training a model across GPUs located on multiple servers

Explanation:

Distributed training can use GPUs across multiple nodes to provide more compute and memory than a single server can provide.

9. What is data parallelism?

  1. Replicating a model across workers and processing different batches of data on each worker
  2. Splitting only the network cable
  3. Replicating storage without running a model
  4. Splitting source code into unrelated applications

Answer: A) Replicating a model across workers and processing different batches of data on each worker

Explanation:

In data parallelism, each worker generally maintains a copy of the model and processes different portions of a training batch, followed by gradient synchronization.

10. What is model parallelism?

  1. Splitting a model or its computation across multiple devices
  2. Running the same model independently without communication
  3. Copying the same dataset to every storage device only
  4. Using multiple programming languages in one application

Answer: A) Splitting a model or its computation across multiple devices

Explanation:

Model parallelism distributes portions of a model or its computations across multiple accelerators when a workload is too large or when additional parallelism is beneficial.

11. What is tensor parallelism?

  1. Splitting tensor computations across multiple devices
  2. Copying tensors to a backup disk only
  3. Compressing all tensors into text files
  4. Removing tensors from a neural network

Answer: A) Splitting tensor computations across multiple devices

Explanation:

Tensor parallelism partitions operations involving tensors across multiple devices so that a layer or computation can use multiple accelerators.

12. What is pipeline parallelism?

  1. Dividing model layers into stages that execute across different devices
  2. Running only one CPU instruction at a time
  3. Copying model files to storage
  4. Routing network packets through a single switch

Answer: A) Dividing model layers into stages that execute across different devices

Explanation:

Pipeline parallelism places different groups of model layers on different devices and can process multiple microbatches through those stages.

13. Why is high-speed networking important for distributed AI training?

  1. Workers frequently exchange data such as gradients and activations
  2. Training never communicates between GPUs
  3. Networking is used only for user authentication
  4. Network bandwidth has no relationship with distributed training

Answer: A) Workers frequently exchange data such as gradients and activations

Explanation:

Distributed training can generate substantial communication between GPUs and nodes. Network bandwidth and latency can therefore strongly affect scaling efficiency.

14. What does RDMA provide for high-performance networking?

  1. Direct memory-to-memory data transfer with reduced CPU involvement
  2. Automatic model training
  3. GPU memory expansion through software only
  4. Database schema migration

Answer: A) Direct memory-to-memory data transfer with reduced CPU involvement

Explanation:

Remote Direct Memory Access can transfer data between systems while bypassing significant portions of the traditional CPU and kernel networking path.

15. Which technology is commonly associated with high-performance GPU cluster networking?

  1. InfiniBand
  2. USB keyboard protocol
  3. SMTP
  4. HTML

Answer: A) InfiniBand

Explanation:

InfiniBand is commonly used in high-performance computing and AI clusters where high bandwidth and low latency are important.

16. What is NVLink primarily designed to provide?

  1. High-bandwidth communication between supported GPUs and related components
  2. Internet access for web browsers
  3. Object storage management
  4. Container image compression

Answer: A) High-bandwidth communication between supported GPUs and related components

Explanation:

NVLink is a high-speed interconnect technology designed to provide fast communication between supported NVIDIA GPUs and systems.

17. What is NCCL commonly used for?

  1. Collective communication between GPUs
  2. Relational database indexing
  3. Container image building
  4. Web application authentication

Answer: A) Collective communication between GPUs

Explanation:

NCCL provides optimized collective communication primitives used in distributed GPU workloads, including operations such as all-reduce and broadcast.

18. What does an all-reduce operation commonly do in distributed training?

  1. Combines values from multiple workers and distributes the result to the workers
  2. Deletes all model weights
  3. Copies only files to object storage
  4. Disables GPU communication

Answer: A) Combines values from multiple workers and distributes the result to the workers

Explanation:

All-reduce is commonly used to synchronize gradients or other distributed values across training workers.

19. What is AI storage infrastructure responsible for?

  1. Providing access to datasets, model artifacts, checkpoints, logs, and other AI data
  2. Executing every neural-network operation
  3. Replacing GPUs
  4. Managing only user passwords

Answer: A) Providing access to datasets, model artifacts, checkpoints, logs, and other AI data

Explanation:

AI workloads require storage for large datasets, model checkpoints, training artifacts, model weights, logs, and other data.

20. Why can storage throughput affect AI training performance?

  1. Slow data delivery can leave expensive accelerators waiting for input
  2. Storage has no relationship with training data
  3. GPUs do not require data
  4. Storage controls only user authentication

Answer: A) Slow data delivery can leave expensive accelerators waiting for input

Explanation:

If the storage system cannot supply training data quickly enough, GPUs can become underutilized while waiting for data.

21. Which storage type is commonly used for large-scale AI datasets in cloud environments?

  1. Object storage
  2. CPU registers
  3. GPU cache only
  4. Keyboard memory

Answer: A) Object storage

Explanation:

Object storage is commonly used for large datasets, model artifacts, checkpoints, and other durable data in cloud-based AI platforms.

22. What is local NVMe storage useful for in AI infrastructure?

  1. Providing high-speed local access to frequently used data or model artifacts
  2. Replacing all network infrastructure
  3. Managing Kubernetes authentication
  4. Training models without compute resources

Answer: A) Providing high-speed local access to frequently used data or model artifacts

Explanation:

Local NVMe storage can provide high I/O performance and low latency for caching datasets, model weights, temporary files, and other frequently accessed data.

23. What is a model checkpoint?

  1. A saved state of a model and associated training information
  2. A network switch
  3. A GPU driver
  4. A database query

Answer: A) A saved state of a model and associated training information

Explanation:

Checkpoints allow training to be resumed or models to be recovered from previously saved states.

24. What is Kubernetes used for in AI infrastructure?

  1. Orchestrating containerized AI workloads and infrastructure services
  2. Replacing all GPUs with CPUs
  3. Creating neural-network weights automatically
  4. Acting as a physical network cable

Answer: A) Orchestrating containerized AI workloads and infrastructure services

Explanation:

Kubernetes provides scheduling, workload management, service discovery, scaling, health management, and resource orchestration for containerized AI workloads.

25. How can Kubernetes expose GPUs as schedulable resources?

  1. Through device plugins
  2. Through HTML metadata
  3. Through DNS records only
  4. Through SQL triggers

Answer: A) Through device plugins

Explanation:

Kubernetes uses its device plugin framework to allow vendor-specific resources such as GPUs and high-performance NICs to be advertised to the kubelet and made available to workloads.

26. In Kubernetes, how is a GPU normally requested by a Pod?

  1. As a GPU resource limit such as nvidia.com/gpu
  2. As an HTML attribute
  3. As a DNS record
  4. As a database index

Answer: A) As a GPU resource limit such as nvidia.com/gpu

Explanation:

After a compatible device plugin exposes the GPU resource, a Pod can request the resource through its container resource specification. Kubernetes documents GPU resources as custom schedulable resources.

27. What is the purpose of a GPU Operator?

  1. Automating deployment and lifecycle management of GPU software components in a cluster
  2. Training every AI model automatically
  3. Replacing Kubernetes scheduling
  4. Converting GPUs into storage devices

Answer: A) Automating deployment and lifecycle management of GPU software components in a cluster

Explanation:

A GPU Operator can automate components such as drivers, container toolkit integration, device plugins, and GPU monitoring components in Kubernetes environments.

28. What does GPU scheduling accomplish?

  1. Assigns available GPU resources to workloads according to resource requirements and scheduling rules
  2. Changes model architecture automatically
  3. Creates training datasets
  4. Converts GPUs into CPUs

Answer: A) Assigns available GPU resources to workloads according to resource requirements and scheduling rules

Explanation:

GPU scheduling helps determine where GPU-enabled workloads should run based on available resources and cluster policies.

29. What is GPU utilization?

  1. A measure of how actively a GPU is being used
  2. The physical number of GPUs in a data center
  3. The amount of disk storage available
  4. The number of Kubernetes Pods only

Answer: A) A measure of how actively a GPU is being used

Explanation:

GPU utilization is an operational metric that helps determine whether allocated GPU resources are actively processing workloads.

30. Why is GPU utilization an important infrastructure metric?

  1. Low utilization can indicate that expensive accelerator resources are being underused
  2. It determines the model's programming language
  3. It removes the need for monitoring
  4. It measures database table size

Answer: A) Low utilization can indicate that expensive accelerator resources are being underused

Explanation:

Accelerators can represent a significant infrastructure cost. Monitoring utilization helps identify bottlenecks, idle resources, and opportunities for better workload placement.

31. What is MIG in NVIDIA GPU infrastructure?

  1. Multi-Instance GPU
  2. Memory Integration Gateway
  3. Machine Inference Group
  4. Model Infrastructure Graph

Answer: A) Multi-Instance GPU

Explanation:

MIG allows supported NVIDIA GPUs to be partitioned into separate GPU instances, enabling multiple workloads to share a physical GPU with hardware-level isolation.

32. What is GPU virtualization useful for?

  1. Allowing multiple virtual machines or workloads to share GPU resources under supported configurations
  2. Increasing model accuracy automatically
  3. Replacing all GPU memory
  4. Removing the need for drivers

Answer: A) Allowing multiple virtual machines or workloads to share GPU resources under supported configurations

Explanation:

GPU virtualization can improve resource sharing and isolation by allowing supported virtualized workloads to access GPU resources.

33. What is an AI inference server responsible for?

  1. Loading AI models and serving prediction requests
  2. Training every model from scratch
  3. Managing only physical power cables
  4. Replacing the data center network

Answer: A) Loading AI models and serving prediction requests

Explanation:

Inference servers host models and provide APIs or other interfaces through which applications can submit inference requests.

34. What is model serving?

  1. Making a trained model available for inference requests
  2. Training a model without data
  3. Backing up physical servers only
  4. Creating a network switch

Answer: A) Making a trained model available for inference requests

Explanation:

Model serving includes loading models, accepting requests, executing inference, returning results, and often handling scaling and observability.

35. What is batching in AI inference?

  1. Processing multiple inference inputs together
  2. Deleting multiple model versions
  3. Creating multiple databases
  4. Splitting one GPU into CPUs

Answer: A) Processing multiple inference inputs together

Explanation:

Batching can improve hardware utilization by processing multiple requests together, although larger batches may increase individual request latency.

36. What is continuous or dynamic batching in an inference server?

  1. Dynamically combining incoming requests into batches during serving
  2. Training models continuously without validation
  3. Copying all model files to every server
  4. Creating a new GPU for each request

Answer: A) Dynamically combining incoming requests into batches during serving

Explanation:

Dynamic batching allows an inference system to combine requests arriving at different times to improve accelerator utilization while serving online traffic.

37. What does autoscaling do in AI infrastructure?

  1. Adjusts available compute resources according to workload demand
  2. Changes model weights automatically
  3. Increases dataset quality automatically
  4. Removes network security controls

Answer: A) Adjusts available compute resources according to workload demand

Explanation:

Autoscaling can increase or decrease compute capacity based on workload demand, resource utilization, queue length, or other signals.

38. Why is observability important for AI infrastructure?

  1. It provides visibility into performance, resource usage, failures, and system health
  2. It eliminates the need for infrastructure
  3. It automatically improves model accuracy
  4. It replaces all security controls

Answer: A) It provides visibility into performance, resource usage, failures, and system health

Explanation:

Observability allows operators to understand infrastructure and workload behavior using metrics, logs, traces, and health signals.

39. Which metric is particularly useful for measuring AI inference responsiveness?

  1. Inference latency
  2. Number of source-code comments
  3. Disk filename length
  4. Number of Kubernetes namespaces only

Answer: A) Inference latency

Explanation:

Inference latency measures how long it takes to process an inference request and is important for applications with response-time requirements.

40. What is throughput in an AI inference system?

  1. The amount of inference work completed per unit of time
  2. The number of GPUs installed physically
  3. The amount of RAM installed
  4. The number of model files stored

Answer: A) The amount of inference work completed per unit of time

Explanation:

Throughput can be measured using metrics such as requests per second or tokens per second, depending on the workload.

41. What is fault tolerance in AI infrastructure?

  1. The ability of a system to continue operating or recover when components fail
  2. The ability to avoid using GPUs
  3. The ability to increase model size without memory
  4. The ability to remove all monitoring

Answer: A) The ability of a system to continue operating or recover when components fail

Explanation:

Fault-tolerant AI infrastructure uses techniques such as redundancy, checkpointing, workload recovery, health monitoring, and automated replacement to handle failures.

42. Why are checkpoints important for fault tolerance during large-scale training?

  1. They allow training to resume from a saved state after a failure
  2. They eliminate the need for GPUs
  3. They prevent all hardware failures
  4. They increase network latency

Answer: A) They allow training to resume from a saved state after a failure

Explanation:

Long-running training jobs can lose significant progress after a failure. Checkpoints provide recovery points from which training can continue.

43. What is a container image used for in AI infrastructure?

  1. Packaging an application, runtime, libraries, and dependencies
  2. Storing only GPU electricity measurements
  3. Replacing physical GPUs
  4. Measuring network latency

Answer: A) Packaging an application, runtime, libraries, and dependencies

Explanation:

Container images provide reproducible application environments and are widely used to package AI training and inference workloads.

44. Why are compatible GPU drivers and runtime libraries important?

  1. AI applications need compatible software to communicate correctly with GPU hardware
  2. Drivers determine the training dataset labels
  3. Drivers replace model weights
  4. Runtime libraries are used only for documentation

Answer: A) AI applications need compatible software to communicate correctly with GPU hardware

Explanation:

GPU drivers, container runtimes, CUDA libraries, and framework components must be compatible with the deployed hardware and software stack.

45. What is infrastructure-as-code (IaC) useful for in AI environments?

  1. Defining and reproducing infrastructure configuration through code or declarative definitions
  2. Training neural networks without datasets
  3. Replacing GPUs with software
  4. Increasing model accuracy automatically

Answer: A) Defining and reproducing infrastructure configuration through code or declarative definitions

Explanation:

IaC can make infrastructure provisioning repeatable, version-controlled, and easier to automate across environments.

46. An AI cluster has powerful GPUs, but distributed training scales poorly because GPUs spend significant time waiting for communication. Which infrastructure component should be investigated?

  1. Interconnect and network fabric
  2. Keyboard layout
  3. Web browser theme
  4. Database column names

Answer: A) Interconnect and network fabric

Explanation:

Distributed training depends heavily on communication between workers. Network bandwidth, latency, topology, congestion, and GPU interconnects can all affect scaling efficiency.

47. A Kubernetes cluster has GPUs installed on its nodes, but Pods cannot request or discover those GPUs. What infrastructure component should be checked first?

  1. GPU device plugin and associated GPU software stack
  2. DNS TXT records
  3. Object-storage bucket naming
  4. Application HTML templates

Answer: A) GPU device plugin and associated GPU software stack

Explanation:

Kubernetes relies on device plugins to advertise specialized hardware such as GPUs as schedulable resources. A missing or malfunctioning device plugin can prevent workloads from consuming the GPU resources.

48. An inference platform has sufficient GPU compute, but requests experience high latency because model weights take too long to become available on serving workers. Which infrastructure area should be investigated?

  1. Model storage, caching, and data-transfer paths
  2. Keyboard drivers
  3. Database table formatting
  4. Source-code indentation

Answer: A) Model storage, caching, and data-transfer paths

Explanation:

Inference startup and model loading can be affected by storage throughput, network bandwidth, cache placement, and the path used to move model artifacts into serving workers.

49. A company needs to run thousands of AI workloads across shared GPU clusters while controlling GPU allocation, networking, storage, workload isolation, and monitoring. Which architecture is most appropriate?

  1. A managed AI platform built on Kubernetes with GPU-aware scheduling, high-performance networking, storage, and observability
  2. A single workstation shared manually by all users
  3. Independent laptops with no centralized management
  4. A database server without accelerators

Answer: A) A managed AI platform built on Kubernetes with GPU-aware scheduling, high-performance networking, storage, and observability

Explanation:

Large shared AI environments require orchestration, resource scheduling, networking, storage, isolation, monitoring, and lifecycle management. Modern AI infrastructure platforms commonly integrate these layers around Kubernetes.

50. An organization is building an AI infrastructure platform for large-scale training and inference. It requires GPU clusters, high-speed networking, scalable storage, Kubernetes orchestration, GPU-aware scheduling, model serving, monitoring, secure updates, and failure recovery. Which design best represents a production AI infrastructure stack?

  1. GPU compute + high-performance network fabric + scalable storage + Kubernetes orchestration + GPU scheduling + model serving + observability + security and recovery
  2. CPUs and local disks without networking
  3. Object storage with no compute layer
  4. A single GPU workstation with manual model execution

Answer: A) GPU compute + high-performance network fabric + scalable storage + Kubernetes orchestration + GPU scheduling + model serving + observability + security and recovery

Explanation:

A production AI infrastructure platform requires multiple coordinated layers. Accelerated compute executes AI workloads, high-performance networking supports distributed communication, storage provides datasets and model artifacts, Kubernetes manages workloads, GPU-aware scheduling allocates accelerators, serving infrastructure exposes models, and observability, security, and recovery provide operational reliability. Current AI infrastructure reference architectures similarly describe compute, networking, storage, Kubernetes, scheduling, serving, health, and lifecycle management as interconnected infrastructure capabilities.

Comments and Discussions!

Load comments ↻



Copyright © 2026 www.includehelp.com. All rights reserved.