Home »
Trending Technologies MCQs
AI Model Optimization MCQs (Multiple-Choice Questions)
Practice AI Model Optimization MCQs to test your knowledge of techniques used to improve artificial intelligence model performance, reduce computational costs, and simplify deployment. These questions cover model quantization, pruning, knowledge distillation, mixed-precision computing, inference acceleration, memory optimization, and computational graph transformations. They are useful for AI engineers, machine learning practitioners, researchers, and students preparing for technical interviews or assessments. The set includes both foundational and practical questions covering modern AI model optimization systems.
AI Model Optimization MCQs
These AI Model Optimization multiple-choice questions cover important concepts such as post-training quantization, quantization-aware training, structured pruning, low-rank approximation, knowledge distillation, operator fusion, compiler optimization, dynamic batching, caching, and hardware-aware inference. This set combines conceptual, technical, and scenario-based questions to help test your understanding of AI model optimization systems.
AI Model Optimization MCQs cover the methods used to reduce model size, improve inference speed, manage memory consumption, and balance computational efficiency with predictive quality. Each question includes an answer and explanation.
List of AI Model Optimization MCQs
The following AI Model Optimization multiple-choice questions cover model compression, numerical precision, optimization algorithms, inference runtimes, large language model optimization, performance benchmarking, and deployment considerations.
1. What is the primary objective of AI model optimization?
- To increase model size regardless of performance
- To improve performance, resource usage, or deployment suitability while maintaining acceptable model quality
- To remove all training data from the development process
- To ensure every model uses the same architecture
Answer: B) To improve performance, resource usage, or deployment suitability while maintaining acceptable model quality
Explanation:
AI model optimization aims to improve metrics such as inference latency, throughput, memory consumption, energy usage, and model size. The appropriate trade-offs depend on the application and its accuracy requirements.
2. What is model quantization?
- Increasing the number of neural network layers
- Duplicating all model parameters
- Converting every model output into a text string
- Representing model weights or activations using lower-precision numerical formats
Answer: D) Representing model weights or activations using lower-precision numerical formats
Explanation:
Quantization maps higher-precision values to lower-precision representations, such as INT8 or INT4. It can reduce model storage and improve inference speed on compatible hardware, although it may introduce accuracy degradation.
3. What is the main difference between post-training quantization (PTQ) and quantization-aware training (QAT)?
- PTQ quantizes a trained model without quantization-aware retraining, while QAT simulates quantization effects during training or fine-tuning
- PTQ requires every model to be trained from scratch, while QAT never uses training data
- PTQ applies only to databases, while QAT applies only to image files
- There is no difference between the two methods
Answer: A) PTQ quantizes a trained model without quantization-aware retraining, while QAT simulates quantization effects during training or fine-tuning
Explanation:
PTQ is generally applied after model training and may require calibration data. QAT exposes the model to simulated quantization effects during training, allowing parameters to adapt and potentially preserving accuracy more effectively at low precision.
4. Why is calibration data used in static quantization?
- To increase the number of model layers automatically
- To determine the final test labels before inference
- To estimate activation ranges and determine quantization parameters
- To replace the trained weights with random values
Answer: C) To estimate activation ranges and determine quantization parameters
Explanation:
Static quantization uses representative inputs to estimate activation ranges and calculate scales and zero points. Poorly representative calibration data can lead to inaccurate quantization parameters and degraded model quality.
5. What is a major advantage of INT8 quantization compared with FP32 inference?
- It guarantees higher accuracy for every model
- It can reduce memory requirements and accelerate supported computations
- It eliminates all numerical approximation errors
- It removes the need for inference hardware
Answer: B) It can reduce memory requirements and accelerate supported computations
Explanation:
INT8 values require fewer bits than FP32 values. Quantized models can use less memory bandwidth and benefit from specialized integer arithmetic, but the actual speedup depends on the hardware, operators, runtime, and workload.
6. What is the purpose of a zero point in affine quantization?
- To force every model weight to zero
- To specify the number of hidden layers
- To calculate the model's training duration
- To represent real-valued zero within the quantized integer representation
Answer: D) To represent real-valued zero within the quantized integer representation
Explanation:
Affine quantization commonly uses a scale and zero point to map real values to integers. The zero point determines which integer represents real-valued zero, while the scale controls the spacing between represented values.
7. What is mixed-precision computing in AI model execution?
- Using different numerical precisions for supported operations or tensors within a model
- Running two unrelated operating systems on one computer
- Training only on integer-valued datasets
- Using multiple datasets without updating model parameters
Answer: A) Using different numerical precisions for supported operations or tensors within a model
Explanation:
Mixed precision may combine FP32, FP16, or BF16 operations depending on numerical requirements and hardware support. It can reduce memory usage and improve throughput while retaining higher precision where needed.
8. Which statement best describes FP16 and BF16 formats?
- Both are 32-bit integer formats
- Both always provide greater precision than FP32
- Both are 16-bit floating-point formats, but BF16 has a wider exponent range than FP16
- Both can represent only positive integers
Answer: C) Both are 16-bit floating-point formats, but BF16 has a wider exponent range than FP16
Explanation:
FP16 has more fraction bits, while BF16 uses more exponent bits and therefore offers a range closer to FP32. Their numerical behavior differs, so the appropriate format depends on model stability, hardware support, and workload requirements.
9. What is model pruning?
- Adding duplicate layers to a trained network
- Removing selected weights, connections, neurons, or structures considered less important
- Increasing every parameter to its maximum value
- Converting all model outputs into probability distributions
Answer: B) Removing selected weights, connections, neurons, or structures considered less important
Explanation:
Pruning reduces model complexity by removing selected parameters or structures. Depending on the pruning method and hardware, it may reduce storage, computation, or inference latency.
10. What distinguishes structured pruning from unstructured pruning?
- Structured pruning changes only the training dataset
- Unstructured pruning always removes entire neural network layers
- Both techniques remove exactly the same parameters in every case
- Structured pruning removes groups such as channels or heads, while unstructured pruning removes individual parameters or connections
Answer: D) Structured pruning removes groups such as channels or heads, while unstructured pruning removes individual parameters or connections
Explanation:
Structured pruning removes organized components such as channels, filters, or attention heads, often producing smaller dense computations. Unstructured pruning creates sparse weight patterns, which require compatible sparse kernels to translate into substantial speed gains.
11. What is knowledge distillation in machine learning?
- Training a smaller student model to learn from outputs or representations of a teacher model
- Deleting the teacher model before collecting any training signals
- Converting model weights into database indexes
- Increasing the size of every layer without changing its behavior
Answer: A) Training a smaller student model to learn from outputs or representations of a teacher model
Explanation:
Knowledge distillation transfers information from a teacher model to a student model. The training objective may use teacher logits, softened probability distributions, intermediate representations, or a combination of these signals.
12. What is the primary purpose of low-rank approximation in model optimization?
- To make every model matrix larger
- To increase the number of output classes
- To approximate a large matrix using lower-dimensional factors
- To eliminate all matrix multiplication operations
Answer: C) To approximate a large matrix using lower-dimensional factors
Explanation:
Low-rank approximation represents a matrix as the product of smaller matrices. When the approximation rank is sufficiently low, it can reduce parameter storage and computation, though the approximation may affect model quality.
13. What is the purpose of operator fusion in an inference runtime?
- To force every operator to execute on a separate device
- To combine compatible operations into a larger execution unit, reducing overhead and potentially improving memory access
- To duplicate every intermediate tensor
- To prevent the runtime from analyzing the computation graph
Answer: B) To combine compatible operations into a larger execution unit, reducing overhead and potentially improving memory access
Explanation:
Operator fusion combines compatible operations, such as certain convolution and activation sequences, into a single kernel or execution unit. It can reduce kernel launches and intermediate memory traffic when supported by the compiler or runtime.
14. What is computational graph optimization?
- Manually increasing every tensor dimension
- Converting all neural networks into recurrent networks
- Removing the model's input and output definitions
- Transforming a model's operation graph to remove redundancy or improve execution
Answer: D) Transforming a model's operation graph to remove redundancy or improve execution
Explanation:
Graph optimizations can include constant folding, redundant operation elimination, operator fusion, and layout transformations. These transformations aim to improve execution while preserving the model's intended semantics within acceptable numerical tolerances.
15. What does constant folding mean in model graph optimization?
- Evaluating operations whose inputs are known constants ahead of runtime
- Changing all learned weights to the same value
- Recomputing every constant for every inference request
- Removing all numerical operations from the graph
Answer: A) Evaluating operations whose inputs are known constants ahead of runtime
Explanation:
Constant folding evaluates computations that depend only on known constants during model preparation or compilation. This reduces work that would otherwise be performed during inference.
16. Why can converting a model to ONNX help with optimized deployment?
- It guarantees that every model runs faster on all hardware
- It removes the need for numerical validation
- It provides an interoperable model representation that compatible inference runtimes can optimize
- It automatically converts every model into INT4
Answer: C) It provides an interoperable model representation that compatible inference runtimes can optimize
Explanation:
ONNX defines a portable representation for supported machine learning operators and model graphs. Compatible runtimes, such as ONNX Runtime, can apply graph optimizations and use hardware-specific execution providers, but operator and feature compatibility must be checked.
17. What is the main purpose of an inference engine such as NVIDIA TensorRT?
- To collect training labels from users automatically
- To optimize and execute supported models for inference on compatible hardware
- To replace the model's learned parameters with random values
- To perform only database transactions
Answer: B) To optimize and execute supported models for inference on compatible hardware
Explanation:
Inference engines can select kernels, optimize computation graphs, support suitable numerical precisions, and generate execution plans for target hardware. Actual gains depend on the model, supported operations, configuration, and input shapes.
18. What is kernel selection in a deep learning compiler?
- Choosing the training dataset's file extension
- Selecting the model's final output label manually
- Removing all hardware-specific execution paths
- Choosing an implementation of an operation that suits the hardware and workload
Answer: D) Choosing an implementation of an operation that suits the hardware and workload
Explanation:
Compilers and inference runtimes may select kernels based on tensor shapes, numerical precision, memory layout, and hardware capabilities. The selected implementation can affect latency, throughput, and resource usage.
19. What is the primary purpose of operator benchmarking during model optimization?
- To measure the performance of individual operations or execution components
- To determine the model's marketing budget
- To ensure all operators have identical execution times
- To eliminate the need for end-to-end testing
Answer: A) To measure the performance of individual operations or execution components
Explanation:
Operator-level benchmarks help identify expensive operations and potential bottlenecks. They should be complemented by end-to-end measurements because launch overhead, data movement, scheduling, and request processing also affect total latency.
20. What is the difference between inference latency and throughput?
- Both always measure model size
- Latency measures storage consumption, while throughput measures training accuracy
- Latency measures the time required for an inference request, while throughput measures the number of requests or samples processed per unit of time
- Throughput is measured only during model training
Answer: C) Latency measures the time required for an inference request, while throughput measures the number of requests or samples processed per unit of time
Explanation:
Latency is important for interactive applications, while throughput measures processing capacity over time. Optimizations that improve throughput through batching or parallelism may increase the latency experienced by individual requests.
21. Why should a GPU inference benchmark include warm-up iterations?
- To increase the number of model parameters
- To allow initialization, compilation, or kernel-selection overhead to settle before timed measurements
- To force the GPU to operate at zero utilization
- To ensure that all model outputs are identical
Answer: B) To allow initialization, compilation, or kernel-selection overhead to settle before timed measurements
Explanation:
Initial executions may include context setup, memory allocation, compilation, or kernel-selection overhead. Warm-up iterations help separate these costs from steady-state inference performance.
22. What is batch inference?
- Executing only one neural network layer at a time
- Training a model without calculating gradients
- Running every prediction on a separate physical server
- Processing multiple input examples together in one inference operation
Answer: D) Processing multiple input examples together in one inference operation
Explanation:
Batch inference groups multiple examples into a batch so the hardware can process them together. This can improve utilization and throughput, although larger batches may require more memory and increase waiting time.
23. What is dynamic batching in an AI inference service?
- Combining incoming requests into batches at runtime, subject to configured limits or waiting policies
- Changing the model architecture after every prediction
- Removing all requests that arrive simultaneously
- Training a new model for every user request
Answer: A) Combining incoming requests into batches at runtime, subject to configured limits or waiting policies
Explanation:
Dynamic batching combines compatible requests as they arrive to improve accelerator utilization. A service can configure maximum batch sizes and queueing delays to balance throughput against latency requirements.
24. What is a common limitation of increasing batch size during inference?
- It always reduces memory usage
- It prevents parallel computation on GPUs
- It can increase memory consumption and request latency
- It guarantees that numerical precision improves
Answer: C) It can increase memory consumption and request latency
Explanation:
Larger batches often improve throughput until hardware resources become saturated. However, they consume more memory and may require requests to wait longer before a batch is dispatched.
25. What is KV caching in autoregressive transformer inference?
- Storing every generated answer permanently in a database
- Reusing previously computed key and value tensors from earlier tokens to avoid recomputing them at every generation step
- Removing the attention mechanism from the model
- Quantizing every input token into a single bit
Answer: B) Reusing previously computed key and value tensors from earlier tokens to avoid recomputing them at every generation step
Explanation:
KV caching stores attention keys and values for previously processed tokens. It reduces repeated computation during autoregressive generation, but cache memory can become a major constraint for long sequences and large concurrent workloads.
26. What is paged attention designed to improve in large language model serving?
- The number of labels in a classification dataset
- The physical resolution of model output images
- The training accuracy of every model without retraining
- KV-cache memory management by organizing cached data into manageable blocks or pages
Answer: D) KV-cache memory management by organizing cached data into manageable blocks or pages
Explanation:
Paged attention techniques manage KV-cache storage in blocks rather than requiring one large contiguous allocation per sequence. This can reduce memory fragmentation and improve the number of concurrent requests a serving system can support.
27. What is speculative decoding in large language model inference?
- Using a smaller draft model to propose tokens that a target model verifies
- Removing the target model from the generation process
- Predicting tokens without using any model parameters
- Replacing every generated token with a random token
Answer: A) Using a smaller draft model to propose tokens that a target model verifies
Explanation:
Speculative decoding uses a draft model to propose multiple tokens and a target model to verify them. With a compatible verification procedure, it can reduce sequential generation overhead while preserving the target model's intended output distribution.
28. What is continuous batching in large language model serving?
- Processing only fixed-size batches that never change
- Combining model weights from unrelated models into one tensor
- Adding and removing requests from execution batches as individual sequences finish or new requests arrive
- Disabling token generation until every request is completed
Answer: C) Adding and removing requests from execution batches as individual sequences finish or new requests arrive
Explanation:
Continuous batching dynamically manages active sequences during generation. Unlike a static batch that waits for all members to finish, it can use freed execution slots for new requests and improve accelerator utilization.
29. What is weight-only quantization in large language models?
- Quantizing only the input text while keeping the model unchanged
- Quantizing model weights while leaving activations at their original precision or handling them separately
- Removing all trainable weights from the model
- Quantizing only the final generated sentence
Answer: B) Quantizing model weights while leaving activations at their original precision or handling them separately
Explanation:
Weight-only quantization reduces the storage requirements of model parameters, commonly using low-bit formats. During execution, weights may be dequantized or processed through specialized kernels, so actual speed gains depend on the implementation and hardware.
30. What is group-wise quantization?
- Applying one shared quantization scale to every model in an organization
- Assigning a unique output label to each weight
- Converting all weights into floating-point numbers with unlimited precision
- Quantizing groups of values using separate quantization parameters for each group
Answer: D) Quantizing groups of values using separate quantization parameters for each group
Explanation:
Group-wise quantization uses separate scales and, where applicable, zero points for groups of weights. It can represent local value ranges more accurately than a single scale for an entire large tensor, at the cost of storing and processing additional quantization metadata.
31. What is the purpose of a memory-efficient attention implementation?
- To reduce unnecessary intermediate memory traffic and memory consumption during attention computation
- To increase the number of model parameters automatically
- To eliminate all query, key, and value computations
- To guarantee constant memory use regardless of sequence length
Answer: A) To reduce unnecessary intermediate memory traffic and memory consumption during attention computation
Explanation:
Optimized attention kernels can process attention in tiles and avoid materializing large intermediate matrices in high-bandwidth memory. FlashAttention is a well-known example designed to improve attention memory efficiency and execution speed.
32. Why can sequence length strongly affect transformer inference memory usage?
- Sequence length changes the model's file extension
- Sequence length determines the physical size of the GPU
- Longer sequences require additional activations or cached key-value states, and some attention computations scale quadratically with sequence length
- Longer sequences automatically reduce every intermediate tensor
Answer: C) Longer sequences require additional activations or cached key-value states, and some attention computations scale quadratically with sequence length
Explanation:
For conventional full self-attention, computation and attention-score storage can scale quadratically with sequence length. During autoregressive inference, KV-cache memory typically grows approximately linearly with the number of cached tokens, model dimensions, and concurrent sequences.
33. What is LoRA primarily used for in large language model optimization?
- Increasing every layer's hidden dimension without limits
- Adapting a pretrained model using trainable low-rank updates while keeping the original weights frozen in the standard approach
- Replacing model training with database replication
- Converting every model parameter to a one-bit integer
Answer: B) Adapting a pretrained model using trainable low-rank updates while keeping the original weights frozen in the standard approach
Explanation:
Low-Rank Adaptation (LoRA) represents weight updates using low-rank matrices. It reduces the number of trainable parameters and optimizer states required for adaptation, although inference and deployment benefits depend on whether adapters are merged or executed separately.
34. What is the main advantage of parameter-efficient fine-tuning methods?
- They eliminate the need for pretrained models in every workflow
- They guarantee better task accuracy than full fine-tuning
- They ensure that no training data is required
- They reduce the number of parameters that must be updated during adaptation
Answer: D) They reduce the number of parameters that must be updated during adaptation
Explanation:
Parameter-efficient fine-tuning methods update only selected parameters, adapters, or low-rank components. They can reduce optimizer-state memory and training costs compared with updating every model parameter.
35. What is the purpose of gradient checkpointing during model training?
- To reduce activation memory by recomputing selected intermediate results during backpropagation
- To store every activation permanently on the GPU
- To eliminate all forward-pass computations
- To replace the optimizer with a checkpoint file
Answer: A) To reduce activation memory by recomputing selected intermediate results during backpropagation
Explanation:
Gradient checkpointing saves selected activations and recomputes other intermediate values when needed for backward propagation. It trades additional computation for reduced peak memory consumption.
36. What is activation memory in a neural network?
- The storage required only for the final model file
- The permanent disk space used by the operating system
- The memory required to hold intermediate outputs produced by network operations
- The number of labels in a training dataset
Answer: C) The memory required to hold intermediate outputs produced by network operations
Explanation:
Activations are intermediate tensor values generated during forward computation. Training often retains many of them for backpropagation, while inference can use different memory-reuse strategies to reduce peak usage.
37. What does memory bandwidth measure in AI hardware?
- The number of model layers available
- The rate at which data can be transferred between memory and processing components
- The number of training labels processed per epoch in every case
- The storage capacity of the model's vocabulary alone
Answer: B) The rate at which data can be transferred between memory and processing components
Explanation:
Memory bandwidth determines how quickly data can move between memory and compute units. Workloads that repeatedly read large weight tensors may be memory-bandwidth-bound even when the hardware has substantial arithmetic capacity.
38. What is a compute-bound model operation?
- An operation that performs no mathematical computation
- An operation that is always limited by network latency
- An operation that requires no hardware resources
- An operation whose performance is primarily limited by available computational throughput
Answer: D) An operation whose performance is primarily limited by available computational throughput
Explanation:
A compute-bound operation is limited mainly by the rate at which arithmetic can be performed. Increasing compute utilization or selecting a more suitable kernel may help, whereas reducing memory traffic is more important for a memory-bound operation.
39. What is the roofline model used for in performance analysis?
- To relate attainable performance to arithmetic intensity and hardware compute and memory-bandwidth limits
- To calculate the number of classes in a classifier
- To select training labels automatically
- To determine the physical height of a server rack
Answer: A) To relate attainable performance to arithmetic intensity and hardware compute and memory-bandwidth limits
Explanation:
The roofline model helps determine whether a workload is likely limited by computational throughput or memory bandwidth. Arithmetic intensity measures operations performed per byte transferred, providing a useful basis for selecting optimization strategies.
40. Why should AI model performance be benchmarked using representative production inputs?
- To guarantee that every request has the same execution time
- To avoid testing the optimized model on actual input shapes
- To reveal performance behavior under realistic sequence lengths, batch sizes, and workload distributions
- To remove the need for monitoring after deployment
Answer: C) To reveal performance behavior under realistic sequence lengths, batch sizes, and workload distributions
Explanation:
Input shapes and request patterns influence memory use, kernel selection, batching, and latency. Representative benchmarks provide more reliable deployment estimates than tests using only convenient synthetic inputs.
41. What is the purpose of model distillation when deploying AI on edge devices?
- To make the student model larger than every available device can support
- To transfer useful behavior to a smaller model that may better fit the device's compute and memory constraints
- To eliminate all evaluation requirements
- To force the device to use cloud inference for every prediction
Answer: B) To transfer useful behavior to a smaller model that may better fit the device's compute and memory constraints
Explanation:
A distilled model can require less memory and computation than its teacher. It is useful for constrained deployments, but the student must still be evaluated for task quality, latency, energy use, and hardware compatibility.
42. What is the purpose of hardware-aware neural architecture search?
- To generate model architectures without evaluating any candidate
- To optimize only the number of training examples
- To ensure every architecture uses the same inference latency
- To search for architectures that satisfy performance or resource constraints on a target hardware platform
Answer: D) To search for architectures that satisfy performance or resource constraints on a target hardware platform
Explanation:
Hardware-aware neural architecture search considers factors such as latency, memory consumption, energy usage, or throughput when selecting model structures. An architecture with fewer parameters is not necessarily faster if its operations are poorly supported by the target hardware.
43. What is the main purpose of caching repeated inference results?
- To avoid recomputing outputs for eligible repeated inputs when cached results remain valid
- To guarantee correct predictions for inputs never seen before
- To increase the number of model parameters
- To replace all model evaluation procedures
Answer: A) To avoid recomputing outputs for eligible repeated inputs when cached results remain valid
Explanation:
Inference caching can reduce latency and computation when identical requests occur frequently. It requires appropriate cache keys, expiration or invalidation rules, and care around user-specific data, stochastic generation, and model-version changes.
44. What is a key consideration when deploying a quantized model to a mobile device?
- Whether the model contains the maximum possible number of layers
- Whether the training dataset is stored in the device's user interface
- Whether the target runtime and processor support the selected precision and operators efficiently
- Whether every parameter is represented using FP32
Answer: C) Whether the target runtime and processor support the selected precision and operators efficiently
Explanation:
Quantization benefits depend on support from the target runtime, CPU, GPU, or neural processing unit. An unsupported or poorly optimized operator may fall back to another implementation and reduce the expected performance gains.
45. What is the main risk of applying aggressive quantization without validating model quality?
- The model automatically gains additional training data
- The model may suffer unacceptable accuracy degradation or changes in task behavior
- The model becomes guaranteed to run faster on every processor
- The model no longer requires inference testing
Answer: B) The model may suffer unacceptable accuracy degradation or changes in task behavior
Explanation:
Lower precision reduces the range or resolution of representable values. Aggressive quantization can affect sensitive layers or outputs, so validation should include task-specific quality metrics and, where relevant, safety and robustness evaluations.
46. What is the purpose of an ablation study during model optimization?
- To remove all model evaluation metrics
- To guarantee that every optimization technique improves performance
- To compare models using different datasets without controlling any variables
- To isolate the contribution of a component or optimization technique by removing or changing it
Answer: D) To isolate the contribution of a component or optimization technique by removing or changing it
Explanation:
An ablation study measures how a model changes when a component or technique is removed or modified. It helps determine whether a proposed optimization actually contributes to quality, speed, memory savings, or other objectives.
47. Which metric is most appropriate for evaluating whether an optimized model meets a strict interactive response-time requirement?
- Inference latency at relevant percentiles, such as p95 or p99
- The number of source-code comments
- The total number of training files alone
- The size of the development team's repository
Answer: A) Inference latency at relevant percentiles, such as p95 or p99
Explanation:
Tail-latency metrics such as p95 and p99 show how slower requests behave, rather than reporting only an average. They are useful for identifying latency spikes caused by batching, scheduling, contention, or variable input sizes.
48. Why is end-to-end benchmarking necessary after optimizing individual model operators?
- Individual operator measurements always predict production latency exactly
- End-to-end measurements are needed only for training models
- Data transfer, preprocessing, scheduling, runtime overhead, and postprocessing can affect total request performance
- Optimizing operators guarantees that every application-level metric improves
Answer: C) Data transfer, preprocessing, scheduling, runtime overhead, and postprocessing can affect total request performance
Explanation:
A model's operators are only part of the deployed inference pipeline. End-to-end benchmarks reveal overheads outside the model computation and show whether local optimizations translate into real application-level benefits.
49. An optimized model uses 4-bit weights instead of 16-bit weights, but its inference latency barely improves. What is the most appropriate next step?
- Assume the quantization implementation must be incorrect and discard all measurements
- Profile the execution to check kernel support, dequantization overhead, memory bottlenecks, and whether the workload is actually compute-bound
- Increase the model's parameter count without further analysis
- Measure only the model file size and declare the optimization successful
Answer: B) Profile the execution to check kernel support, dequantization overhead, memory bottlenecks, and whether the workload is actually compute-bound
Explanation:
Reducing weight precision can decrease model storage without delivering a proportional latency improvement. The runtime may use inefficient kernels, perform additional conversions, or be limited by another part of the execution pipeline. Profiling helps identify the actual bottleneck.
50. A production AI service must reduce GPU memory consumption and p99 latency while keeping model quality within an established tolerance. Which optimization strategy is the most appropriate?
- Apply the most aggressive quantization available and deploy without testing
- Increase batch size indefinitely and ignore queueing delays
- Combine representative benchmarking, suitable quantization, runtime and kernel optimization, and controlled batching, then validate quality, memory usage, throughput, and tail latency
- Optimize only the model's file size and assume that runtime performance will improve automatically
Answer: C) Combine representative benchmarking, suitable quantization, runtime and kernel optimization, and controlled batching, then validate quality, memory usage, throughput, and tail latency
Explanation:
Production optimization requires balancing multiple objectives rather than minimizing model size alone. A measured workflow identifies bottlenecks, applies techniques supported by the target hardware, and validates model quality alongside peak memory, throughput, and p99 latency before deployment.