×

Trending Technologies MCQs

Small Language Models (SLMs) MCQs (Multiple-Choice Questions)

Practice Small Language Models (SLMs) MCQs to test your knowledge of compact language models, efficient transformer architectures, model compression, and lightweight AI deployment. These questions cover parameter counts, tokenization, quantization, knowledge distillation, fine-tuning, inference optimization, and on-device language processing. They are useful for AI developers, machine learning practitioners, researchers, and students preparing for technical interviews or assessments. The set includes both foundational and practical questions covering modern small language model systems.

Small Language Models (SLMs) MCQs

These Small Language Models multiple-choice questions cover important concepts such as transformer architectures, model size, attention mechanisms, pretraining, supervised fine-tuning, parameter-efficient adaptation, quantization, knowledge distillation, retrieval-augmented generation, and edge deployment. This set combines conceptual, technical, and scenario-based questions to help test your understanding of small language model systems.

Small Language Models MCQs cover the architectures and techniques used to develop, adapt, optimize, and deploy compact language models under limited computational and memory resources. Each question includes an answer and explanation.

List of Small Language Models (SLMs) MCQs

The following Small Language Models multiple-choice questions cover model architecture, training, compression, inference, evaluation, privacy, and practical deployment scenarios.

1. What is a Small Language Model (SLM)?

  1. A database that stores only short text documents
  2. A language model designed with relatively compact computational and parameter requirements
  3. A model that can process only one word at a time
  4. A system that performs no machine learning computations

Answer: B) A language model designed with relatively compact computational and parameter requirements

Explanation:

An SLM is a language model designed to operate with comparatively modest model size, memory, or compute requirements. There is no universally fixed parameter-count threshold separating small models from large language models; the classification depends on context and deployment needs.

2. What is a major advantage of deploying an SLM instead of a much larger language model?

  1. It guarantees superior reasoning on every task
  2. It eliminates the need for model evaluation
  3. It always supports every language equally well
  4. It can reduce inference cost, memory requirements, and deployment latency

Answer: D) It can reduce inference cost, memory requirements, and deployment latency

Explanation:

Compact models can require fewer computational resources and may run on local devices or less expensive servers. Their actual latency and cost advantages depend on architecture, hardware, workload, and serving configuration.

3. Which neural network architecture is commonly used to build modern autoregressive small language models?

  1. Transformer decoder
  2. Decision tree
  3. K-means clustering
  4. Relational database engine

Answer: A) Transformer decoder

Explanation:

Many modern SLMs use decoder-only transformer architectures to predict the next token in a sequence. Other language-model designs, including encoder-only and encoder-decoder transformers, are also used for different language tasks.

4. What does the parameter count of a language model generally indicate?

  1. The number of documents in its training dataset
  2. The number of tokens it can generate in every request
  3. The number of learned numerical parameters in the model
  4. The number of users who can access the model

Answer: C) The number of learned numerical parameters in the model

Explanation:

Model parameters include learned weights and, where applicable, biases and other trainable values. Parameter count influences model storage and computation but does not independently determine intelligence, accuracy, or real-world performance.

5. Why can a well-trained SLM outperform a larger model on a narrowly defined task?

  1. Smaller models automatically contain more knowledge
  2. The SLM may be trained or fine-tuned with high-quality, task-specific data
  3. Larger models cannot process domain-specific terminology
  4. Parameter count has no relationship to model capacity

Answer: B) The SLM may be trained or fine-tuned with high-quality, task-specific data

Explanation:

A compact model can perform strongly on a specific task when its training data, objective, and adaptation process match the application. However, specialization does not guarantee that it will outperform larger models on broader tasks.

6. What is tokenization in an SLM?

  1. Converting a trained model into a smaller physical device
  2. Removing all punctuation from every generated answer
  3. Converting model weights into floating-point values
  4. Splitting input text into tokens that the model can process

Answer: D) Splitting input text into tokens that the model can process

Explanation:

Tokenization maps text into units such as subwords, characters, or other token types. The tokenizer's vocabulary and segmentation rules affect sequence length, memory use, and how efficiently the model processes different languages and domains.

7. What is the purpose of an embedding layer in a language model?

  1. To map token identifiers to learned vector representations
  2. To convert every output into a database record
  3. To eliminate the attention mechanism
  4. To store the complete training corpus inside each input token

Answer: A) To map token identifiers to learned vector representations

Explanation:

An embedding layer maps discrete token IDs into dense numerical vectors. These representations provide the continuous inputs used by the model's subsequent neural network layers.

8. What is the primary purpose of self-attention in a transformer-based SLM?

  1. To remove every token that appears more than once
  2. To guarantee correct factual responses
  3. To allow token representations to incorporate information from relevant positions in the sequence
  4. To convert text into images

Answer: C) To allow token representations to incorporate information from relevant positions in the sequence

Explanation:

Self-attention computes relationships among token representations and uses attention weights to combine information. In a causal language model, the attention mask prevents a position from attending to future tokens during next-token prediction.

9. What is the purpose of causal masking in an autoregressive SLM?

  1. To prevent the model from using token embeddings
  2. To prevent each token position from attending to future positions during next-token training and generation
  3. To remove all previously generated tokens from context
  4. To reduce every input sequence to one token

Answer: B) To prevent each token position from attending to future positions during next-token training and generation

Explanation:

Causal masking preserves the autoregressive property by restricting attention to the current and preceding positions. Without this restriction during standard next-token training, a model could access information that should not be available when predicting the next token.

10. What does context length represent in a language model?

  1. The number of model parameters
  2. The size of the model's training team
  3. The number of GPUs required for every deployment
  4. The number of tokens the model can consider within its supported input context

Answer: D) The number of tokens the model can consider within its supported input context

Explanation:

Context length defines the token window available to the model for a particular operation. Longer contexts can increase memory and computational requirements, and the usable context depends on the model architecture and deployment implementation.

11. What is pretraining in the development of an SLM?

  1. Training a model on a broad dataset to learn general language patterns or other representations
  2. Evaluating only the final production interface
  3. Compressing the model without learning any parameters
  4. Running a model without providing text inputs

Answer: A) Training a model on a broad dataset to learn general language patterns or other representations

Explanation:

Pretraining develops general capabilities using a large collection of training examples and an objective such as next-token prediction. The resulting model can then be evaluated, fine-tuned, or adapted for particular applications.

12. What is supervised fine-tuning (SFT) commonly used for?

  1. Deleting the pretrained model's learned parameters
  2. Increasing the vocabulary size without changing any model configuration
  3. Adapting a pretrained model using examples containing desired inputs and outputs
  4. Replacing training examples with random hardware measurements

Answer: C) Adapting a pretrained model using examples containing desired inputs and outputs

Explanation:

SFT trains a pretrained model on curated input-output examples to improve instruction following, formatting, domain behavior, or other target capabilities. Data quality and diversity strongly influence the resulting behavior.

13. What is instruction tuning in an SLM?

  1. Reducing every instruction to a single character
  2. Fine-tuning the model on examples that teach it to respond appropriately to natural-language instructions
  3. Increasing the number of physical processors without retraining
  4. Removing all user prompts from model inputs

Answer: B) Fine-tuning the model on examples that teach it to respond appropriately to natural-language instructions

Explanation:

Instruction tuning exposes a model to prompts paired with appropriate responses. It helps the model learn how to follow a range of tasks expressed through natural language, although the quality of instruction following depends on the data and training process.

14. What is knowledge distillation in small language model development?

  1. Copying a teacher model's entire training dataset into the deployment device
  2. Increasing the student model's parameter count without changing its training
  3. Removing all output probabilities from the teacher model
  4. Training a smaller student model using information or outputs provided by a larger teacher model

Answer: D) Training a smaller student model using information or outputs provided by a larger teacher model

Explanation:

Knowledge distillation transfers useful behavior from a teacher to a student model. Training may use teacher-generated responses, logits, softened probability distributions, or intermediate representations, depending on the distillation method.

15. What is the purpose of temperature scaling in knowledge distillation?

  1. To soften the teacher's output probability distribution and reveal relative probabilities among alternative tokens or classes
  2. To physically cool the GPU during training
  3. To increase the student's context length automatically
  4. To remove all low-probability tokens from the training dataset

Answer: A) To soften the teacher's output probability distribution and reveal relative probabilities among alternative tokens or classes

Explanation:

In distillation, a higher softmax temperature produces a softer probability distribution. These distributions can provide richer training signals than hard labels alone, helping the student learn similarities among alternative outputs.

16. What is parameter-efficient fine-tuning (PEFT)?

  1. A technique that requires retraining every model parameter in every case
  2. A method that eliminates the need for training examples
  3. A family of methods that adapts a model by updating a relatively small subset of parameters or additional components
  4. A method for increasing the number of output tokens without computation

Answer: C) A family of methods that adapts a model by updating a relatively small subset of parameters or additional components

Explanation:

PEFT methods reduce the number of parameters updated during fine-tuning. They can lower optimizer-state memory and training costs, making adaptation more practical when computational resources are limited.

17. What is LoRA commonly used for when adapting an SLM?

  1. Replacing the tokenizer with a database index
  2. Learning low-rank weight updates while typically keeping the original model weights frozen
  3. Increasing every weight matrix to full rank
  4. Removing the need for pretrained weights

Answer: B) Learning low-rank weight updates while typically keeping the original model weights frozen

Explanation:

Low-Rank Adaptation (LoRA) represents parameter updates using smaller trainable matrices. It can significantly reduce the trainable parameter count compared with full fine-tuning, while preserving the base model for reuse across tasks.

18. What is quantization in small language models?

  1. Increasing the precision of every parameter without changing its representation
  2. Converting all generated tokens into integers without modifying model computation
  3. Expanding every layer to increase model capacity
  4. Representing model weights or activations with lower-precision numerical formats

Answer: D) Representing model weights or activations with lower-precision numerical formats

Explanation:

Quantization can reduce model memory requirements and sometimes improve inference speed. Common formats include INT8 and lower-bit representations, although quality and performance depend on the quantization method, model, and hardware.

19. Why is 4-bit quantization attractive for deploying SLMs on memory-constrained devices?

  1. It can substantially reduce the storage required for model weights compared with 16-bit or 32-bit representations
  2. It guarantees zero loss in model accuracy
  3. It eliminates the need for RAM during inference
  4. It makes all model operations execute in a single processor cycle

Answer: A) It can substantially reduce the storage required for model weights compared with 16-bit or 32-bit representations

Explanation:

Four-bit quantization uses fewer bits per quantized weight than higher-precision formats. Additional memory is needed for scales, metadata, runtime buffers, and activations, so total memory usage is greater than the raw weight-bit calculation alone.

20. What is post-training quantization?

  1. Training the model exclusively with integer-valued input text
  2. Increasing model precision after every inference request
  3. Quantizing a trained model without performing quantization-aware retraining
  4. Removing every model parameter after pretraining

Answer: C) Quantizing a trained model without performing quantization-aware retraining

Explanation:

Post-training quantization is applied after a model has been trained. Depending on the method, it may use calibration data to determine quantization ranges and can be simpler than conducting quantization-aware training.

21. What is the primary purpose of quantization-aware training?

  1. To ensure every layer uses FP32 permanently
  2. To expose training to simulated quantization effects so the model can adapt to reduced numerical precision
  3. To remove the need for an optimization objective
  4. To convert the training dataset into a compressed archive

Answer: B) To expose training to simulated quantization effects so the model can adapt to reduced numerical precision

Explanation:

Quantization-aware training simulates the effects of quantization during training or fine-tuning. This can help preserve task quality when converting the resulting model to supported lower-precision formats.

22. What is structured pruning in an SLM?

  1. Removing only words from the tokenizer vocabulary without changing model structure
  2. Deleting every attention layer regardless of its role
  3. Increasing the number of hidden dimensions
  4. Removing organized model components, such as attention heads, channels, or neurons

Answer: D) Removing organized model components, such as attention heads, channels, or neurons

Explanation:

Structured pruning removes entire groups of parameters or computations. Because it can create smaller dense operations, it may translate into practical speed improvements more readily than arbitrary sparse weights on hardware without sparse-computation support.

23. What is the main benefit of using an SLM for on-device inference?

  1. It can process supported tasks locally, potentially reducing network dependence and external data transmission
  2. It guarantees that the device never consumes battery power
  3. It removes the need to update the model or application
  4. It guarantees better answers than all cloud-hosted models

Answer: A) It can process supported tasks locally, potentially reducing network dependence and external data transmission

Explanation:

Local inference can improve responsiveness and support operation without a continuous network connection. It may also reduce data sharing with remote services, but privacy still depends on application logging, local storage, permissions, and data-handling practices.

24. Which factor is especially important when selecting an SLM for a smartphone?

  1. The number of data centers owned by the model developer
  2. The number of pages in the model's documentation
  3. Available RAM, model storage size, processor support, latency, and power consumption
  4. The physical size of the training dataset's archive label

Answer: C) Available RAM, model storage size, processor support, latency, and power consumption

Explanation:

Mobile deployment is constrained by memory, compute capacity, thermal limits, and battery life. Model size alone is insufficient because runtime buffers, context length, and concurrent application activity also influence resource requirements.

25. What is the purpose of a model runtime in SLM deployment?

  1. To create the original training dataset automatically
  2. To load the model and execute its operations on supported hardware
  3. To replace every model parameter with a random value
  4. To eliminate the need for input tokenization

Answer: B) To load the model and execute its operations on supported hardware

Explanation:

A model runtime manages model execution, tensor operations, memory, and supported hardware acceleration. The selected runtime can significantly affect compatibility, performance, and resource consumption.

26. What is a key benefit of using a dedicated neural processing unit (NPU) for SLM inference?

  1. It guarantees unlimited model memory
  2. It removes all numerical differences between model formats
  3. It makes model evaluation unnecessary
  4. It can accelerate supported neural network operations with potentially lower energy consumption

Answer: D) It can accelerate supported neural network operations with potentially lower energy consumption

Explanation:

NPUs are designed to accelerate supported machine learning workloads. The benefits depend on operator compatibility, precision support, memory movement, hardware design, and the model execution framework.

27. What is the primary purpose of retrieval-augmented generation (RAG) when used with an SLM?

  1. To provide relevant external information to the model at inference time
  2. To increase the model's parameter count automatically
  3. To eliminate the need for document retrieval
  4. To convert every retrieved document into a model weight

Answer: A) To provide relevant external information to the model at inference time

Explanation:

RAG retrieves relevant documents or passages and supplies them as context for generation. It can help a compact model answer domain-specific questions without requiring all information to be stored in its parameters, although retrieval quality and context limits remain important.

28. Why can retrieval-augmented generation be useful for a domain-specific SLM?

  1. It guarantees that retrieved information is always correct
  2. It eliminates the need to evaluate the generated answer
  3. It can supply updated or specialized knowledge without retraining the entire model for every content change
  4. It removes the need for a tokenizer

Answer: C) It can supply updated or specialized knowledge without retraining the entire model for every content change

Explanation:

Knowledge can be updated in the retrieval corpus independently of the model weights. The system still needs appropriate document indexing, access controls, retrieval evaluation, and safeguards against irrelevant or misleading source material.

29. What is the purpose of prompt templates when deploying an instruction-tuned SLM?

  1. To increase the number of learned model parameters
  2. To structure system instructions, user input, and contextual information in the expected format
  3. To replace model inference with a static database query
  4. To guarantee correct output regardless of the prompt content

Answer: B) To structure system instructions, user input, and contextual information in the expected format

Explanation:

Prompt templates organize inputs according to the model's expected instruction or chat format. Correct formatting can improve reliability, while malformed templates may produce unexpected behavior or poorer responses.

30. What is the purpose of a chat template in a conversational SLM?

  1. To define the physical layout of the inference server
  2. To replace all conversation messages with random tokens
  3. To ensure that every model uses identical token IDs and control tokens
  4. To serialize messages and role information into the token sequence expected by the model

Answer: D) To serialize messages and role information into the token sequence expected by the model

Explanation:

Chat templates specify how roles, messages, and control markers are represented for a particular model. Templates differ across model families, so using the correct format is important for instruction following and turn management.

31. What does perplexity measure in language modeling?

  1. How well a language model predicts a sequence of tokens, commonly derived from average negative log-likelihood
  2. The physical memory capacity of a GPU
  3. The number of users who have downloaded the model
  4. The total number of layers in a transformer

Answer: A) How well a language model predicts a sequence of tokens, commonly derived from average negative log-likelihood

Explanation:

Perplexity is an intrinsic language-model evaluation metric derived from token prediction probabilities. Lower perplexity generally indicates better predictive performance on the evaluated text, but it does not directly guarantee better instruction following or real-world task performance.

32. Why should SLM evaluation include task-specific benchmarks in addition to perplexity?

  1. Perplexity directly measures every safety and reasoning capability
  2. Task-specific benchmarks are unnecessary for pretrained models
  3. Perplexity alone does not fully measure instruction following, reasoning, factual accuracy, or application-specific success
  4. Perplexity can be calculated only for image classification

Answer: C) Perplexity alone does not fully measure instruction following, reasoning, factual accuracy, or application-specific success

Explanation:

A model can achieve good next-token prediction performance without meeting the requirements of a particular application. Task-specific tests help assess whether the model follows instructions, produces correct outputs, and meets the deployment's quality standards.

33. What is catastrophic forgetting during fine-tuning?

  1. The model automatically increases its context length
  2. The model loses some previously learned capabilities while adapting to new training data
  3. The tokenizer becomes independent of the model vocabulary
  4. The inference runtime permanently removes all weights

Answer: B) The model loses some previously learned capabilities while adapting to new training data

Explanation:

Catastrophic forgetting can occur when fine-tuning shifts model parameters toward a new task at the expense of earlier capabilities. Data mixing, parameter-efficient adaptation, careful learning rates, and evaluation across multiple tasks can help identify or mitigate this problem.

34. Why is high-quality training data especially important when developing a compact language model?

  1. It guarantees that the model requires no computation
  2. It ensures the model can answer every possible question
  3. It eliminates the need for validation
  4. It helps limited model capacity learn useful patterns and task-relevant behavior

Answer: D) It helps limited model capacity learn useful patterns and task-relevant behavior

Explanation:

Compact models have fewer parameters or computational resources available than many larger alternatives. Well-curated, representative data can help them learn useful behaviors, while noisy, duplicated, or poorly matched data may waste training resources.

35. What is synthetic data generation in SLM development?

  1. Creating training examples using automated methods, including teacher models or programmatic generators
  2. Collecting only physical sensor readings without converting them into examples
  3. Deleting every training example that contains text
  4. Generating model weights without any computational process

Answer: A) Creating training examples using automated methods, including teacher models or programmatic generators

Explanation:

Synthetic data can provide additional instruction, reasoning, or domain-specific examples. It should be checked for correctness, diversity, bias, duplication, and teacher-generated errors before being used for training.

36. What is a major risk of training an SLM primarily on low-quality synthetic data?

  1. The model always becomes larger than its teacher
  2. The tokenizer automatically learns every language
  3. The model may learn incorrect patterns, repetitive outputs, or limitations inherited from the generated data
  4. The model becomes independent of its training objective

Answer: C) The model may learn incorrect patterns, repetitive outputs, or limitations inherited from the generated data

Explanation:

Synthetic examples can contain factual errors, stylistic repetition, or systematic biases. Data filtering, independent verification, diversity controls, and evaluation on trusted examples help reduce these risks.

37. What is the purpose of a system prompt in an instruction-following SLM?

  1. To permanently rewrite the model's pretrained weights for every request
  2. To provide high-level instructions that guide the model's behavior within the conversation
  3. To guarantee that the model cannot generate incorrect information
  4. To replace the need for input tokens

Answer: B) To provide high-level instructions that guide the model's behavior within the conversation

Explanation:

A system prompt provides behavioral or task-level guidance to the model. It influences generation but is not a substitute for robust access controls, output validation, or other application-level safety measures.

38. What is the purpose of constrained decoding in an SLM application?

  1. To increase model size during generation
  2. To make the model ignore all output requirements
  3. To eliminate the tokenizer from the inference pipeline
  4. To restrict generated tokens or outputs according to specified rules or a grammar

Answer: D) To restrict generated tokens or outputs according to specified rules or a grammar

Explanation:

Constrained decoding limits the generation process to outputs that satisfy defined constraints, such as a JSON schema or formal grammar. It can improve structural validity but does not guarantee that the generated content is factually correct.

39. Why might an SLM produce repetitive text during generation?

  1. Because decoding settings, training limitations, or model behavior can encourage repeated token sequences
  2. Because every language model is required to repeat each sentence
  3. Because quantization prevents all new tokens from being generated
  4. Because tokenization removes all semantic information in every case

Answer: A) Because decoding settings, training limitations, or model behavior can encourage repeated token sequences

Explanation:

Repetition may arise from model limitations, prompt design, or decoding choices. Repetition penalties, suitable sampling settings, improved training data, and application-specific stopping criteria can help, but their effects should be tested rather than assumed.

40. What is the role of temperature during text generation?

  1. To control the physical temperature of the inference processor
  2. To determine the model's parameter count
  3. To adjust the sharpness of the next-token probability distribution used for sampling
  4. To determine the number of transformer layers

Answer: C) To adjust the sharpness of the next-token probability distribution used for sampling

Explanation:

Lower sampling temperatures make the distribution more concentrated around higher-probability tokens, while higher temperatures generally make it flatter. Temperature changes generation behavior but does not change the model's learned weights.

41. What does top-p sampling do during SLM text generation?

  1. Always selects the lowest-probability token
  2. Samples from a set of high-probability tokens whose cumulative probability reaches a specified threshold
  3. Limits the model to exactly p parameters
  4. Removes all tokens that occur in the prompt

Answer: B) Samples from a set of high-probability tokens whose cumulative probability reaches a specified threshold

Explanation:

Top-p, or nucleus sampling, dynamically selects a candidate set based on cumulative probability. The method balances diversity and probability concentration during generation, with results influenced by the selected threshold and other decoding settings.

42. What is KV caching used for during autoregressive SLM inference?

  1. To store every conversation permanently in the model weights
  2. To eliminate all attention computations for new tokens
  3. To replace the token embedding layer
  4. To reuse previously computed attention keys and values instead of recomputing them for the entire prefix at each generation step

Answer: D) To reuse previously computed attention keys and values instead of recomputing them for the entire prefix at each generation step

Explanation:

KV caching avoids recomputing attention keys and values for previously processed tokens during incremental generation. The cache speeds up generation but consumes memory that grows with context length and the number of active sequences.

43. What is speculative decoding intended to improve?

  1. Generation speed by using a draft model to propose tokens that a target model verifies
  2. Training-data collection by removing all model inference
  3. Parameter count by adding extra transformer layers
  4. Output accuracy by guaranteeing that every draft token is accepted

Answer: A) Generation speed by using a draft model to propose tokens that a target model verifies

Explanation:

Speculative decoding can reduce the number of sequential target-model decoding steps by verifying draft tokens in parallel or grouped operations. A compatible verification algorithm can preserve the target model's intended sampling distribution, although speed gains depend on draft quality and acceptance rates.

44. Why can batching improve SLM inference throughput?

  1. It removes every memory allocation from the runtime
  2. It guarantees the same latency for all requests
  3. It allows compatible requests to share hardware execution and improve resource utilization
  4. It eliminates the need to load model weights

Answer: C) It allows compatible requests to share hardware execution and improve resource utilization

Explanation:

Batching can improve throughput by increasing the amount of work performed in each device operation. However, it can also increase memory usage and waiting time, so batch size should be selected according to latency and capacity requirements.

45. What is the main difference between an SLM and a traditional rule-based chatbot?

  1. An SLM cannot generate new sentences
  2. An SLM generates outputs using learned statistical patterns in model parameters, while a rule-based chatbot follows explicitly programmed rules and flows
  3. A rule-based chatbot must always use a GPU
  4. An SLM can only respond using fixed templates

Answer: B) An SLM generates outputs using learned statistical patterns in model parameters, while a rule-based chatbot follows explicitly programmed rules and flows

Explanation:

An SLM uses learned representations to generate context-dependent text. Rule-based systems follow predefined logic and templates, although practical applications may combine language models with deterministic rules and external tools.

46. Which consideration is most important when deploying an SLM in an offline industrial environment?

  1. Whether the model can access the public internet continuously
  2. Whether the model has the largest parameter count available
  3. Whether the system can generate an unlimited number of tokens without resource constraints
  4. Whether the model and runtime fit local hardware, operate reliably offline, and meet the application's accuracy and safety requirements

Answer: D) Whether the model and runtime fit local hardware, operate reliably offline, and meet the application's accuracy and safety requirements

Explanation:

Offline deployments need compatible model files, local dependencies, sufficient compute and memory, and reliable application behavior without remote services. Validation should reflect real operational conditions and the consequences of incorrect outputs.

47. Why should a small language model be evaluated for hallucinations?

  1. To determine whether it generates unsupported or incorrect information presented as fact
  2. To measure only its model file size
  3. To calculate the number of processors in the deployment system
  4. To guarantee that every generated answer is deterministic

Answer: A) To determine whether it generates unsupported or incorrect information presented as fact

Explanation:

SLMs can generate plausible but incorrect statements. Evaluation with representative factual tasks, grounded generation, retrieval, uncertainty handling, and suitable output checks can help reduce the impact of hallucinations.

48. What is a major privacy advantage of local SLM inference?

  1. It guarantees that application data is never stored anywhere
  2. It eliminates the need for device security
  3. It can keep prompts and generated outputs on the device instead of sending them to an external inference service
  4. It guarantees that the model cannot disclose sensitive information

Answer: C) It can keep prompts and generated outputs on the device instead of sending them to an external inference service

Explanation:

Local inference can reduce the need to transmit sensitive prompts to remote servers. However, application logs, backups, malicious software, model behavior, and device access still need appropriate privacy and security controls.

49. An SLM fits within the available device memory but generates responses too slowly. What should an engineer do first?

  1. Increase the number of model layers without collecting performance data
  2. Profile prefill and token-generation latency, inspect hardware utilization, and test supported runtime and quantization optimizations
  3. Remove all evaluation tests from the development process
  4. Increase the context length regardless of the workload

Answer: B) Profile prefill and token-generation latency, inspect hardware utilization, and test supported runtime and quantization optimizations

Explanation:

Slow generation may be caused by memory bandwidth, inefficient kernels, unsupported precision, excessive context length, or runtime overhead. Profiling helps identify the actual bottleneck before applying optimizations and verifying that quality remains acceptable.

50. A company wants to deploy an SLM that answers questions about internal technical documentation on employee laptops without sending confidential documents to an external service. Which design is most appropriate?

  1. Use the largest cloud-hosted model for every request and upload all internal documents
  2. Store the documentation in the model's prompt permanently without controlling access
  3. Run a compatible quantized SLM locally, retrieve relevant documents from an access-controlled local index, and evaluate answer quality, resource usage, and data handling
  4. Remove all source documents after indexing and assume that retrieval will always return correct answers

Answer: C) Run a compatible quantized SLM locally, retrieve relevant documents from an access-controlled local index, and evaluate answer quality, resource usage, and data handling

Explanation:

A locally deployed SLM combined with retrieval-augmented generation can answer questions using internal documents without routinely transmitting those documents to an external inference service. Quantization can help the model fit laptop resources, while access-controlled retrieval, source grounding, quality evaluation, and local data-security controls help address confidentiality and reliability requirements.

Comments and Discussions!

Load comments ↻



Copyright © 2026 www.includehelp.com. All rights reserved.