Large Language Models (LLMs) MCQs (Multiple-Choice Questions)

These Large Language Models (LLMs) multiple-choice questions cover fundamental and advanced concepts involved in understanding, developing, fine-tuning, evaluating, and deploying modern language models. The questions cover topics such as Transformer architecture, tokenization, embeddings, attention mechanisms, pretraining, inference, prompting, fine-tuning, parameter-efficient fine-tuning, decoding strategies, context windows, RAG, quantization, model evaluation, and LLM deployment.

Large Language Models (LLMs) MCQs

These LLM MCQs are useful for students, developers, AI/ML engineers, researchers, and professionals preparing for technical interviews or looking to strengthen their understanding of modern language models.

List of Large Language Models (LLMs) MCQs

Below is a list of 50 Large Language Models (LLMs) multiple-choice questions with answers and explanations.

1. What does LLM stand for in the context of artificial intelligence?

  1. Logical Learning Machine
  2. Large Language Model
  3. Language Logic Module
  4. Large Learning Mechanism

Answer: B) Large Language Model

Explanation:

LLM stands for Large Language Model. It refers to a machine learning model trained on large amounts of text data to understand and generate language.

2. Which neural network architecture forms the foundation of most modern LLMs?

  1. Decision Tree
  2. Transformer
  3. Random Forest
  4. Naive Bayes

Answer: B) Transformer

Explanation:

The Transformer architecture introduced self-attention as a central mechanism and has become the dominant architecture for many modern language models.

3. What is the primary purpose of tokenization in an LLM?

  1. To convert text into model-processable tokens
  2. To translate text into another language
  3. To remove all punctuation
  4. To train the neural network

Answer: A) To convert text into model-processable tokens

Explanation:

Tokenization converts input text into tokens, which may represent words, subwords, characters, or other units. These tokens are then mapped to numerical IDs that the model can process.

4. What is an embedding in an LLM?

  1. A compressed text file
  2. A numerical vector representation of tokens or other information
  3. A database table
  4. A decoding algorithm

Answer: B) A numerical vector representation of tokens or other information

Explanation:

Embeddings represent tokens or other information as numerical vectors. These vectors allow neural networks to operate on language information mathematically.

5. What is the main purpose of the self-attention mechanism in a Transformer?

  1. To remove tokens from an input
  2. To determine relationships between tokens in a sequence
  3. To compress the model into a smaller file
  4. To replace tokenization

Answer: B) To determine relationships between tokens in a sequence

Explanation:

Self-attention allows each token representation to incorporate information from other relevant tokens in the sequence, helping the model capture contextual relationships.

6. Which three components are used to calculate scaled dot-product attention?

  1. Input, output, and loss
  2. Query, Key, and Value
  3. Token, vocabulary, and embedding
  4. Encoder, decoder, and optimizer

Answer: B) Query, Key, and Value

Explanation:

Scaled dot-product attention uses Query, Key, and Value representations. Attention scores are calculated from the relationship between queries and keys and then used to weight the values.

7. Why is the dot product in scaled dot-product attention divided by the square root of the key dimension?

  1. To increase vocabulary size
  2. To reduce excessively large attention scores
  3. To remove padding tokens
  4. To increase the number of layers

Answer: B) To reduce excessively large attention scores

Explanation:

Scaling by the square root of the key dimension helps keep attention scores at a suitable magnitude, which improves the behavior of the subsequent softmax operation.

8. What is multi-head attention designed to provide?

  1. Multiple independent tokenizers
  2. Multiple attention representations that can capture different relationships
  3. Multiple datasets during inference
  4. Multiple output files

Answer: B) Multiple attention representations that can capture different relationships

Explanation:

Multi-head attention performs attention through multiple heads. Different heads can learn different patterns or relationships within the input sequence.

9. Why does a Transformer need positional information?

  1. Attention alone does not inherently encode token order
  2. It increases vocabulary size
  3. It removes duplicate tokens
  4. It converts text into images

Answer: A) Attention alone does not inherently encode token order

Explanation:

Because Transformer attention processes token relationships without recurrence, positional information is needed so the model can distinguish different positions in a sequence.

10. What is a causal language model primarily trained to predict?

  1. The previous token only
  2. The next token based on preceding context
  3. The document title only
  4. The vocabulary size

Answer: B) The next token based on preceding context

Explanation:

A causal language model predicts the next token using the tokens that appear before it. Decoder-only LLMs commonly use this objective for autoregressive generation.

11. What is autoregressive generation?

  1. Generating all output tokens independently at the same time
  2. Generating tokens sequentially using previously generated tokens as context
  3. Generating only numerical vectors
  4. Generating training labels randomly

Answer: B) Generating tokens sequentially using previously generated tokens as context

Explanation:

In autoregressive generation, the model predicts a token and then uses that token as part of the context for subsequent predictions.

12. What does the vocabulary of a tokenizer represent?

  1. The maximum number of model layers
  2. The collection of token-to-ID mappings known by the tokenizer
  3. The number of training examples
  4. The number of attention heads

Answer: B) The collection of token-to-ID mappings known by the tokenizer

Explanation:

A tokenizer vocabulary contains the tokens and their associated numerical IDs used to convert text into model inputs.

13. What is a subword tokenizer designed to do?

  1. Represent text only as complete sentences
  2. Represent words using smaller reusable token units
  3. Remove all uncommon words
  4. Convert text directly into images

Answer: B) Represent words using smaller reusable token units

Explanation:

Subword tokenization breaks words into reusable pieces. This allows tokenizers to represent common words efficiently while still handling previously unseen or rare words.

14. What is an attention mask commonly used for?

  1. Controlling which positions can participate in attention
  2. Increasing GPU memory
  3. Changing the model vocabulary
  4. Removing model parameters

Answer: A) Controlling which positions can participate in attention

Explanation:

An attention mask can prevent the model from attending to padding tokens or, in causal models, prevent a token from attending to future positions.

15. What is pretraining in the context of LLMs?

  1. Training a model on a large general-purpose corpus before specialized adaptation
  2. Only testing the final model
  3. Compressing a trained model
  4. Deploying a model to a web server

Answer: A) Training a model on a large general-purpose corpus before specialized adaptation

Explanation:

Pretraining exposes an LLM to a large and diverse dataset so that it learns general language patterns and representations before later adaptation or instruction tuning.

16. Which objective is commonly associated with decoder-only LLM pretraining?

  1. Next-token prediction
  2. Image segmentation
  3. Clustering without labels
  4. Object detection

Answer: A) Next-token prediction

Explanation:

Decoder-only language models commonly use causal language modeling, where the model learns to predict the next token from the preceding context.

17. What does the loss function measure during LLM training?

  1. The size of the tokenizer
  2. The discrepancy between model predictions and target values
  3. The number of GPUs available
  4. The number of tokens in the vocabulary

Answer: B) The discrepancy between model predictions and target values

Explanation:

The training loss quantifies how different the model's predictions are from the expected targets. Optimization algorithms use this signal to update model parameters.

18. Which loss is commonly used for next-token language modeling?

  1. Cross-entropy loss
  2. Mean absolute error only
  3. K-means loss
  4. Hinge geometry loss

Answer: A) Cross-entropy loss

Explanation:

Cross-entropy is commonly used to measure the difference between the probability distribution predicted by a language model and the target token distribution.

19. What is perplexity commonly used to measure for a language model?

  1. GPU temperature
  2. How well a language model predicts a sequence of tokens
  3. The number of model layers
  4. The number of training datasets

Answer: B) How well a language model predicts a sequence of tokens

Explanation:

Perplexity is derived from language-model likelihood or cross-entropy and is commonly used to evaluate how well a model predicts text. Lower perplexity generally indicates better predictive performance on the same evaluation setup.

20. What is fine-tuning?

  1. Training an existing pretrained model further on a more specific dataset
  2. Deleting model parameters
  3. Changing the GPU driver
  4. Increasing the tokenizer vocabulary automatically

Answer: A) Training an existing pretrained model further on a more specific dataset

Explanation:

Fine-tuning adapts a pretrained model to a particular task, domain, style, or instruction-following behavior using additional training data.

21. What is instruction tuning intended to improve?

  1. The model's ability to follow natural-language instructions
  2. The physical size of the GPU
  3. The number of tokenizer files
  4. The operating system kernel

Answer: A) The model's ability to follow natural-language instructions

Explanation:

Instruction tuning trains a pretrained model on examples containing instructions and desired responses, helping it respond more appropriately to user requests.

22. What is Reinforcement Learning from Human Feedback (RLHF) intended to achieve?

  1. Align model behavior with human preferences using feedback
  2. Replace tokenization with reinforcement learning
  3. Remove the Transformer architecture
  4. Increase the vocabulary automatically

Answer: A) Align model behavior with human preferences using feedback

Explanation:

RLHF uses human preference information as part of an alignment process to encourage model outputs that better match desired behaviors.

23. What is parameter-efficient fine-tuning (PEFT)?

  1. A family of methods that adapts a model while updating fewer parameters than full fine-tuning
  2. A method that removes all model parameters
  3. A tokenizer compression algorithm
  4. A GPU scheduling protocol

Answer: A) A family of methods that adapts a model while updating fewer parameters than full fine-tuning

Explanation:

PEFT methods reduce the number of trainable parameters required for adaptation. This can substantially reduce memory and computational requirements compared with updating the entire model.

24. What does LoRA primarily do during model adaptation?

  1. Adds trainable low-rank updates to selected model weights
  2. Replaces the tokenizer with a dictionary
  3. Removes all attention layers
  4. Converts text directly into audio

Answer: A) Adds trainable low-rank updates to selected model weights

Explanation:

Low-Rank Adaptation, or LoRA, introduces trainable low-rank matrices while keeping the original model weights frozen in the typical setup.

25. What is prompt engineering?

  1. Designing input instructions to guide model behavior
  2. Changing the Transformer source code
  3. Replacing the model's GPU
  4. Deleting the training corpus

Answer: A) Designing input instructions to guide model behavior

Explanation:

Prompt engineering involves designing instructions, context, examples, constraints, and output requirements to improve the usefulness and reliability of an LLM's responses.

26. What is zero-shot prompting?

  1. Providing no task-specific examples in the prompt
  2. Providing exactly zero tokens
  3. Training a model without parameters
  4. Removing all model weights

Answer: A) Providing no task-specific examples in the prompt

Explanation:

Zero-shot prompting asks a pretrained or instruction-tuned model to perform a task using instructions without providing task-specific demonstrations.

27. What is few-shot prompting?

  1. Providing a small number of examples within the prompt
  2. Training the model for only a few seconds
  3. Using only a few model parameters
  4. Using fewer than two tokens

Answer: A) Providing a small number of examples within the prompt

Explanation:

Few-shot prompting provides a small number of input-output examples that demonstrate the desired task or response pattern.

28. What does the context window of an LLM define?

  1. The maximum amount of supported input and relevant generated context under the model's token limit
  2. The number of GPUs used during training
  3. The vocabulary size only
  4. The number of fine-tuning datasets

Answer: A) The maximum amount of supported input and relevant generated context under the model's token limit

Explanation:

A model's context window specifies how many tokens can be handled within a single model context according to that model's supported limits.

29. What is temperature used for during text generation?

  1. Controlling the randomness of token sampling
  2. Controlling GPU physical temperature
  3. Changing the vocabulary size
  4. Changing the model architecture

Answer: A) Controlling the randomness of token sampling

Explanation:

Temperature modifies the distribution of token probabilities during sampling. Higher values generally produce more varied outputs, while lower values make sampling more concentrated around high-probability tokens.

30. What is top-k sampling?

  1. Sampling only from the k highest-probability candidate tokens
  2. Selecting the k largest model layers
  3. Keeping only k training examples
  4. Using exactly k attention heads

Answer: A) Sampling only from the k highest-probability candidate tokens

Explanation:

Top-k sampling restricts the candidate set for the next token to the k tokens with the highest probabilities before sampling.

31. What is top-p sampling?

  1. Sampling from the smallest set of tokens whose cumulative probability reaches a specified threshold
  2. Selecting the p-th model layer
  3. Keeping only p training batches
  4. Selecting tokens based only on length

Answer: A) Sampling from the smallest set of tokens whose cumulative probability reaches a specified threshold

Explanation:

Top-p, also called nucleus sampling, dynamically selects a set of high-probability tokens whose cumulative probability reaches the configured threshold.

32. What happens in greedy decoding?

  1. The highest-probability next token is selected at each generation step
  2. A random token is always selected
  3. All possible sequences are generated
  4. The model performs another training epoch

Answer: A) The highest-probability next token is selected at each generation step

Explanation:

Greedy decoding selects the token with the highest probability at each step rather than sampling from a probability distribution.

33. What are logits in an LLM output?

  1. Raw prediction scores for possible output tokens
  2. Compressed model files
  3. Tokenized input strings
  4. Training datasets

Answer: A) Raw prediction scores for possible output tokens

Explanation:

Logits are unnormalized scores produced by the model for possible output tokens. A softmax operation can convert them into a probability distribution.

34. What does softmax typically produce from language-model logits?

  1. A probability distribution over candidate tokens
  2. A new tokenizer
  3. A hidden layer
  4. A training dataset

Answer: A) A probability distribution over candidate tokens

Explanation:

Softmax transforms a set of scores into non-negative values that sum to one, allowing them to be interpreted as probabilities over candidate tokens.

35. What is RAG in an LLM application?

  1. Retrieval-Augmented Generation
  2. Recursive AI Generation
  3. Randomized Attention Generation
  4. Runtime Algorithmic Grouping

Answer: A) Retrieval-Augmented Generation

Explanation:

RAG combines information retrieval with generation. Relevant external information is retrieved and supplied to the model as context before generating a response.

36. Why are embeddings commonly used in RAG systems?

  1. To represent text as vectors for similarity-based retrieval
  2. To increase the number of Transformer layers
  3. To replace the language model entirely
  4. To remove all document metadata

Answer: A) To represent text as vectors for similarity-based retrieval

Explanation:

Embedding models convert text into vectors. These vectors can be compared to identify semantically similar documents or chunks for retrieval.

37. What is hallucination in an LLM?

  1. When a model generates information that is unsupported, incorrect, or fabricated
  2. When a model refuses every request
  3. When tokenization fails completely
  4. When a GPU shuts down

Answer: A) When a model generates information that is unsupported, incorrect, or fabricated

Explanation:

LLMs can generate fluent responses that are not supported by reliable evidence. Such unsupported or fabricated outputs are commonly referred to as hallucinations.

38. Which technique can help provide an LLM with information from a private document collection without retraining the base model?

  1. RAG
  2. Changing the activation function
  3. Increasing vocabulary size
  4. Removing attention layers

Answer: A) RAG

Explanation:

RAG can retrieve relevant information from a private document collection at inference time and provide it to the LLM as contextual input.

39. What is quantization used for when deploying LLMs?

  1. Representing model values using lower numerical precision to reduce memory and computation requirements
  2. Increasing the number of training examples
  3. Adding more Transformer layers
  4. Replacing text with images

Answer: A) Representing model values using lower numerical precision to reduce memory and computation requirements

Explanation:

Quantization represents model parameters or computations using lower-precision numerical formats. This can reduce memory usage and may improve inference efficiency, although quality and compatibility depend on the method.

40. What is KV caching used for during autoregressive LLM inference?

  1. Reusing previously computed key and value representations
  2. Storing the entire training dataset
  3. Changing the tokenizer vocabulary
  4. Fine-tuning the model during generation

Answer: A) Reusing previously computed key and value representations

Explanation:

KV caching stores previously computed attention key and value representations so they do not need to be recomputed for every newly generated token.

41. What is inference in the context of an LLM?

  1. Using a trained model to produce predictions or generated outputs
  2. Training the model from random weights
  3. Creating a tokenizer vocabulary
  4. Collecting the training corpus

Answer: A) Using a trained model to produce predictions or generated outputs

Explanation:

Inference is the process of running a trained model on input data to obtain predictions, representations, or generated outputs.

42. Which factor directly affects the amount of GPU memory required to load an LLM?

  1. Model parameter count and numerical precision
  2. Only the prompt's punctuation
  3. The user's operating system wallpaper
  4. The number of web pages in a browser

Answer: A) Model parameter count and numerical precision

Explanation:

Larger models require more memory to store their parameters. Higher numerical precision generally requires more memory per parameter than lower-precision representations.

43. What is batch size in LLM training or inference?

  1. The number of examples processed together in one batch
  2. The number of model layers
  3. The number of vocabulary tokens
  4. The number of attention heads

Answer: A) The number of examples processed together in one batch

Explanation:

Batch size specifies how many examples or sequences are processed together during a model operation. Its practical effect depends on whether the workload is training or inference.

44. What is gradient descent used for when training an LLM?

  1. Updating model parameters to reduce the training loss
  2. Tokenizing the input text
  3. Retrieving documents from a vector database
  4. Generating a user interface

Answer: A) Updating model parameters to reduce the training loss

Explanation:

Gradient-based optimization computes parameter gradients from the loss and uses them to update model weights in an attempt to minimize the training objective.

45. What is an optimizer such as Adam used for during neural-network training?

  1. Updating model parameters based on gradients
  2. Splitting text into tokens
  3. Creating HTML pages
  4. Searching the internet

Answer: A) Updating model parameters based on gradients

Explanation:

Optimizers determine how model parameters should be updated using gradient information. Adam is a widely used adaptive optimization algorithm for neural networks.

46. Which approach is generally appropriate when a company wants an LLM to answer questions using its frequently changing internal documents?

  1. RAG with an updated document retrieval system
  2. Training the model from scratch every day
  3. Removing the model's context window
  4. Replacing all tokens with numbers manually

Answer: A) RAG with an updated document retrieval system

Explanation:

RAG allows an application to retrieve current information from an external knowledge source at inference time, avoiding the need to retrain the entire model whenever documents change.

47. Which metric is particularly useful for evaluating the factual accuracy of an LLM application that answers questions from retrieved documents?

  1. Groundedness or faithfulness to the provided evidence
  2. GPU clock speed only
  3. Tokenizer vocabulary size only
  4. Number of Transformer layers only

Answer: A) Groundedness or faithfulness to the provided evidence

Explanation:

For retrieval-based applications, evaluating whether the generated answer is supported by the retrieved evidence is important in addition to general language quality.

48. A developer wants to reduce LLM response latency while generating tokens sequentially. Which optimization can help reuse computations from earlier generated tokens?

  1. KV caching
  2. Increasing the vocabulary indefinitely
  3. Retraining the tokenizer for every request
  4. Disabling all model weights

Answer: A) KV caching

Explanation:

KV caching stores previously calculated attention key and value states during autoregressive generation. Reusing these states avoids repeating some computations for earlier tokens.

49. A team wants to adapt a pretrained LLM to a specialized customer-support style but has limited GPU memory. Which approach is particularly suitable?

  1. Parameter-efficient fine-tuning such as LoRA
  2. Training a new LLM from scratch
  3. Increasing the model size without changing hardware
  4. Deleting the pretrained weights

Answer: A) Parameter-efficient fine-tuning such as LoRA

Explanation:

PEFT methods such as LoRA can adapt a pretrained model while training a relatively small number of additional parameters, making them useful when compute or memory resources are constrained.

50. A developer is building an LLM application that must answer questions from a private knowledge base, cite the retrieved documents, and avoid inventing unsupported information. Which architecture is most appropriate?

  1. RAG with document retrieval, contextual prompting, citation generation, and answer-grounding evaluation
  2. Greedy decoding without access to external information
  3. Increasing the temperature and removing retrieved context
  4. Training a tokenizer without a language model

Answer: A) RAG with document retrieval, contextual prompting, citation generation, and answer-grounding evaluation

Explanation:

A RAG architecture can retrieve relevant private documents and provide them as context to the LLM. The application can then require citations and evaluate whether the generated response is supported by the retrieved evidence. This combines retrieval, generation, and evaluation rather than relying solely on the model's pretrained knowledge.

Advertisement
Advertisement

Comments and Discussions!

Load comments ↻


Advertisement
Advertisement
Advertisement

Copyright © 2026 www.includehelp.com. All rights reserved.