Home »
Trending Technologies MCQs
Transfer Learning MCQs (Multiple-Choice Questions)
Practice Transfer Learning MCQs to test your understanding of machine learning techniques that reuse knowledge from previously trained models to solve new tasks. These multiple-choice questions cover pretrained models, feature extraction, fine-tuning, domain adaptation, knowledge transfer, and practical applications. They are useful for students, AI developers, machine learning engineers, and candidates preparing for technical interviews. The collection includes foundational and practical questions to strengthen your understanding of Transfer Learning.
Transfer Learning MCQs
These Transfer Learning multiple-choice questions cover important concepts such as pretrained models, source and target domains, feature extraction, fine-tuning, domain adaptation, knowledge transfer, neural network layers, model initialization, transferability, and evaluation techniques. The questions explore how previously learned representations can improve performance on new tasks with limited data through conceptual, technical, and scenario-based questions.
These Transfer Learning MCQs help learners understand how models reuse learned features, adapt to new datasets, and reduce the need for training from scratch. Each question includes an answer and explanation.
List of Transfer Learning MCQs
Explore the following 50 MCQs covering Transfer Learning fundamentals, methods, architectures, benefits, limitations, evaluation techniques, and real-world implementation scenarios.
1. What is Transfer Learning in machine learning?
- A technique that trains every model without using previous knowledge
- A method that reuses knowledge learned from one task or domain to improve another task
- A process that converts all labeled data into random values
- A technique used only for database migration
Answer: B) A method that reuses knowledge learned from one task or domain to improve another task
Explanation:
Transfer Learning uses knowledge acquired from a source task or dataset to help solve a related target task. It can reduce training time and the amount of target-task data required, depending on how closely the tasks are related.
2. What is a pretrained model?
- A model whose parameters have already been learned through training on a dataset
- A model that has never processed any training data
- A model that can only store database records
- A model that cannot be modified after creation
Answer: A) A model whose parameters have already been learned through training on a dataset
Explanation:
A pretrained model has learned parameters from an earlier training process, often using a large dataset. Developers can reuse its learned representations for a new task through feature extraction, fine-tuning, or other adaptation methods.
3. What are the source task and target task in Transfer Learning?
- The source task is always smaller than the target task
- The source task is the final deployment stage, while the target task is data storage
- The source task is where transferable knowledge is learned, and the target task is where that knowledge is applied
- Both terms refer exclusively to the same dataset
Answer: C) The source task is where transferable knowledge is learned, and the target task is where that knowledge is applied
Explanation:
The source task provides the initial knowledge, while the target task is the new problem that benefits from that knowledge. The source and target tasks may differ, but useful shared patterns can make transfer possible.
4. Which approach reuses a pretrained model as a fixed feature extractor?
- Random initialization of every layer
- Training the entire network from scratch
- Deleting all pretrained parameters
- Keeping the pretrained feature-extraction layers frozen and training a new task-specific classifier
Answer: D) Keeping the pretrained feature-extraction layers frozen and training a new task-specific classifier
Explanation:
In feature extraction, the pretrained backbone generates representations for the target dataset while its parameters remain unchanged. A new task-specific head is trained to map those representations to the required output labels.
5. What is fine-tuning in Transfer Learning?
- Removing all layers from a pretrained model
- Continuing training of a pretrained model on target-task data
- Converting a neural network into a spreadsheet
- Using only random predictions during inference
Answer: B) Continuing training of a pretrained model on target-task data
Explanation:
Fine-tuning adapts some or all pretrained model parameters to the target task. It can improve task-specific performance, but careful learning-rate selection and validation are important to reduce overfitting and damage to useful pretrained representations.
6. Why are pretrained models often useful when target-task training data is limited?
- They provide previously learned representations that can reduce dependence on large target datasets
- They guarantee perfect accuracy without evaluation
- They automatically create correct labels for every target example
- They eliminate the need to define the target task
Answer: A) They provide previously learned representations that can reduce dependence on large target datasets
Explanation:
Pretrained models may already capture useful patterns from large or diverse datasets. Reusing these patterns can help a target model learn from fewer examples than training an equivalent model from random initialization, although the benefit depends on the task and data.
7. Which of the following is a common Transfer Learning application?
- Sorting files alphabetically without learning
- Compressing a document using a fixed encoding
- Adapting a pretrained image recognition model to identify plant diseases
- Converting a text file into a PDF without analysis
Answer: C) Adapting a pretrained image recognition model to identify plant diseases
Explanation:
A model pretrained on a broad image dataset can provide reusable visual features for a plant disease classification task. The model can be adapted using labeled plant images and evaluated on independent examples from the target domain.
8. What does freezing a layer mean during Transfer Learning?
- Deleting the layer from the architecture
- Preventing the layer from producing outputs
- Converting its weights into text labels
- Keeping its trainable parameters unchanged during optimization
Answer: D) Keeping its trainable parameters unchanged during optimization
Explanation:
A frozen layer continues to participate in the forward computation but its parameters are not updated by gradient-based optimization. Freezing layers can preserve pretrained features and reduce the number of trainable parameters.
9. What is a major advantage of feature extraction over full fine-tuning?
- It usually requires fewer trainable parameters and less computation during training
- It always produces higher accuracy on every target dataset
- It eliminates the need for target-task labels in supervised classification
- It changes every pretrained parameter automatically
Answer: A) It usually requires fewer trainable parameters and less computation during training
Explanation:
Feature extraction keeps the pretrained backbone fixed and trains only a task-specific head. This can make training less expensive and reduce overfitting on small datasets, although it may be less adaptable when the target domain differs substantially from the source domain.
10. Which optimizer is commonly used when fine-tuning a neural network?
- File sorting
- Adam
- HTML parsing
- Database indexing
Answer: B) Adam
Explanation:
Adam is a gradient-based optimization algorithm commonly used to fine-tune neural networks. Other optimizers, such as SGD with momentum, can also work well depending on the architecture, dataset, and training configuration.
11. What is domain adaptation?
- A method that prevents a model from processing new data
- A technique for increasing the number of network layers without training
- A Transfer Learning approach that addresses differences between source and target data distributions
- A method for converting images into database tables
Answer: C) A Transfer Learning approach that addresses differences between source and target data distributions
Explanation:
Domain adaptation aims to make a model perform well when the target data distribution differs from the source distribution. Depending on the setting, it may use labeled target examples, unlabeled target data, or both.
12. What is negative transfer in Transfer Learning?
- Transferring a model to a computer with a different operating system
- Using fewer layers than the original architecture
- Reducing the size of a dataset without changing its distribution
- When transferred knowledge reduces performance on the target task
Answer: D) When transferred knowledge reduces performance on the target task
Explanation:
Negative transfer occurs when knowledge from the source task is poorly suited to the target task and harms performance. It can arise from substantial domain differences, incompatible representations, or inappropriate fine-tuning choices.
13. Which neural network architecture is commonly used for Transfer Learning in computer vision?
- ResNet
- Binary search tree
- Relational database schema
- Queue data structure
Answer: A) ResNet
Explanation:
ResNet is a convolutional neural network architecture that uses residual connections to support the training of deep networks. Pretrained ResNet models are commonly adapted for image classification, object recognition, and other visual tasks.
14. Which pretrained model family is commonly used for Transfer Learning in Natural Language Processing?
- Merge sort
- BERT
- Binary search
- Quick sort
Answer: B) BERT
Explanation:
BERT is a transformer-based language model pretrained on text. Its representations can be adapted to tasks such as sentiment classification, question answering, and named entity recognition through fine-tuning or feature extraction.
15. What is the purpose of replacing the classification head of a pretrained model?
- To remove the model's ability to extract features
- To guarantee identical outputs for every input
- To match the output layer to the target task's labels or prediction requirements
- To eliminate the need for model evaluation
Answer: C) To match the output layer to the target task's labels or prediction requirements
Explanation:
A pretrained classification head may predict classes from the original dataset. Replacing it with a task-specific head allows the model to produce outputs for the target task, such as a new set of categories or a regression value.
16. What is a feature representation in a pretrained neural network?
- A list of training filenames
- A description of the computer's hardware
- A set of manually written instructions unrelated to the input
- A numerical representation of input characteristics learned by the model
Answer: D) A numerical representation of input characteristics learned by the model
Explanation:
Feature representations encode useful patterns in input data. For example, intermediate layers of an image model may represent edges, textures, shapes, or more complex visual structures that can support a target classification task.
17. Why is a lower learning rate often used when fine-tuning a pretrained model?
- To make smaller parameter updates that help preserve useful pretrained knowledge
- To prevent the model from calculating gradients
- To force the model to forget all previous training immediately
- To ensure that every batch contains the same example
Answer: A) To make smaller parameter updates that help preserve useful pretrained knowledge
Explanation:
A smaller learning rate can reduce the risk of making large updates that damage useful pretrained parameters. Fine-tuning strategies may use different learning rates for the backbone and task-specific head, but the best values should be selected through validation.
18. What is catastrophic forgetting during fine-tuning?
- The permanent loss of a computer's storage device
- The loss of previously acquired capabilities as a model adapts to a new task
- The automatic creation of additional training examples
- The conversion of model weights into class labels
Answer: B) The loss of previously acquired capabilities as a model adapts to a new task
Explanation:
Catastrophic forgetting occurs when learning a new task significantly degrades performance on previously learned tasks. Depending on the application, methods such as replay, regularization, or parameter-efficient adaptation can help preserve earlier capabilities.
19. What is the role of a validation dataset in Transfer Learning?
- It replaces all training examples
- It is used exclusively to store model parameters
- It helps select hyperparameters and monitor generalization during development
- It guarantees that the test results will be perfect
Answer: C) It helps select hyperparameters and monitor generalization during development
Explanation:
A validation dataset helps developers compare training configurations, select learning rates, and determine when to stop training. A separate test set should be reserved for final evaluation to avoid biased performance estimates.
20. Which statement about Transfer Learning and training from scratch is correct?
- Transfer Learning always requires more data than training from scratch
- Training from scratch always produces better results
- Transfer Learning can only be applied to language models
- Transfer Learning can reuse learned parameters, while training from scratch typically initializes model parameters without task-specific pretrained knowledge
Answer: D) Transfer Learning can reuse learned parameters, while training from scratch typically initializes model parameters without task-specific pretrained knowledge
Explanation:
Transfer Learning starts with knowledge learned previously, whereas training from scratch generally learns parameters from an initial initialization using the target training process. Transfer Learning can be more practical with limited data, but the best approach depends on the task, dataset, and available resources.
21. What is the main purpose of domain generalization?
- To improve performance on new domains without relying on target-domain adaptation data during training
- To memorize a single training domain
- To force all input examples to have identical features
- To eliminate the need for testing on unseen distributions
Answer: A) To improve performance on new domains without relying on target-domain adaptation data during training
Explanation:
Domain generalization aims to produce models that remain useful on unseen target domains. It differs from domain adaptation, which can use information from the target domain to adjust the model during development or training.
22. What is a major benefit of using Transfer Learning in computer vision?
- It eliminates all image preprocessing requirements in every task
- It allows reusable visual features from pretrained models to support new image tasks
- It guarantees that image classification will never produce errors
- It prevents models from detecting objects not included in the original dataset
Answer: B) It allows reusable visual features from pretrained models to support new image tasks
Explanation:
Pretrained vision models can learn general visual features from large datasets. These features can be reused for tasks such as medical image classification, defect detection, and wildlife recognition, provided the representations are appropriate for the target domain.
23. Which situation is most suitable for using Transfer Learning?
- A task that must not use any previous model or data
- A problem that cannot be expressed as a prediction or learning task
- A new classification task with limited labeled data and access to a relevant pretrained model
- A task where the output is predetermined and no computation is needed
Answer: C) A new classification task with limited labeled data and access to a relevant pretrained model
Explanation:
Transfer Learning is often valuable when labeled target data is limited and a pretrained model offers useful representations. A suitable source model can reduce the amount of target-task training required, although its relevance should be verified.
24. What does layer-wise fine-tuning mean?
- Removing every layer except the output layer
- Changing the input data into a different file format
- Training each layer with identical labels regardless of the task
- Unfreezing and adapting selected layers of a pretrained model, often in stages
Answer: D) Unfreezing and adapting selected layers of a pretrained model, often in stages
Explanation:
Layer-wise fine-tuning selectively updates parts of a pretrained network. For example, a developer may first train a new output head, then unfreeze some deeper backbone layers to adapt the model more closely to the target dataset.
25. Which of the following can indicate that negative transfer has occurred?
- A transferred model performs worse than a suitable baseline trained without that transferred knowledge
- A model uses a pretrained backbone
- A dataset contains both training and validation examples
- A neural network has more than one layer
Answer: A) A transferred model performs worse than a suitable baseline trained without that transferred knowledge
Explanation:
Negative transfer is indicated when reusing source-task knowledge harms target-task performance compared with an appropriate alternative. Comparing models under fair evaluation conditions can help determine whether transfer is beneficial.
26. What is parameter-efficient fine-tuning (PEFT)?
- A method that requires updating every parameter in a large model
- A collection of techniques that adapts a model by training a relatively small subset of parameters or additional parameters
- A process that converts neural networks into decision trees automatically
- A technique that prevents pretrained models from accepting new inputs
Answer: B) A collection of techniques that adapts a model by training a relatively small subset of parameters or additional parameters
Explanation:
PEFT methods reduce the number of parameters that must be trained for a target task. They can lower memory and storage requirements compared with updating an entire large model. LoRA and adapter-based methods are common examples.
27. What does LoRA stand for in the context of Transfer Learning?
- Layer Output Reduction Algorithm
- Learning Optimization and Regression Architecture
- Low-Rank Adaptation
- Logical Representation Assessment
Answer: C) Low-Rank Adaptation
Explanation:
Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning technique that introduces trainable low-rank updates to selected model weights. It allows adaptation while keeping many of the original pretrained parameters frozen.
28. What is the purpose of adapters in Transfer Learning?
- To replace the training dataset with hardware specifications
- To remove all pretrained knowledge from a model
- To guarantee perfect accuracy on unseen data
- To add small trainable modules that adapt a pretrained model to a target task
Answer: D) To add small trainable modules that adapt a pretrained model to a target task
Explanation:
Adapter methods insert small trainable components into a pretrained architecture. Instead of updating every original parameter, training focuses on these additional modules, which can reduce the cost of adapting large models to multiple tasks.
29. Why is data distribution important when selecting a pretrained model?
- A large mismatch between source and target distributions can reduce the usefulness of transferred representations
- Distribution has no effect on any machine learning model
- Every pretrained model performs equally well on all datasets
- A model's source dataset determines the target labels automatically
Answer: A) A large mismatch between source and target distributions can reduce the usefulness of transferred representations
Explanation:
A pretrained model learns patterns from its source data. When the target data differs significantly in appearance, language, environment, or task requirements, the source representations may be less suitable and additional adaptation may be needed.
30. What is the purpose of data augmentation when fine-tuning a model?
- To duplicate only incorrect predictions without changing the data
- To create useful variations of training examples that may improve generalization
- To remove all variation from the training dataset
- To replace model evaluation with random sampling
Answer: B) To create useful variations of training examples that may improve generalization
Explanation:
Data augmentation applies transformations such as image cropping, flipping, or color variation when appropriate for the task. It can help a fine-tuned model generalize better, but transformations must preserve the meaning and labels of the examples.
31. Which practice helps reduce overfitting during Transfer Learning?
- Evaluating performance only on training examples
- Removing all validation data
- Using validation-based early stopping and suitable regularization
- Increasing the number of training parameters without monitoring results
Answer: C) Using validation-based early stopping and suitable regularization
Explanation:
Early stopping can halt training when validation performance stops improving. Regularization, appropriate augmentation, and limiting the number of trainable parameters can also help control overfitting, especially when the target dataset is small.
32. How is Transfer Learning commonly used in sentiment analysis?
- By converting sentiment labels into computer hardware instructions
- By requiring a separate language model for every sentence
- By removing all text before classification
- By adapting a pretrained language model to classify text as positive, negative, or another sentiment category
Answer: D) By adapting a pretrained language model to classify text as positive, negative, or another sentiment category
Explanation:
A language model pretrained on a large text corpus can provide useful linguistic representations. Fine-tuning it on labeled sentiment examples allows the model to learn the target classification task.
33. What is a domain shift in machine learning?
- A difference between the data distribution used for training and the distribution encountered during deployment or evaluation
- A change in the computer's desktop wallpaper
- A method for increasing the number of class labels automatically
- A process that makes every dataset identical
Answer: A) A difference between the data distribution used for training and the distribution encountered during deployment or evaluation
Explanation:
Domain shift occurs when the characteristics or distribution of incoming data differ from those used during training. It can reduce model accuracy and is a key consideration in Transfer Learning and domain adaptation.
34. Which statement about transferability of learned features is correct?
- Features from early layers are always useless for new tasks
- The usefulness of transferred features depends on the source model, target task, and level of similarity between domains
- Every feature learned from one dataset transfers perfectly to every other dataset
- Only the output layer can contain reusable information
Answer: B) The usefulness of transferred features depends on the source model, target task, and level of similarity between domains
Explanation:
Learned features can be reused when they capture patterns relevant to the target task. Their usefulness depends on the model architecture, source training data, target distribution, and adaptation method. Some features may require substantial modification.
35. What is the main purpose of a learning-rate scheduler during fine-tuning?
- To change the target labels at every training step
- To determine the number of classes without examining the task
- To adjust the learning rate during training according to a chosen schedule or rule
- To eliminate the need for an optimizer
Answer: C) To adjust the learning rate during training according to a chosen schedule or rule
Explanation:
A learning-rate scheduler changes the step size used by an optimizer over time. A suitable schedule can support stable fine-tuning, but the best configuration depends on the model, dataset, and optimization objective.
36. What is zero-shot transfer in machine learning?
- Training a model on the target task with millions of labeled examples
- Transferring a dataset between storage devices without changing it
- Fine-tuning a model using a large target-task labeled dataset
- Applying learned knowledge to a target task without target-task training examples, depending on the model and task setup
Answer: D) Applying learned knowledge to a target task without target-task training examples, depending on the model and task setup
Explanation:
Zero-shot transfer uses knowledge learned previously to attempt a new task without labeled examples specifically supplied for that target task. It can rely on pretrained representations, semantic descriptions, or instructions, but its success depends on how well the learned knowledge matches the target problem.
37. Why might a developer freeze early layers but fine-tune later layers of a pretrained vision model?
- To preserve broadly useful low-level features while adapting higher-level representations to the target task
- To prevent the model from processing images
- To guarantee that the model will use fewer input pixels
- To eliminate the need for labeled target data in all cases
Answer: A) To preserve broadly useful low-level features while adapting higher-level representations to the target task
Explanation:
Early layers often learn general visual patterns, while deeper layers may represent more task-specific features. Freezing early layers and fine-tuning later ones can be a useful compromise, although the best layer configuration depends on the dataset and model.
38. What is a potential disadvantage of fine-tuning all layers of a large pretrained model?
- It makes the model incapable of receiving input data
- It can require substantial computational resources and may overfit limited target data
- It guarantees that every pretrained feature remains unchanged
- It removes the need for gradient calculations
Answer: B) It can require substantial computational resources and may overfit limited target data
Explanation:
Full fine-tuning updates many parameters and may require significant memory and computation. When target data is limited, it can also overfit or alter useful pretrained representations, making feature extraction or parameter-efficient fine-tuning attractive alternatives.
39. Which metric is appropriate for evaluating a Transfer Learning model on a classification task?
- Number of source-code comments
- Size of the model's filename
- Accuracy, precision, recall, or F1-score, depending on the evaluation requirements
- Number of folders in the project
Answer: C) Accuracy, precision, recall, or F1-score, depending on the evaluation requirements
Explanation:
Classification metrics measure different aspects of predictive performance. Accuracy can be useful for balanced datasets, while precision, recall, and F1-score can provide more informative results for imbalanced data or applications with different error costs.
40. What is an important consideration when using a pretrained model from an external source?
- Whether its file name is short enough
- Whether the model has the largest possible number of parameters
- Whether its predictions can be accepted without validation
- Its license, training-data characteristics, intended use, and suitability for the target task
Answer: D) Its license, training-data characteristics, intended use, and suitability for the target task
Explanation:
Developers should review licensing conditions, model documentation, known limitations, training-data details, and intended use before deployment. They should also evaluate the model for target-domain performance, fairness, privacy, and security requirements.
41. How can Transfer Learning support speech recognition systems?
- By adapting pretrained speech representations to a new language, accent, or acoustic environment
- By converting speech into random labels without analysis
- By requiring every speaker to record all possible sentences
- By eliminating the need to process audio signals
Answer: A) By adapting pretrained speech representations to a new language, accent, or acoustic environment
Explanation:
Pretrained speech models can learn useful acoustic and linguistic representations from large datasets. Adapting them to a specific language, accent, or environment can improve recognition when suitable target-domain data is available.
42. Which statement best describes multi-task learning compared with Transfer Learning?
- Multi-task learning can only train one task at a time and cannot share information
- Multi-task learning commonly learns multiple tasks together, while Transfer Learning often reuses knowledge from a source task for a target task
- Transfer Learning never uses pretrained parameters
- Both terms always refer to exactly the same training procedure
Answer: B) Multi-task learning commonly learns multiple tasks together, while Transfer Learning often reuses knowledge from a source task for a target task
Explanation:
Multi-task learning generally optimizes a model across multiple tasks, often with shared representations. Transfer Learning focuses on using previously learned knowledge to improve a different target task, which may be trained later and separately.
43. What is the purpose of benchmarking a transferred model against a baseline?
- To ensure that all models have identical parameter values
- To remove all differences between training and testing data
- To determine whether transfer provides a meaningful improvement over an alternative approach
- To eliminate the need for evaluation metrics
Answer: C) To determine whether transfer provides a meaningful improvement over an alternative approach
Explanation:
A baseline provides a reference for measuring the value of Transfer Learning. Comparing a pretrained model with alternatives such as a simple classifier or a model trained from scratch can reveal whether transferred knowledge improves accuracy, training cost, or generalization.
44. What is a common use of Transfer Learning with large language models?
- Converting all text into images before processing
- Removing the pretrained model's language representations
- Training a model without defining any task or objective
- Adapting a pretrained language model to specialized tasks such as summarization, classification, or question answering
Answer: D) Adapting a pretrained language model to specialized tasks such as summarization, classification, or question answering
Explanation:
Large language models can be adapted to specialized applications through supervised fine-tuning, parameter-efficient methods, or other approaches. The chosen method depends on the task, available examples, compute resources, and deployment constraints.
45. Which factor is important when deciding whether to freeze or unfreeze pretrained layers?
- The size of the target dataset, domain similarity, computational budget, and validation performance
- The color of the computer monitor
- The number of comments in the training script
- The operating system's wallpaper settings
Answer: A) The size of the target dataset, domain similarity, computational budget, and validation performance
Explanation:
Freezing and unfreezing decisions depend on how much target data is available, how relevant the pretrained representations are, and the resources available for training. Developers should compare configurations using a suitable validation dataset.
46. What is a major advantage of Transfer Learning for small organizations developing AI applications?
- It eliminates every requirement for computing resources
- It can reduce the data, time, and computing resources needed compared with training a comparable model from scratch
- It guarantees commercial success for every AI product
- It removes the need to test the model before deployment
Answer: B) It can reduce the data, time, and computing resources needed compared with training a comparable model from scratch
Explanation:
Reusing pretrained models can make AI development more accessible to organizations with limited datasets or compute budgets. The actual savings depend on the model size, adaptation method, licensing terms, and requirements of the target application.
47. Why should a test dataset remain separate from the data used for fine-tuning?
- To ensure the model never produces an output
- To force the model to memorize every example
- To obtain a more reliable estimate of performance on unseen examples
- To eliminate the need for training labels
Answer: C) To obtain a more reliable estimate of performance on unseen examples
Explanation:
If test examples influence training or repeated model selection, the final evaluation can become overly optimistic. Keeping the test set separate helps measure how well the adapted model generalizes to data not used during development.
48. A developer has 300 labeled images of damaged electronic components and access to a relevant pretrained image classification model. Which strategy is a reasonable starting point?
- Discard the pretrained model and assume 300 images are sufficient for every architecture
- Evaluate the model only on the same 300 images used for training
- Train a model with random labels and skip validation
- Replace the classification head, train it on the target dataset, and compare feature extraction with selective fine-tuning using a held-out validation set
Answer: D) Replace the classification head, train it on the target dataset, and compare feature extraction with selective fine-tuning using a held-out validation set
Explanation:
A relevant pretrained model can provide useful visual features for the small target dataset. Training a new classification head is a sensible baseline, while selective fine-tuning may improve adaptation. A separate validation set helps choose the approach and identify overfitting.
49. A company fine-tunes a language model on customer support messages, but its performance drops when deployed on messages written in a different regional dialect. Which issue should the team investigate first?
- Domain shift between the training messages and deployment messages
- The alphabetical order of the training filenames
- The number of comments in the source code
- The color scheme used in the user interface
Answer: A) Domain shift between the training messages and deployment messages
Explanation:
The regional dialect may introduce vocabulary, spelling, and linguistic patterns that differ from the training data. The team should evaluate performance on representative deployment data and consider collecting suitable examples, adapting the model, or applying domain adaptation techniques.
50. An organization needs to deploy a pretrained transformer for legal document classification. It has a small labeled legal dataset, limited GPU memory, and strict requirements for validating performance before deployment. Which approach is most appropriate?
- Train a very large transformer from scratch and evaluate it only on the training data
- Start with a suitable pretrained model, consider parameter-efficient fine-tuning or a frozen feature extractor, and select the approach using validation data before final testing
- Use the pretrained model without checking whether its representations suit legal documents
- Fine-tune every parameter with the largest possible learning rate and deploy without evaluation
Answer: B) Start with a suitable pretrained model, consider parameter-efficient fine-tuning or a frozen feature extractor, and select the approach using validation data before final testing
Explanation:
A pretrained transformer can provide useful language representations, while parameter-efficient fine-tuning can reduce memory requirements. A frozen feature extractor is another potential baseline. The organization should compare approaches on representative validation data, reserve a separate test set for final evaluation, and verify privacy, security, and domain-specific performance before deployment.