×

Trending Technologies MCQs

Self-Supervised Learning MCQs (Multiple-Choice Questions)

Practice Self-Supervised Learning MCQs to test your knowledge of machine learning techniques that generate supervisory signals directly from unlabeled data. These questions cover pretext tasks, contrastive learning, masked prediction, representation learning, and model pretraining. They are useful for students, machine learning engineers, AI developers, and professionals preparing for technical interviews or assessments. The set includes both foundational and practical questions covering modern Self-Supervised Learning methods.

Self-Supervised Learning MCQs

These Self-Supervised Learning multiple-choice questions cover important concepts such as positive and negative samples, contrastive objectives, masked language modeling, autoencoders, embeddings, augmentation strategies, and transfer learning. They also explore loss functions, representation quality, training pipelines, evaluation methods, and applications in computer vision, natural language processing, and speech recognition. This set combines conceptual, technical, and scenario-based questions to test your understanding of Self-Supervised Learning systems.

Self-Supervised Learning MCQs cover the technologies used to learn useful representations from unlabeled data by creating training signals from the data itself. Each question includes an answer and explanation.

List of Self-Supervised Learning MCQs

The following Self-Supervised Learning multiple-choice questions cover learning fundamentals, pretext tasks, contrastive learning, masked prediction, representation evaluation, and real-world applications.

1. What is Self-Supervised Learning?

  1. A learning method that requires manually labeled data for every example
  2. A technique used only for reinforcement learning
  3. A machine learning approach that generates supervisory signals from the data itself
  4. A method that prevents models from learning representations

Answer: C) A machine learning approach that generates supervisory signals from the data itself

Explanation:

Self-Supervised Learning creates training targets from the structure or content of unlabeled data. The model learns by solving tasks such as predicting masked tokens, reconstructing missing information, or identifying relationships between different views of the same example.

2. How does Self-Supervised Learning differ from traditional supervised learning?

  1. It derives training signals from the data rather than requiring manually assigned labels for every training example
  2. It cannot use neural networks
  3. It always requires fewer computational resources
  4. It does not use a training objective

Answer: A) It derives training signals from the data rather than requiring manually assigned labels for every training example

Explanation:

Traditional supervised learning generally relies on labeled examples containing target outputs supplied by a dataset creator or another process. Self-Supervised Learning constructs targets from the data, reducing dependence on manually labeled datasets for pretraining.

3. What is a pretext task in Self-Supervised Learning?

  1. A task that requires a human to label every input
  2. A task used only to measure hardware performance
  3. A task that prevents the model from learning patterns
  4. A training task designed to generate supervision from unlabeled data

Answer: D) A training task designed to generate supervision from unlabeled data

Explanation:

A pretext task provides a learning objective derived from the input data. Examples include predicting masked words, reconstructing image patches, and determining whether two augmented views originated from the same example.

4. What is the main purpose of representation learning?

  1. To increase the number of labels in a dataset
  2. To learn useful numerical representations of input data
  3. To remove all meaningful information from inputs
  4. To replace training with random parameter selection

Answer: B) To learn useful numerical representations of input data

Explanation:

Representation learning transforms raw data into features or embeddings that capture useful structure. These representations can support downstream tasks such as classification, retrieval, clustering, and prediction.

5. Which of the following is an example of Self-Supervised Learning in natural language processing?

  1. Predicting masked words in a sentence using the surrounding context
  2. Assigning every sentence a manually written category before training
  3. Sorting documents alphabetically without training a model
  4. Deleting words from a corpus without defining a learning objective

Answer: A) Predicting masked words in a sentence using the surrounding context

Explanation:

Masked language modeling hides selected tokens and trains a model to predict the original tokens from the remaining context. The original text supplies the target information, so manual labels for each masked token are unnecessary.

6. What is contrastive learning?

  1. A method that trains models only on manually labeled categories
  2. A technique that removes similarity information from embeddings
  3. A learning approach that encourages related representations to be similar and, in many methods, unrelated ones to be distinguishable
  4. A method that prevents models from comparing examples

Answer: C) A learning approach that encourages related representations to be similar and, in many methods, unrelated ones to be distinguishable

Explanation:

Contrastive learning trains representations using relationships between examples or views. Many approaches pull positive pairs closer in representation space and distinguish them from negative pairs, helping the model capture useful similarities.

7. What is a positive pair in contrastive Self-Supervised Learning?

  1. Two examples that must belong to different semantic categories
  2. Two views or examples treated as related under the learning objective
  3. Two randomly generated labels with no relationship to the data
  4. Two model checkpoints with identical file names

Answer: B) Two views or examples treated as related under the learning objective

Explanation:

A positive pair often consists of two augmented views of the same original image or two related representations of a sample. The training objective encourages the model to capture their shared information, although the exact definition of a positive pair depends on the method.

8. What is a negative pair in many contrastive learning methods?

  1. Two views that are always identical at the pixel level
  2. Two examples that have both been manually assigned the same label
  3. Two inputs that cannot be represented by neural networks
  4. Two examples treated as unrelated by the contrastive objective

Answer: D) Two examples treated as unrelated by the contrastive objective

Explanation:

Negative pairs are treated as contrasting examples in many contrastive objectives. The model learns to distinguish them from positive pairs, but incorrectly treating semantically similar examples as negatives can harm representation quality.

9. What is the purpose of data augmentation in Self-Supervised Learning?

  1. To create transformed views that help the model learn useful and robust representations
  2. To ensure that every training image receives a manually assigned class
  3. To remove all variation from the training data
  4. To replace the learning objective with a file conversion process

Answer: A) To create transformed views that help the model learn useful and robust representations

Explanation:

Data augmentation creates modified versions of training examples through transformations such as cropping, flipping, or color changes. In Self-Supervised Learning, these views can provide related training examples without requiring manual labels.

10. What is masked language modeling?

  1. A technique for encrypting all words in a document
  2. A method that trains models only on speech signals
  3. A training objective in which selected tokens are hidden or corrupted and the model learns to predict the original tokens
  4. A method that deletes text permanently before training

Answer: C) A training objective in which selected tokens are hidden or corrupted and the model learns to predict the original tokens

Explanation:

Masked language modeling trains a model to recover selected tokens from surrounding context. It is used in models such as BERT and helps them learn contextual representations of language.

11. Which model is well known for using masked language modeling during pretraining?

  1. Linear regression
  2. BERT
  3. K-means clustering
  4. Decision stump

Answer: B) BERT

Explanation:

BERT uses masked language modeling as a major pretraining objective. The model learns to predict masked tokens using their surrounding context, producing contextual representations useful for many natural language processing tasks.

12. What is the main idea behind autoregressive Self-Supervised Learning?

  1. Predicting every input token from future tokens only
  2. Removing sequence order before training
  3. Training without any prediction target
  4. Predicting the next token or sequence element from preceding context

Answer: D) Predicting the next token or sequence element from preceding context

Explanation:

Autoregressive training uses preceding tokens to predict the next token in a sequence. The original sequence provides the targets automatically, allowing language models to learn from large collections of unlabeled text.

13. What is an autoencoder?

  1. A neural network trained to encode input data into a representation and reconstruct the input
  2. A system that always classifies images using manually supplied labels
  3. A database that stores only model predictions
  4. A model that cannot learn compressed representations

Answer: A) A neural network trained to encode input data into a representation and reconstruct the input

Explanation:

An autoencoder typically contains an encoder that maps an input to a latent representation and a decoder that reconstructs the input from that representation. Reconstruction objectives can provide self-generated supervision for learning useful features.

14. What is the role of the encoder in an autoencoder?

  1. To assign human labels to all training samples
  2. To reconstruct the original input directly without a representation
  3. To transform the input into a latent representation
  4. To calculate only the final evaluation accuracy

Answer: C) To transform the input into a latent representation

Explanation:

The encoder maps the input into a latent feature space. The decoder uses this representation to reconstruct the input, while the reconstruction loss encourages the latent representation to preserve information relevant to the chosen objective.

15. What is reconstruction loss in Self-Supervised Learning?

  1. A measure of the number of human labels in a dataset
  2. A measure of the difference between an original target and its reconstructed or predicted version
  3. A measure of network bandwidth only
  4. A method for preventing model optimization

Answer: B) A measure of the difference between an original target and its reconstructed or predicted version

Explanation:

Reconstruction loss quantifies how closely a model's reconstruction matches the target input. Depending on the data and objective, it may use mean squared error, cross-entropy, or another suitable measure.

16. What is a latent representation?

  1. The complete raw dataset stored without modification
  2. A manually written description of every input
  3. A file containing only the model's evaluation scores
  4. A learned internal representation that captures features of the input

Answer: D) A learned internal representation that captures features of the input

Explanation:

A latent representation is a numerical encoding learned by a model. It can capture useful structure in the input and serve as a foundation for downstream tasks such as classification, clustering, retrieval, or generation.

17. What is the InfoNCE loss commonly used for?

  1. Training a model to distinguish a positive example from contrasting examples using a normalized similarity-based objective
  2. Removing all relationships between embeddings
  3. Measuring only the number of parameters in a model
  4. Replacing all training data with manually written rules

Answer: A) Training a model to distinguish a positive example from contrasting examples using a normalized similarity-based objective

Explanation:

InfoNCE is a contrastive objective that scores a positive pair relative to competing examples. It encourages representations of related views to be distinguishable from other candidates, depending on the sampling strategy and similarity function.

18. What is representation collapse in Self-Supervised Learning?

  1. A condition in which the model learns an ideal representation for every task
  2. A process that increases the diversity of all learned features
  3. A failure mode in which different inputs map to identical or nearly identical representations
  4. A method for improving the resolution of training images

Answer: C) A failure mode in which different inputs map to identical or nearly identical representations

Explanation:

Representation collapse occurs when a model produces nearly the same representation for many different inputs, losing useful distinctions. Some Self-Supervised Learning methods use architectural choices, normalization, regularization, or specialized objectives to reduce the risk of collapse.

19. What is a projection head in many contrastive learning architectures?

  1. A component used only to display training graphs
  2. A small neural network that maps learned features into the space used by the training objective
  3. A tool for manually labeling every image
  4. A storage system for raw audio recordings

Answer: B) A small neural network that maps learned features into the space used by the training objective

Explanation:

A projection head transforms encoder representations into a space where the self-supervised loss is applied. In some methods, the encoder representation before the projection head is used for downstream tasks rather than the projected vector.

20. What is the main goal of SimCLR?

  1. To train models only on manually labeled image categories
  2. To replace neural networks with decision trees
  3. To predict future stock prices using reinforcement learning
  4. To learn visual representations through contrastive learning on augmented image views

Answer: D) To learn visual representations through contrastive learning on augmented image views

Explanation:

SimCLR is a Self-Supervised Learning framework for visual representation learning. It creates two augmented views of each image and uses a contrastive objective to encourage their representations to be similar relative to representations from other images.

21. What is the primary idea behind BYOL?

  1. Learning representations by predicting the target network's representation of another augmented view without explicit negative pairs
  2. Training only with manually assigned image labels
  3. Using negative examples as the only training signal
  4. Eliminating the need for a neural encoder

Answer: A) Learning representations by predicting the target network's representation of another augmented view without explicit negative pairs

Explanation:

BYOL uses online and target networks to learn representations from different augmented views of the same input. Its training objective does not rely on explicit negative pairs, and its design uses mechanisms such as a slowly updated target network to support learning without collapse.

22. What is the role of a target network in methods such as BYOL?

  1. To assign permanent labels to the original dataset
  2. To generate random labels at every training step
  3. To provide target representations that are updated using a controlled mechanism
  4. To remove the online network from training

Answer: C) To provide target representations that are updated using a controlled mechanism

Explanation:

In BYOL, the target network provides representations for the online network to predict. Its parameters are commonly updated through an exponential moving average of the online network parameters rather than by direct gradient updates from the same objective.

23. What is the purpose of a momentum encoder in MoCo?

  1. To remove the need for image augmentation
  2. To produce relatively consistent key representations using a slowly updated encoder
  3. To replace the contrastive learning objective with supervised labels
  4. To guarantee that all embeddings have identical values

Answer: B) To produce relatively consistent key representations using a slowly updated encoder

Explanation:

Momentum Contrast (MoCo) uses a momentum-updated encoder to generate key representations and a queue of keys for contrastive learning. The slower parameter updates help maintain consistency among representations used as contrastive candidates.

24. What is the purpose of masking in image-based Self-Supervised Learning?

  1. To prevent the model from processing visual information
  2. To replace every image with a class label
  3. To guarantee perfect image reconstruction without training
  4. To hide selected image regions and train the model to predict or reconstruct their content

Answer: D) To hide selected image regions and train the model to predict or reconstruct their content

Explanation:

Masked image modeling hides some image patches and trains the model to infer missing information from the visible regions. This encourages learning useful visual representations from unlabeled images.

25. Which model is strongly associated with masked image modeling using a vision transformer?

  1. Masked Autoencoders (MAE)
  2. K-means clustering
  3. Naive Bayes
  4. Linear discriminant analysis

Answer: A) Masked Autoencoders (MAE)

Explanation:

Masked Autoencoders use a vision transformer encoder to process visible image patches and a decoder to reconstruct the masked content. The training process allows visual representations to be learned from unlabeled images.

26. What is the purpose of a teacher-student architecture in Self-Supervised Learning?

  1. To require manual labels for every example
  2. To make both networks produce unrelated outputs
  3. To train a student network using targets or representations produced by a teacher network
  4. To remove representation learning from the training process

Answer: C) To train a student network using targets or representations produced by a teacher network

Explanation:

In teacher-student approaches, the teacher generates target predictions or representations that guide the student. Depending on the method, the teacher may be a separately trained model or a slowly updated version of the student.

27. What is self-distillation?

  1. A technique that requires two unrelated datasets and no model
  2. A method in which a model learns from targets generated by another version or branch of itself
  3. A process that permanently removes learned model parameters
  4. A method used exclusively for database indexing

Answer: B) A method in which a model learns from targets generated by another version or branch of itself

Explanation:

Self-distillation uses predictions or representations from a teacher-like version of a model to guide another version or branch. Some Self-Supervised Learning methods use this strategy to learn consistent representations without relying on manually labeled targets.

28. What is the main purpose of learning an embedding from unlabeled data?

  1. To store the complete training dataset in every vector
  2. To make all examples indistinguishable
  3. To eliminate numerical features from the model
  4. To encode useful properties of an input in a numerical vector

Answer: D) To encode useful properties of an input in a numerical vector

Explanation:

An embedding maps an input into a numerical vector that captures selected properties or relationships. Self-Supervised Learning aims to make these vectors useful for downstream tasks, such as finding similar documents or recognizing visual patterns.

29. How can Self-Supervised Learning help when labeled data is limited?

  1. By pretraining a model on a large collection of unlabeled data before fine-tuning it with a smaller labeled dataset
  2. By eliminating the need for all forms of evaluation
  3. By converting unlabeled data into guaranteed-correct human labels
  4. By preventing the model from learning general features

Answer: A) By pretraining a model on a large collection of unlabeled data before fine-tuning it with a smaller labeled dataset

Explanation:

Self-Supervised Learning can use large unlabeled datasets to learn general-purpose representations. These representations can then be adapted to a downstream task using a smaller labeled dataset, although the benefits depend on data quality and how well the pretraining task matches the target application.

30. What is transfer learning?

  1. Training every model from scratch without reusing learned features
  2. Moving training data between storage devices only
  3. Reusing knowledge or representations learned from one task or dataset for another task
  4. Replacing the training objective with a random function

Answer: C) Reusing knowledge or representations learned from one task or dataset for another task

Explanation:

Transfer learning adapts knowledge from a previously trained model to a new task. Self-Supervised Learning is often used for pretraining, after which a model can be fine-tuned or used as a feature extractor for downstream applications.

31. What is linear probing in representation learning?

  1. Retraining every layer of a pretrained model from scratch
  2. Training a linear classifier on frozen representations to evaluate their usefulness
  3. Removing the encoder before evaluation
  4. Measuring only the number of training examples

Answer: B) Training a linear classifier on frozen representations to evaluate their usefulness

Explanation:

Linear probing freezes the pretrained encoder and trains a simple linear classifier using its representations. It provides an indication of how much task-relevant information is accessible through a linear decision boundary, although it does not measure every aspect of representation quality.

32. What is fine-tuning after Self-Supervised Learning?

  1. Deleting the pretrained model and discarding all learned representations
  2. Evaluating a model without using any downstream data
  3. Replacing all learned weights with random values in every training step
  4. Adapting pretrained model parameters to a particular downstream task

Answer: D) Adapting pretrained model parameters to a particular downstream task

Explanation:

Fine-tuning continues training a pretrained model using data and objectives relevant to a target task. Depending on the approach, the entire model or only selected layers may be updated.

33. What is a potential disadvantage of contrastive learning with poorly chosen negative samples?

  1. Semantically similar examples may be pushed apart even though their representations should remain related
  2. The model automatically becomes a supervised classifier
  3. The encoder can no longer produce numerical vectors
  4. All training losses become zero immediately

Answer: A) Semantically similar examples may be pushed apart even though their representations should remain related

Explanation:

False negatives occur when examples treated as unrelated are actually semantically similar. Pushing their representations apart can damage the learned feature space, so negative sampling and objective design can significantly affect representation quality.

34. Why is the choice of data augmentation important in Self-Supervised Learning?

  1. Any augmentation always improves representation quality
  2. Augmentation is relevant only to supervised learning
  3. Transformations should preserve information relevant to the learning objective
  4. Augmentation must remove all semantic information from an input

Answer: C) Transformations should preserve information relevant to the learning objective

Explanation:

Augmentations define which variations the model is encouraged to treat as equivalent. If a transformation removes information essential to the task, the model may learn representations that ignore important distinctions.

35. What is cross-modal Self-Supervised Learning?

  1. Learning only from one numerical feature without context
  2. Learning relationships between different data modalities, such as images and text, using automatically derived supervision
  3. Training a model using only manually entered labels
  4. Removing alignment between audio and video data

Answer: B) Learning relationships between different data modalities, such as images and text, using automatically derived supervision

Explanation:

Cross-modal Self-Supervised Learning uses relationships between modalities to create training signals. For example, an image and its accompanying caption can form a positive pair for learning aligned visual and textual representations.

36. What is a common Self-Supervised Learning approach for speech data?

  1. Using only manually transcribed speech for every training objective
  2. Discarding the audio waveform before training
  3. Assigning random transcripts to all audio recordings
  4. Predicting masked or quantized audio representations from surrounding context

Answer: D) Predicting masked or quantized audio representations from surrounding context

Explanation:

Self-Supervised speech models can hide portions of an audio representation and predict targets derived from the signal. This allows models to learn useful speech features from large amounts of audio without requiring a manual transcript for every recording.

37. What is the purpose of a codebook in some Self-Supervised Learning methods?

  1. To map representations to a finite set of learned discrete codes or prototypes
  2. To store only the source code of the training software
  3. To prevent representations from being quantized
  4. To assign permanent human labels to all examples

Answer: A) To map representations to a finite set of learned discrete codes or prototypes

Explanation:

Some self-supervised methods use a codebook to map continuous representations to discrete codes or learned prototypes. These codes can provide prediction targets or compact representations for training, as seen in certain approaches to visual and speech representation learning.

38. What is the main purpose of clustering-based Self-Supervised Learning?

  1. To require a human to define every cluster assignment in advance
  2. To eliminate the need for feature representations
  3. To use cluster assignments or prototypes as training signals for learning representations
  4. To prevent similar examples from being grouped together

Answer: C) To use cluster assignments or prototypes as training signals for learning representations

Explanation:

Clustering-based approaches group representations or compare them with learned prototypes. The resulting assignments can provide pseudo-targets that guide further representation learning without requiring manually assigned class labels.

39. What is a pseudo-label in machine learning?

  1. A label that must always be supplied by a human expert
  2. A target label generated by a model or algorithm rather than directly provided as a ground-truth annotation
  3. A label that cannot be used during training
  4. A numerical value that contains no relationship to the input

Answer: B) A target label generated by a model or algorithm rather than directly provided as a ground-truth annotation

Explanation:

Pseudo-labels are automatically generated targets that can be used to train a model. They are common in semi-supervised learning and may also appear in self-supervised pipelines. Their quality depends on the method used to generate them, and errors can be reinforced during training.

40. What is the difference between Self-Supervised Learning and semi-supervised learning?

  1. Self-Supervised Learning cannot use unlabeled data
  2. Semi-supervised learning never uses labeled examples
  3. Both approaches require all examples to have manually assigned labels
  4. Self-Supervised Learning derives targets from data, while semi-supervised learning typically combines labeled and unlabeled examples

Answer: D) Self-Supervised Learning derives targets from data, while semi-supervised learning typically combines labeled and unlabeled examples

Explanation:

Self-Supervised Learning creates its own training signals from the data, often using unlabeled examples. Semi-supervised learning typically uses a smaller labeled dataset together with a larger unlabeled dataset, although practical pipelines may combine both approaches.

41. Why can Self-Supervised Learning reduce dependence on manual annotation?

  1. Because it creates training targets from patterns, context, or relationships within the data
  2. Because it makes data quality irrelevant
  3. Because it eliminates all model evaluation requirements
  4. Because it guarantees that every learned feature is useful

Answer: A) Because it creates training targets from patterns, context, or relationships within the data

Explanation:

Self-supervised objectives derive targets from the original data or transformed views. This makes it possible to pretrain models on large datasets without manually annotating every example, although downstream tasks may still require labeled data for evaluation or fine-tuning.

42. What is data leakage in the evaluation of a Self-Supervised Learning model?

  1. Using data augmentation during training
  2. Learning embeddings from unlabeled training examples
  3. Allowing information from evaluation data to improperly influence training or model selection
  4. Using a pretrained encoder for a downstream task

Answer: C) Allowing information from evaluation data to improperly influence training or model selection

Explanation:

Data leakage occurs when information from a validation or test set improperly affects model training or selection, making evaluation results overly optimistic. Dataset splits, preprocessing, duplicate detection, and benchmark contamination should be managed carefully.

43. What is a common challenge when pretraining a model using a very large unlabeled dataset?

  1. Unlabeled datasets cannot contain useful information
  2. Computational cost, data quality, duplication, and dataset bias can affect training
  3. The model cannot use GPU acceleration
  4. Large datasets always lead to worse representations

Answer: B) Computational cost, data quality, duplication, and dataset bias can affect training

Explanation:

Large-scale pretraining can require substantial computation and data management. Duplicates, low-quality examples, harmful content, or unrepresentative data can affect the learned representations, so dataset curation and evaluation remain important.

44. How can Self-Supervised Learning be applied to anomaly detection?

  1. By guaranteeing that every unusual example is malicious
  2. By ignoring patterns in normal data
  3. By requiring all anomalies to be labeled before any training
  4. By learning representations or normal patterns and identifying examples that differ substantially from them

Answer: D) By learning representations or normal patterns and identifying examples that differ substantially from them

Explanation:

A model can learn patterns from largely unlabeled data and use reconstruction error, representation distances, or other scores to identify unusual examples. An unusual example is not necessarily an anomaly of practical concern, so thresholds and performance should be validated on representative data.

45. What is the purpose of evaluating a pretrained representation on downstream tasks?

  1. To determine whether the learned features transfer to practical tasks
  2. To ensure that the encoder never needs training again
  3. To measure only the amount of unlabeled data collected
  4. To replace all training objectives with random predictions

Answer: A) To determine whether the learned features transfer to practical tasks

Explanation:

Downstream evaluation tests whether a pretrained representation supports tasks such as classification, retrieval, segmentation, or question answering. Strong performance on a pretext task alone does not guarantee that the learned representation will be useful for every target application.

46. Why is batch normalization or other normalization sometimes used in Self-Supervised Learning architectures?

  1. To remove all learned representations from the network
  2. To guarantee perfect accuracy regardless of the dataset
  3. To help control activation distributions and support stable optimization in suitable architectures
  4. To convert every unlabeled example into a human-labeled example

Answer: C) To help control activation distributions and support stable optimization in suitable architectures

Explanation:

Normalization techniques can improve optimization behavior by controlling the scale or distribution of activations. Their role varies by architecture and training method, and normalization alone does not guarantee stable learning or prevent representation collapse.

47. What is a key advantage of using Self-Supervised Learning for medical imaging research?

  1. It removes the need for privacy protections in medical datasets
  2. It can exploit large collections of unlabeled scans to learn useful representations before task-specific training
  3. It guarantees correct diagnoses without clinical evaluation
  4. It prevents models from learning anatomical features

Answer: B) It can exploit large collections of unlabeled scans to learn useful representations before task-specific training

Explanation:

Medical imaging datasets may contain many scans without detailed expert annotations. Self-Supervised Learning can use these scans for pretraining and potentially reduce the labeled data required for downstream tasks, but clinical validation, privacy protections, and representative evaluation remain essential.

48. Which metric is commonly used to evaluate image representations on a classification task after pretraining?

  1. The file size of the unlabeled dataset alone
  2. The number of augmentation operations applied
  3. The number of parameters in the tokenizer
  4. Classification accuracy on a held-out labeled dataset

Answer: D) Classification accuracy on a held-out labeled dataset

Explanation:

Classification accuracy measures the fraction of correctly classified examples in a labeled evaluation set. Depending on the task and class distribution, metrics such as precision, recall, F1 score, or area under a curve may also be appropriate.

49. A developer wants to train an image encoder using a large collection of unlabeled photographs. Which approach is most appropriate?

  1. Use an objective such as contrastive learning or masked image modeling to derive training signals from the photographs
  2. Assign every image a random class and treat it as ground truth
  3. Train the model without defining any learning objective
  4. Discard all photographs that do not contain manually supplied labels

Answer: A) Use an objective such as contrastive learning or masked image modeling to derive training signals from the photographs

Explanation:

Contrastive learning can train an encoder to recognize relationships between augmented views, while masked image modeling trains it to predict hidden image information. Both approaches can learn useful visual representations without requiring a manually assigned class label for every photograph.

50. A company has millions of unlabeled product images but only a small number of labeled examples for product classification. Which training strategy best uses Self-Supervised Learning?

  1. Train only on the small labeled dataset and discard all unlabeled images
  2. Assign random product categories to the unlabeled images and use them as verified labels
  3. Pretrain an image encoder on the unlabeled images using a suitable self-supervised objective, then fine-tune or evaluate it on the labeled classification dataset
  4. Skip representation learning and classify images using their file names alone

Answer: C) Pretrain an image encoder on the unlabeled images using a suitable self-supervised objective, then fine-tune or evaluate it on the labeled classification dataset

Explanation:

Self-Supervised Learning allows the encoder to learn visual patterns from the large unlabeled collection. The pretrained encoder can then be fine-tuned with the smaller labeled dataset to adapt its representations to product categories. The company should evaluate the resulting model on a separate held-out dataset and check for duplicates or leakage between training and evaluation data.

Comments and Discussions!

Load comments ↻



Copyright © 2026 www.includehelp.com. All rights reserved.