Home »
MCQs
ML with Python MCQs (Multiple-Choice Questions)
Machine Learning with Python is the use of Python and its libraries to build, train, evaluate, and deploy machine learning models. Python provides a large ecosystem for machine learning, including libraries such as scikit-learn, NumPy, pandas, SciPy, and Matplotlib.
ML with Python MCQs
These ML with Python MCQs cover important concepts such as machine learning algorithms, NumPy, pandas, scikit-learn, data preprocessing, feature scaling, encoding, missing values, training and testing datasets, model fitting, prediction, cross-validation, pipelines, hyperparameter tuning, classification, regression, clustering, and model evaluation.
List of ML with Python MCQs
The following Machine Learning with Python multiple-choice questions are useful for students, Python programmers, data science learners, machine learning beginners, and candidates preparing for technical examinations and interviews.
1. Which Python library is widely used for implementing traditional machine learning algorithms?
- scikit-learn
- Beautiful Soup
- Flask
- Requests
Answer: A) scikit-learn
Explanation:
scikit-learn is a Python machine learning library that provides tools for supervised and unsupervised learning, preprocessing, model selection, and evaluation.
2. Which library provides multidimensional numerical arrays commonly used in Python machine learning workflows?
- NumPy
- Flask
- Requests
- Beautiful Soup
Answer: A) NumPy
Explanation:
NumPy provides the multidimensional ndarray and numerical operations that are widely used in scientific computing and machine learning.
3. Which Python library provides DataFrame and Series data structures?
- pandas
- NumPy
- Matplotlib
- scikit-learn
Answer: A) pandas
Explanation:
pandas provides data structures and data analysis tools for Python, including Series and DataFrame objects.
4. Which alias is conventionally used when importing pandas?
pd
ps
pandaslib
pds
Answer: A) pd
Explanation:
The conventional import statement is import pandas as pd.
5. Which alias is conventionally used when importing NumPy?
np
num
numpyml
npy
Answer: A) np
Explanation:
The standard convention is import numpy as np. The NumPy documentation uses this convention throughout its examples.
6. What is the primary purpose of machine learning?
- Allowing systems to learn patterns from data for prediction or decision-making
- Creating only static web pages
- Replacing all database systems
- Formatting Python source code
Answer: A) Allowing systems to learn patterns from data for prediction or decision-making
Explanation:
Machine learning uses data to train models that can make predictions, classify observations, identify patterns, or perform other data-driven tasks.
7. Which is a major category of machine learning?
- Supervised learning
- HTML learning
- Database learning
- Markup learning
Answer: A) Supervised learning
Explanation:
Supervised learning is one of the major machine learning categories. Other common categories include unsupervised learning and reinforcement learning.
8. What is provided to a supervised learning model during training?
- Features and target values
- Only feature names
- Only test predictions
- Only visualization images
Answer: A) Features and target values
Explanation:
Supervised learning uses input features and known target values to learn a relationship that can later be used for prediction.
9. What is the purpose of a training dataset?
- To train the machine learning model
- To permanently replace the test dataset
- To store only model predictions
- To visualize the final model
Answer: A) To train the machine learning model
Explanation:
The training dataset is used by the model's learning algorithm to estimate parameters or otherwise learn patterns from the data.
10. What is the main purpose of a test dataset?
- To evaluate model performance on unseen data
- To train the model repeatedly
- To replace the target variable
- To perform feature encoding only
Answer: A) To evaluate model performance on unseen data
Explanation:
A test set provides data that was not used to fit the model and can therefore be used to assess its generalization performance.
11. Which scikit-learn function is commonly used to split data into training and test subsets?
train_test_split()
split_train_test()
dataset_split()
divide_data()
Answer: A) train_test_split()
Explanation:
train_test_split() is provided by sklearn.model_selection and splits arrays or matrices into random train and test subsets.
12. Which scikit-learn module contains train_test_split()?
sklearn.model_selection
sklearn.preprocessing
sklearn.metrics
sklearn.datasets
Answer: A) sklearn.model_selection
Explanation:
The train_test_split() utility belongs to the model_selection module.
13. What does the test_size parameter of train_test_split() control?
- The size of the test subset
- The number of model features
- The learning rate
- The number of classes
Answer: A) The size of the test subset
Explanation:
test_size specifies the proportion or absolute number of samples that should be included in the test split.
14. What is the purpose of random_state in train_test_split()?
- To make the split reproducible
- To increase the number of features
- To normalize the target values
- To select the machine learning algorithm
Answer: A) To make the split reproducible
Explanation:
Providing an integer to random_state controls the randomness used in the split and allows reproducible results.
15. What does the stratify parameter help preserve during a train-test split?
- Class proportions
- Feature names
- Column data types only
- Model coefficients
Answer: A) Class proportions
Explanation:
When supplied with class labels, stratify causes the split to be performed in a stratified manner using those labels.
16. What is an estimator in scikit-learn?
- An object that can fit a model or transform data according to its purpose
- A CSV file
- A Python interpreter
- A visualization library
Answer: A) An object that can fit a model or transform data according to its purpose
Explanation:
Scikit-learn uses estimators as its central API concept. Estimators are objects that can learn from data through methods such as fit().
17. Which method is commonly used to train a scikit-learn estimator?
fit()
train()
learn()
build()
Answer: A) fit()
Explanation:
Scikit-learn estimators use the fit() method to learn from training data.
18. Which method is commonly used to generate predictions from a trained scikit-learn estimator?
predict()
forecast_model()
generate()
classify_model()
Answer: A) predict()
Explanation:
Estimators such as classifiers and regressors commonly provide predict() for generating predictions from new input data.
19. In a typical supervised learning dataset, what does X represent?
- Input features
- Target labels only
- Model predictions
- Evaluation scores
Answer: A) Input features
Explanation:
In the conventional scikit-learn notation, X represents the samples and their features, while y represents target values.
20. In a typical supervised learning dataset, what does y represent?
- Target values
- Input feature names
- Test set size
- Model parameters
Answer: A) Target values
Explanation:
In supervised learning, y generally contains the target values corresponding to the samples in X.
21. Which scikit-learn transformer is commonly used for feature standardization?
StandardScaler
StandardModel
FeatureScalerModel
NormalizeFeaturesOnly
Answer: A) StandardScaler
Explanation:
StandardScaler is commonly used to standardize numerical features before applying machine learning algorithms.
22. What does feature standardization generally do?
- Centers features and scales them according to their variability
- Deletes all features
- Converts all features into text
- Changes classification into clustering
Answer: A) Centers features and scales them according to their variability
Explanation:
Standardization commonly transforms numerical features so that their values are centered around zero and scaled according to their standard deviation.
23. Which scikit-learn transformer is commonly used to handle missing values?
SimpleImputer
MissingValueModel
NullScaler
DataCleaner
Answer: A) SimpleImputer
Explanation:
SimpleImputer provides a straightforward way to replace missing values using a selected strategy.
24. What is the purpose of imputation in machine learning?
- To replace missing values using a defined strategy
- To create a new machine learning algorithm
- To calculate model accuracy
- To remove the target variable
Answer: A) To replace missing values using a defined strategy
Explanation:
Imputation replaces missing data with values determined by a strategy such as the mean, median, most frequent value, or a constant.
25. Which scikit-learn transformer is commonly used to convert categorical features into one-hot encoded columns?
OneHotEncoder
CategoryScaler
CategoricalModel
FeatureDecoder
Answer: A) OneHotEncoder
Explanation:
OneHotEncoder transforms categorical features into a numerical representation using separate indicator columns.
26. Why is categorical encoding often required before training a machine learning model?
- Many algorithms require numerical input features
- It increases the number of target labels automatically
- It removes all missing values
- It converts the model into a neural network
Answer: A) Many algorithms require numerical input features
Explanation:
Many machine learning estimators operate on numerical feature representations, so categorical values may need to be encoded before training.
27. What does the transform() method of a scikit-learn transformer generally do?
- Transforms input data using learned or specified transformation rules
- Trains a classifier from scratch
- Calculates only accuracy
- Deletes the training dataset
Answer: A) Transforms input data using learned or specified transformation rules
Explanation:
Transformers apply a transformation to input data through transform(). Examples include scaling and encoding.
28. What does fit_transform() generally combine?
- Fitting a transformer and transforming the data
- Predicting and evaluating a model
- Training and deleting a model
- Splitting and shuffling only
Answer: A) Fitting a transformer and transforming the data
Explanation:
fit_transform() combines learning the transformation parameters from data with applying the transformation.
29. What is data leakage in machine learning?
- When information from data that should be unavailable influences model training
- When a CSV file is too large
- When a model has no target variable
- When a Python script contains comments
Answer: A) When information from data that should be unavailable influences model training
Explanation:
Data leakage occurs when information that should not be available during training influences the fitted model, potentially producing overly optimistic evaluation results.
30. Why can a scikit-learn Pipeline help prevent data leakage?
- It keeps preprocessing and modeling steps together and fits transformations using the appropriate training data
- It permanently deletes the test set
- It removes the need for evaluation
- It guarantees perfect predictions
Answer: A) It keeps preprocessing and modeling steps together and fits transformations using the appropriate training data
Explanation:
A Pipeline chains transformations and an estimator so that preprocessing can be fitted as part of the training workflow rather than being independently fitted on the complete dataset.
31. Which scikit-learn class is used to create a sequence of preprocessing steps and an estimator?
Pipeline
Workflow
ModelChain
EstimatorSequence
Answer: A) Pipeline
Explanation:
scikit-learn's Pipeline chains transformers and a final estimator into a single object that follows the estimator API.
32. Which machine learning algorithm is commonly used for binary classification?
- Logistic Regression
- Linear Regression only
- K-Means only
- PCA
Answer: A) Logistic Regression
Explanation:
Logistic Regression is a supervised learning algorithm commonly used for classification problems, particularly binary classification.
33. Which scikit-learn class implements Logistic Regression?
LogisticRegression
LogisticModel
ClassificationRegression
BinaryRegression
Answer: A) LogisticRegression
Explanation:
scikit-learn provides Logistic Regression through the LogisticRegression estimator.
34. Which algorithm is primarily used to predict continuous numerical values?
- Linear Regression
- Logistic Regression
- K-Means
- K-Nearest Neighbors Classifier
Answer: A) Linear Regression
Explanation:
Linear Regression is a supervised learning algorithm commonly used to predict continuous target values.
35. Which scikit-learn class implements ordinary linear regression?
LinearRegression
LinearModel
RegressionModel
OrdinaryRegression
Answer: A) LinearRegression
Explanation:
scikit-learn provides ordinary least squares linear regression through the LinearRegression estimator.
36. Which algorithm groups data points into clusters without requiring target labels?
- K-Means
- Linear Regression
- Logistic Regression
- Decision Tree Regression
Answer: A) K-Means
Explanation:
K-Means is an unsupervised clustering algorithm that partitions observations into a specified number of clusters.
37. Which scikit-learn class is commonly used for K-Means clustering?
KMeans
KCluster
KMeansModel
ClusterK
Answer: A) KMeans
Explanation:
The scikit-learn clustering module provides the KMeans estimator for K-Means clustering.
38. Which algorithm predicts the class of a sample using nearby training examples?
- K-Nearest Neighbors
- Linear Regression
- Principal Component Analysis
- K-Means only
Answer: A) K-Nearest Neighbors
Explanation:
K-Nearest Neighbors predicts a sample based on nearby training examples according to a selected distance measure and neighbor count.
39. Which scikit-learn class implements K-Nearest Neighbors classification?
KNeighborsClassifier
KNNClassifier
KNearestClassifier
NeighborClassifier
Answer: A) KNeighborsClassifier
Explanation:
scikit-learn provides the KNeighborsClassifier estimator for classification using the nearest-neighbor approach.
40. Which algorithm uses a tree structure to make predictions?
- Decision Tree
- K-Means
- Principal Component Analysis
- StandardScaler
Answer: A) Decision Tree
Explanation:
Decision Tree models recursively split data according to selected criteria and use the resulting tree structure for prediction.
41. Which ensemble algorithm combines multiple decision trees?
- Random Forest
- Linear Regression
- K-Means
- StandardScaler
Answer: A) Random Forest
Explanation:
Random Forest is an ensemble method that combines predictions from multiple decision trees.
42. Which scikit-learn class is used for a random forest classifier?
RandomForestClassifier
RandomTreeClassifier
ForestClassifier
RandomModelClassifier
Answer: A) RandomForestClassifier
Explanation:
The RandomForestClassifier estimator implements a random forest for classification.
43. What is cross-validation used for in machine learning?
- Evaluating model performance across multiple data splits
- Converting text into images
- Removing all features
- Replacing the training dataset permanently
Answer: A) Evaluating model performance across multiple data splits
Explanation:
Cross-validation repeatedly divides available training data into training and validation portions to provide a more robust estimate of model performance.
44. What does K-fold cross-validation divide the data into?
- K folds
- K target variables
- K machine learning libraries
- K feature names
Answer: A) K folds
Explanation:
In K-fold cross-validation, the data is divided into K subsets, with different subsets used for validation across different iterations.
45. What does cross_val_score() return?
- Scores obtained from the cross-validation folds
- Only the trained model
- Only the feature names
- A list of missing values
Answer: A) Scores obtained from the cross-validation folds
Explanation:
cross_val_score() evaluates an estimator using cross-validation and returns an array containing the score from each cross-validation split.
46. What is hyperparameter tuning?
- Finding suitable values for model configuration parameters
- Changing the target labels after prediction
- Removing all preprocessing
- Converting numerical features to images
Answer: A) Finding suitable values for model configuration parameters
Explanation:
Hyperparameters are configuration values set before or outside the model's fitting process. Hyperparameter tuning searches for values that provide good model performance.
47. Which scikit-learn tool can perform an exhaustive search over specified hyperparameter combinations?
GridSearchCV
ParameterSearch
GridModel
SearchCVModel
Answer: A) GridSearchCV
Explanation:
GridSearchCV evaluates specified combinations of hyperparameter values using cross-validation.
48. Which metric measures the proportion of correct predictions among all predictions?
- Accuracy
- Recall
- Precision
- Mean absolute error
Answer: A) Accuracy
Explanation:
For classification, accuracy is the proportion of predictions that match the actual class labels.
49. Which scikit-learn function is commonly used to calculate classification accuracy?
accuracy_score()
calculate_accuracy()
accuracy()
score_accuracy()
Answer: A) accuracy_score()
Explanation:
scikit-learn provides accuracy_score() in its metrics module for calculating classification accuracy.
50. Which statement best describes Machine Learning with Python?
- It combines Python programming with libraries and tools for building machine learning workflows
- It is a single Python function
- It is a Python data type
- It is a replacement for the Python interpreter
Answer: A) It combines Python programming with libraries and tools for building machine learning workflows
Explanation:
Machine Learning with Python is an ecosystem and development approach rather than a single technology. Python libraries such as scikit-learn, NumPy, and pandas provide tools for preparing data, training models, evaluating results, and building machine learning workflows.
Advertisement
Advertisement