Home »
Trending Technologies MCQs
Mixture of Experts (MoE) MCQs (Multiple-Choice Questions)
Practice Mixture of Experts (MoE) MCQs to test your knowledge of neural network architectures that use multiple specialized expert networks to process input data. These questions cover expert networks, gating mechanisms, token routing, sparse computation, load balancing, and model training. They are useful for students, AI developers, machine learning engineers, and professionals preparing for technical interviews or assessments. The set includes both foundational and practical questions covering modern Mixture of Experts architectures.
Mixture of Experts (MoE) MCQs
These Mixture of Experts (MoE) multiple-choice questions cover important concepts such as expert selection, gating networks, top-k routing, sparse activation, expert parallelism, training objectives, and computational efficiency. They also explore routing strategies, load balancing, distributed inference, model scaling, and common implementation challenges. This set combines conceptual, technical, and scenario-based questions to test your understanding of Mixture of Experts systems.
Mixture of Experts MCQs cover the technologies used to build scalable neural networks by routing inputs to selected specialized components. Each question includes an answer and explanation.
List of Mixture of Experts (MoE) MCQs
The following Mixture of Experts multiple-choice questions cover MoE fundamentals, architecture, routing, training, optimization, deployment, and real-world applications.
1. What is the primary purpose of a Mixture of Experts (MoE) architecture?
- To process every input using only one fixed neuron
- To replace neural networks with relational databases
- To combine multiple expert networks and select suitable experts for each input
- To eliminate the need for model training
Answer: C) To combine multiple expert networks and select suitable experts for each input
Explanation:
A Mixture of Experts architecture contains multiple expert networks and a mechanism that routes inputs to selected experts. This allows different inputs or tokens to use different subsets of the model's parameters instead of activating every expert for every input.
2. What is an expert in a Mixture of Experts model?
- A neural network component that processes inputs routed to it
- A dataset used exclusively for model evaluation
- A hardware device responsible only for storing files
- A fixed rule that prevents model parameters from changing
Answer: A) A neural network component that processes inputs routed to it
Explanation:
An expert is a learnable neural network that processes the inputs assigned to it by the routing mechanism. In transformer-based MoE models, experts often replace the feed-forward network in selected transformer layers.
3. What is the main function of a gating network in an MoE model?
- To store training data permanently
- To increase the size of every input sequence
- To remove the model's attention mechanism
- To determine which experts should process an input
Answer: D) To determine which experts should process an input
Explanation:
The gating network, often called the router in modern implementations, computes scores for available experts. A routing strategy uses these scores to select one or more experts and may assign weights to their outputs.
4. What does sparse activation mean in a Mixture of Experts model?
- All experts must process every token
- Only a subset of available experts is activated for a given input
- The model does not contain trainable parameters
- Only the input embedding layer is trained
Answer: B) Only a subset of available experts is activated for a given input
Explanation:
Sparse activation means that a routing mechanism selects only some experts for each input or token. This can reduce the computation required per token compared with activating every expert, even when the total number of model parameters is large.
5. What is top-k routing in an MoE architecture?
- Selecting the k experts with the highest routing scores
- Selecting k training datasets at random
- Activating all experts regardless of their scores
- Removing the k largest model parameters
Answer: A) Selecting the k experts with the highest routing scores
Explanation:
Top-k routing selects the k experts with the highest scores for a particular input or token. For example, top-2 routing selects two experts, although the exact implementation may also apply capacity limits, normalization, or other routing constraints.
6. How does a sparse MoE model differ from a dense neural network?
- A sparse MoE model cannot learn from training data
- A dense neural network always contains more parameters
- A sparse MoE model activates selected experts rather than all expert networks for each input
- A dense neural network cannot process language
Answer: C) A sparse MoE model activates selected experts rather than all expert networks for each input
Explanation:
A conventional dense layer generally applies its learned transformation to every relevant input. A sparse MoE layer routes each input to selected experts, allowing a model to maintain a large total parameter count without requiring all expert computations for every token.
7. What is a major potential benefit of using MoE in large language models?
- It guarantees zero inference latency
- It allows a large total parameter capacity with comparatively limited active computation per token
- It eliminates the need for GPUs or other accelerators
- It guarantees that every expert specializes in a unique subject
Answer: B) It allows a large total parameter capacity with comparatively limited active computation per token
Explanation:
MoE architectures can scale the total number of parameters while activating only a subset of experts for each token. This can improve model capacity relative to active computation, although communication, memory requirements, and routing overhead can reduce the practical benefits.
8. What does the term "active parameters" mean in a sparse MoE model?
- The total number of parameters ever created during training
- The number of parameters stored on disk
- The number of parameters that have been permanently frozen
- The parameters participating in the computation for a particular input or token
Answer: D) The parameters participating in the computation for a particular input or token
Explanation:
Active parameters are the parameters used in a particular forward computation. In an MoE model, only the selected experts contribute their expert-specific parameters for a given token, while shared layers and other common components may also be active.
9. What is expert load balancing in an MoE model?
- Distributing inputs across experts to avoid excessive concentration on a few experts
- Ensuring that all experts have identical parameter values
- Assigning every token to the same expert
- Reducing the number of input tokens to zero
Answer: A) Distributing inputs across experts to avoid excessive concentration on a few experts
Explanation:
Load balancing aims to prevent some experts from receiving too many tokens while others remain underused. Balanced routing can improve device utilization, reduce congestion, and help training proceed without severe bottlenecks.
10. What can happen if an MoE router sends most tokens to a single expert?
- The model automatically becomes a dense neural network
- Every expert receives exactly the same number of tokens
- The heavily selected expert can become a bottleneck while other experts are underutilized
- The model no longer needs to perform backpropagation
Answer: C) The heavily selected expert can become a bottleneck while other experts are underutilized
Explanation:
Uneven routing can overload a small number of experts and leave other experts with little useful work. This can reduce parallelism, create communication bottlenecks, and make it more difficult for the experts to develop useful, complementary behavior.
11. What is an auxiliary load-balancing loss used for in MoE training?
- To increase the vocabulary size of the tokenizer
- To encourage a more balanced distribution of tokens across experts
- To replace the main language modeling objective completely
- To remove all expert-specific parameters
Answer: B) To encourage a more balanced distribution of tokens across experts
Explanation:
An auxiliary load-balancing loss encourages the router to distribute tokens more evenly across experts. It is commonly combined with the main training objective, with its weight chosen carefully because excessive balancing pressure can interfere with useful routing decisions.
12. In a transformer-based MoE model, where are experts commonly used?
- Only in the computer's operating system
- Exclusively in the tokenizer's vocabulary file
- Only in the network interface controller
- In place of feed-forward network blocks in selected transformer layers
Answer: D) In place of feed-forward network blocks in selected transformer layers
Explanation:
Many transformer-based MoE architectures replace the standard feed-forward sublayer in selected layers with multiple expert feed-forward networks and a router. Attention and other shared transformer components can remain dense and operate as usual.
13. What is a token router in a transformer-based MoE model?
- A component that determines which expert or experts receive each token
- A component that converts all tokens into images
- A storage system for model checkpoints
- A process that permanently deletes input tokens
Answer: A) A component that determines which expert or experts receive each token
Explanation:
A token router calculates expert-selection scores for token representations. Depending on the architecture, it selects one or more experts, dispatches tokens to them, and combines their outputs according to routing weights or another defined aggregation rule.
14. How are the outputs of multiple selected experts commonly combined?
- By discarding every expert output except the first one in all architectures
- By concatenating the outputs into a new vocabulary in every case
- By combining expert outputs using routing weights or another architecture-defined rule
- By replacing the outputs with randomly generated values
Answer: C) By combining expert outputs using routing weights or another architecture-defined rule
Explanation:
When multiple experts are selected, their outputs are commonly weighted and summed to produce the MoE layer output. The exact aggregation method depends on the model design, and some implementations use specialized routing or combination rules.
15. What is top-1 routing?
- Routing every token to every expert
- Routing each token to the expert with the highest routing score
- Selecting the expert with the fewest parameters regardless of the input
- Using the first expert in a fixed list for all tokens
Answer: B) Routing each token to the expert with the highest routing score
Explanation:
Top-1 routing selects a single expert with the highest routing score for each token. It can limit expert computation and simplify dispatch, but the model must still manage routing balance and possible expert capacity constraints.
16. What is expert capacity in an MoE model?
- The total number of training examples in the dataset
- The maximum context window supported by the tokenizer
- The number of GPUs connected to the network, regardless of routing
- The maximum number of tokens an expert is configured to process in a routing batch
Answer: D) The maximum number of tokens an expert is configured to process in a routing batch
Explanation:
Expert capacity limits how many tokens an expert can accept within a batch or routing group. Capacity limits help control memory use and computation, but they may cause token overflow when too many inputs are assigned to the same expert.
17. What is token dropping in an MoE architecture?
- Discarding or skipping tokens that cannot be processed because routing capacity is exceeded
- Removing all tokens from the training vocabulary permanently
- Deleting every token before it reaches the router
- Replacing all expert outputs with padding tokens
Answer: A) Discarding or skipping tokens that cannot be processed because routing capacity is exceeded
Explanation:
Some MoE implementations drop or skip tokens that exceed expert capacity. This can reduce overload, but it may harm model quality if important tokens are lost. Alternative approaches include increasing capacity or using routing and dispatch methods designed to avoid token dropping.
18. What is expert parallelism?
- Replicating every training example on the same processor
- Running all model operations sequentially on one CPU core
- Distributing different experts across multiple devices to execute their computations in parallel
- Storing all expert parameters in a single text document
Answer: C) Distributing different experts across multiple devices to execute their computations in parallel
Explanation:
Expert parallelism places different experts on different GPUs or other accelerators. Tokens are dispatched to the devices hosting their selected experts, allowing expert computations to run in parallel while introducing communication and synchronization requirements.
19. Why can communication overhead be significant in distributed MoE models?
- Expert parallelism eliminates all movement of data
- Tokens may need to move between devices to reach their selected experts
- Each expert must be stored on a separate public website
- Routing does not require any input representations
Answer: B) Tokens may need to move between devices to reach their selected experts
Explanation:
In distributed MoE systems, token representations often travel from the device performing routing to the devices hosting selected experts. These all-to-all communication operations can consume bandwidth and time, especially when expert placement and routing patterns are not well matched to the hardware.
20. What is an all-to-all communication operation in distributed MoE training?
- A method for deleting unused expert parameters
- A technique that guarantees no network communication is required
- A process that sends every token only to the first device
- A communication pattern in which multiple devices exchange data with one another
Answer: D) A communication pattern in which multiple devices exchange data with one another
Explanation:
All-to-all communication allows each participating device to send data to other devices and receive data from them. MoE implementations use this pattern to dispatch tokens to expert-owning devices and return the expert outputs for subsequent computation.
21. What is a shared expert in some MoE architectures?
- An expert component that is available to process tokens independently of the ordinary routed expert selection
- An expert that is permanently disabled during inference
- A component that stores only the training labels
- An expert that must be copied manually for every token
Answer: A) An expert component that is available to process tokens independently of the ordinary routed expert selection
Explanation:
Some MoE designs include shared experts that process tokens in addition to the routed experts. They can capture broadly useful patterns while routed experts handle other computations, although the exact arrangement varies by architecture.
22. What is a key difference between dense and sparse MoE layers in terms of expert computation?
- Dense layers cannot contain trainable parameters
- Sparse MoE layers always require fewer total parameters
- Sparse MoE layers select a subset of experts for each input rather than evaluating all experts
- Dense layers can process only a single token in their lifetime
Answer: C) Sparse MoE layers select a subset of experts for each input rather than evaluating all experts
Explanation:
A sparse MoE layer routes inputs to a limited number of experts, whereas a dense mixture evaluates every expert in the mixture before combining their outputs. Sparse routing can reduce expert computation per input, but it adds routing and dispatch complexity.
23. What does expert specialization mean in an MoE model?
- Every expert is forced to learn the same input-output mapping
- Different experts develop different useful behaviors or respond to different input patterns
- Each expert processes only one training example throughout training
- All experts use fixed parameters that never change
Answer: B) Different experts develop different useful behaviors or respond to different input patterns
Explanation:
Expert specialization occurs when different experts learn useful differences in the inputs they process or the transformations they perform. Specialization is an emergent result of training and routing; it is not guaranteed that experts will align neatly with human-defined topics.
24. How are routing decisions commonly computed in a learned MoE router?
- By assigning every input to an expert using a manually typed list only
- By measuring the physical temperature of each GPU
- By choosing experts solely according to their names
- By applying a learned transformation to the input representation to obtain expert scores
Answer: D) By applying a learned transformation to the input representation to obtain expert scores
Explanation:
A typical router applies a learned projection to a token representation to produce scores for the available experts. A routing rule then selects experts from these scores, sometimes using a softmax distribution, top-k selection, or additional balancing mechanisms.
25. What is the role of a softmax function in many MoE routers?
- It converts routing scores into normalized nonnegative weights that sum to one
- It permanently freezes the expert networks
- It removes all low-level features from the input
- It guarantees perfect load balancing for every batch
Answer: A) It converts routing scores into normalized nonnegative weights that sum to one
Explanation:
Softmax converts a vector of scores into a probability-like distribution whose entries are nonnegative and sum to one. MoE routers may use these values to rank experts or weight their outputs, although some architectures use alternative scoring or normalization methods.
26. Why can training an MoE model be more challenging than training a comparable dense model?
- MoE models cannot use gradient-based optimization
- MoE models have no trainable weights
- Training must manage routing, load imbalance, expert capacity, and distributed communication
- Every expert must be trained using a different programming language
Answer: C) Training must manage routing, load imbalance, expert capacity, and distributed communication
Explanation:
MoE training introduces challenges beyond ordinary neural network optimization. Routing decisions can produce uneven token assignments, capacity limits can affect token processing, and expert parallelism may require expensive communication across devices.
27. What is one reason expert utilization metrics are monitored during MoE training?
- To determine the color of the model's user interface
- To identify whether some experts receive too few or too many tokens
- To replace the model's training objective
- To guarantee that the model needs no evaluation
Answer: B) To identify whether some experts receive too few or too many tokens
Explanation:
Expert utilization metrics reveal how tokens and computational work are distributed among experts. They can expose routing collapse, persistent imbalance, or unused capacity and help engineers adjust routing strategies, loss coefficients, or system configuration.
28. What is routing collapse in a Mixture of Experts model?
- A condition in which every expert develops a unique skill automatically
- A process that increases the number of active experts without limits
- A technique for reducing model checkpoints to text files
- A failure mode in which routing becomes concentrated on a small subset of experts
Answer: D) A failure mode in which routing becomes concentrated on a small subset of experts
Explanation:
Routing collapse occurs when the router repeatedly favors a small number of experts, leaving others underused. This can reduce the benefits of the architecture and may require training adjustments, load-balancing objectives, or changes to routing behavior.
29. What is a routing logit in an MoE model?
- A raw score produced by the router for a potential expert assignment
- A final natural-language answer produced by the model
- A log file containing operating system errors only
- A fixed label assigned to each training document
Answer: A) A raw score produced by the router for a potential expert assignment
Explanation:
Routing logits are unnormalized scores associated with experts for a given input. The model uses these scores, often after normalization or ranking, to decide which experts should process the token and how their outputs may be weighted.
30. What is a potential drawback of increasing the number of experts in an MoE model?
- It always reduces total model storage requirements
- It eliminates the need for routing decisions
- It can increase parameter storage, communication demands, and deployment complexity
- It guarantees that inference will become slower by exactly the same percentage
Answer: C) It can increase parameter storage, communication demands, and deployment complexity
Explanation:
Adding experts increases the model's total parameter capacity but also requires more memory to store their weights. Depending on placement and routing, it can increase communication overhead and complicate deployment, even if the number of active experts per token remains fixed.
31. What is the difference between total parameters and active parameters in an MoE model?
- Total parameters count only the router, while active parameters count the tokenizer
- Total parameters include all model components, while active parameters refer to those used in a particular computation
- Total parameters refer only to the current input, while active parameters refer to every expert ever trained
- The two terms always refer to exactly the same number
Answer: B) Total parameters include all model components, while active parameters refer to those used in a particular computation
Explanation:
Total parameter count includes the model's complete set of learned parameters, including experts that are not selected for a given token. Active parameter count measures the parameters participating in the computation for a particular input, including shared layers and selected experts as applicable.
32. How can MoE architectures support model scaling?
- By removing all expert networks when the model grows
- By ensuring that every token activates every parameter in the model
- By replacing the training process with a fixed rule table
- By increasing the number or size of experts while routing each input to a limited subset
Answer: D) By increasing the number or size of experts while routing each input to a limited subset
Explanation:
MoE models can increase their total parameter capacity by adding experts or enlarging expert networks. Sparse routing limits the number of experts used for each token, helping manage active computation, although the benefits depend on hardware, training quality, and communication costs.
33. What is the purpose of router regularization in MoE training?
- To encourage desirable routing behavior or control undesirable routing patterns
- To force the model to remove all expert networks
- To make all expert outputs random
- To eliminate the need for a training dataset
Answer: A) To encourage desirable routing behavior or control undesirable routing patterns
Explanation:
Router regularization can help shape expert assignment patterns and discourage undesirable behavior such as extreme imbalance. Different approaches target different objectives, and regularization must be chosen carefully so that balanced routing does not override useful input-dependent specialization.
34. Why might an MoE model use different experts for different tokens in the same sentence?
- Every token must have a unique expert that is never reused
- The router assigns experts based only on token position
- Each token representation can receive different routing scores based on its contextual information
- Experts cannot process more than one token during training
Answer: C) Each token representation can receive different routing scores based on its contextual information
Explanation:
In token-level MoE routing, each token representation is evaluated by the router and can be assigned to a different subset of experts. As token representations reflect different words and contexts, routing decisions can vary within the same sequence.
35. What is a common advantage of using mixed precision during MoE inference?
- It eliminates the need to store expert weights
- It can reduce memory use and improve computational throughput on supported hardware
- It guarantees that numerical results are identical to full precision
- It removes the router from the model architecture
Answer: B) It can reduce memory use and improve computational throughput on supported hardware
Explanation:
Mixed-precision computation uses lower-precision representations for suitable operations, potentially reducing memory requirements and improving throughput. Its impact depends on the hardware, numerical format, model implementation, and the need to preserve model accuracy.
36. What is one challenge of serving a large MoE model on multiple GPUs?
- Every GPU must generate a different user response
- The model cannot use batching under any circumstances
- All experts must have identical learned parameters
- Expert placement and token communication must be managed to avoid performance bottlenecks
Answer: D) Expert placement and token communication must be managed to avoid performance bottlenecks
Explanation:
Large MoE deployments distribute expert weights across devices and route tokens to the appropriate locations. Poor placement, limited network bandwidth, imbalanced expert workloads, and insufficient memory can reduce throughput and increase latency.
37. What is expert replication in an MoE deployment?
- Creating additional copies of selected expert weights to improve availability or reduce communication bottlenecks
- Replacing every expert with a random number generator
- Copying the training dataset into the tokenizer vocabulary
- Forcing every input token to use all expert copies simultaneously
Answer: A) Creating additional copies of selected expert weights to improve availability or reduce communication bottlenecks
Explanation:
Expert replication places copies of an expert on multiple devices or locations. This can help serve frequently selected experts and reduce some communication or load bottlenecks, but it consumes additional memory and requires a suitable routing and scheduling strategy.
38. How does batch size affect MoE execution?
- Batch size never affects routing or expert utilization
- A larger batch always makes every expert equally busy
- Batch size changes the number of tokens available for routing and can affect utilization, memory, and communication efficiency
- Batch size determines the number of learned parameters in each expert
Answer: C) Batch size changes the number of tokens available for routing and can affect utilization, memory, and communication efficiency
Explanation:
The batch size influences how many tokens are routed through the experts in a processing step. Larger batches may improve hardware utilization and amortize some overhead, but they can also increase memory use, affect capacity requirements, and change the distribution of tokens across experts.
39. Why is the router important to the overall quality of an MoE model?
- The router is responsible only for naming the experts
- Its decisions determine which expert computations contribute to processing each input
- The router replaces all learned parameters during inference
- The router ensures that experts never make mistakes
Answer: B) Its decisions determine which expert computations contribute to processing each input
Explanation:
The router influences which experts process each token and, in many architectures, how their outputs are combined. Poor routing can send tokens to unsuitable experts or create load imbalance, limiting the benefits of having a large collection of expert networks.
40. What is one reason MoE models may require more memory than dense models with similar active computation?
- MoE models must store every possible input sequence permanently
- MoE models cannot compress their training datasets
- MoE models always store duplicate copies of every token
- They may contain many expert parameters, including experts not activated for a particular token
Answer: D) They may contain many expert parameters, including experts not activated for a particular token
Explanation:
Sparse routing limits active expert computation but does not remove the need to store or otherwise access the model's expert weights. A large MoE model can therefore require substantial accelerator or distributed memory even when only a fraction of its parameters are active for each token.
41. What is a key difference between token-choice routing and expert-choice routing?
- Token-choice routing selects experts for each token, while expert-choice routing lets experts select tokens to process
- Token-choice routing cannot use learned scores
- Expert-choice routing requires every token to be processed by every expert
- Both methods require all routing decisions to be entered manually
Answer: A) Token-choice routing selects experts for each token, while expert-choice routing lets experts select tokens to process
Explanation:
In token-choice routing, each token selects a limited number of experts based on its routing scores. In expert-choice routing, each expert selects tokens from a group of candidates, which can help control expert workload but may result in different token-processing patterns.
42. What is one potential benefit of expert-choice routing?
- It eliminates all model parameters outside the router
- It guarantees that every token receives exactly the same computation
- It can provide more direct control over the number of tokens assigned to each expert
- It prevents experts from specializing
Answer: C) It can provide more direct control over the number of tokens assigned to each expert
Explanation:
Expert-choice routing can impose a fixed or controlled token capacity for each expert, which may improve workload balance. However, it changes how tokens are assigned and may require additional mechanisms to ensure that token coverage and model quality meet the task's requirements.
43. What is a potential risk of reducing the number of experts selected per token?
- The model must always increase its total parameter count
- The model may lose useful contributions from experts that could have helped process the input
- The tokenizer automatically stops producing tokens
- The router becomes unnecessary in every architecture
Answer: B) The model may lose useful contributions from experts that could have helped process the input
Explanation:
Selecting fewer experts can reduce active computation and communication, but it may also restrict the range of expert transformations contributing to each token. The appropriate routing value depends on the architecture, task, training method, and performance constraints.
44. Why should MoE models be evaluated on multiple types of tasks?
- To ensure that every expert receives the same training examples
- To avoid measuring the model's output quality
- To prove that all experts have identical capabilities
- To assess whether the model performs well across different tasks and input distributions
Answer: D) To assess whether the model performs well across different tasks and input distributions
Explanation:
Performance on one benchmark does not establish general effectiveness across language understanding, mathematics, coding, reasoning, or other tasks. Broader evaluation can reveal strengths, weaknesses, and differences in how routing and expert specialization affect model behavior.
45. What is a practical method for diagnosing an MoE routing bottleneck?
- Monitor per-expert token counts, utilization, latency, and communication metrics
- Measure only the size of the model's tokenizer vocabulary
- Disable all performance monitoring tools
- Check only the name of the model checkpoint
Answer: A) Monitor per-expert token counts, utilization, latency, and communication metrics
Explanation:
Per-expert metrics can show whether some experts are overloaded or underutilized. Combining routing statistics with latency, memory, and communication measurements helps distinguish routing imbalance from hardware or network bottlenecks.
46. What is a common trade-off when choosing the expert capacity factor?
- A larger capacity factor always decreases memory use
- A smaller capacity factor guarantees that no tokens are dropped
- A larger capacity factor can reduce overflow but may increase memory and computation requirements
- Capacity factors affect only the tokenizer and never affect experts
Answer: C) A larger capacity factor can reduce overflow but may increase memory and computation requirements
Explanation:
The capacity factor influences the token capacity allocated to each expert relative to a baseline estimate. Increasing it can reduce the chance that an expert exceeds its capacity, but may require more memory or computation; the appropriate value depends on routing balance and implementation details.
47. Why can MoE inference latency remain high even when only a few experts are activated per token?
- Activating fewer experts always requires more arithmetic operations
- Routing overhead, expert communication, memory access, and synchronization can offset computation savings
- MoE models cannot run on modern accelerators
- Every token must be stored permanently before inference begins
Answer: B) Routing overhead, expert communication, memory access, and synchronization can offset computation savings
Explanation:
Sparse activation reduces the number of expert computations, but serving a distributed MoE model may require token dispatch, communication, and synchronization. If these operations are costly, the system may not achieve the latency improvements suggested by active parameter counts alone.
48. Which statement best describes the relationship between MoE and transfer learning?
- MoE prevents pretrained models from being adapted to new tasks
- Transfer learning can only be used with models that contain no experts
- MoE and transfer learning are identical concepts
- An MoE model can be pretrained and later adapted to new tasks through fine-tuning or other adaptation methods
Answer: D) An MoE model can be pretrained and later adapted to new tasks through fine-tuning or other adaptation methods
Explanation:
Mixture of Experts describes a model architecture, while transfer learning describes reusing learned knowledge for a new task or domain. An MoE model can be adapted through fine-tuning or other techniques, depending on which parameters are updated and how routing behavior is managed.
49. An MoE model has 16 experts and uses top-2 routing for each token. How many experts are selected for a token before any capacity-related adjustments?
- 2 experts
- 8 experts
- 16 experts
- 32 experts
Answer: A) 2 experts
Explanation:
Top-2 routing selects two experts for each token from the available set of 16 experts. The total number of available experts does not change the top-k selection count, although capacity limits or other implementation rules may affect the final assignments.
50. A company deploys a large MoE language model across multiple GPUs. During inference, a few experts receive most of the tokens, network communication increases, and response times become inconsistent. Which solution is most appropriate?
- Increase the number of experts without measuring utilization or communication
- Route every token to every expert and disable performance monitoring
- Analyze routing balance and per-expert workloads, optimize expert placement and communication, and adjust routing or capacity settings
- Remove all expert networks and keep the router unchanged
Answer: C) Analyze routing balance and per-expert workloads, optimize expert placement and communication, and adjust routing or capacity settings
Explanation:
The symptoms suggest both routing imbalance and distributed execution bottlenecks. Engineers should inspect per-expert token counts, capacity overflow, device utilization, communication time, and latency. They can then tune load-balancing mechanisms, expert placement, routing capacity, or expert replication and verify the impact through controlled performance tests.