Note
This page is a reference documentation. It only explains the class signature, and not how to use it. Please refer to the user guide for the big picture.
nidl.backbones.volume.MoEParams¶
- class nidl.backbones.volume.MoEParams(dim=768, n_shared_experts=2, n_routed_experts=16, n_activated_experts=6, moe_inter_dim=384, score_func='softmax', route_scale=4.0, bias_clip=0.3, bias_update_rate=0.0001, moe_layer_indices=(1, 3, 5, 7, 9, 11))[source]¶
Bases:
objectHyperparameters of the sparse Mixture-of-Experts (MoE) layers used by VisionTransformer3DMoE when use_moe=True.
- Parameters:
- dimint, default=768
Token embedding dimension (must match the encoder’s embed_dim).
- n_shared_expertsint, default=2
Number of “shared” experts, always active for every token.
- n_routed_expertsint, default=16
Number of routed experts to choose from.
- n_activated_expertsint, default=6
Number of routed experts activated (top-k) per token.
- moe_inter_dimint, default=384
Hidden dimension of each expert’s MLP.
- score_func{“softmax”, “sigmoid”}, default=”softmax”
Function used to turn router logits into routing scores.
- route_scalefloat, default=4.0
Scalar applied to the routing weights.
- bias_clipfloat, default=0.3
Passed to moe_bias_update as bias_clip.
- bias_update_ratefloat, default=1e-4
Passed to moe_bias_update as update_rate.
- moe_layer_indicestuple of int, default=(1, 3, 5, 7, 9, 11)
0-indexed block indices that use a sparse MoE instead of a dense MLP.
- __init__(dim=768, n_shared_experts=2, n_routed_experts=16, n_activated_experts=6, moe_inter_dim=384, score_func='softmax', route_scale=4.0, bias_clip=0.3, bias_update_rate=0.0001, moe_layer_indices=(1, 3, 5, 7, 9, 11))¶
- bias_clip = 0.3¶
- bias_update_rate = 0.0001¶
- dim = 768¶
- moe_inter_dim = 384¶
- moe_layer_indices = (1, 3, 5, 7, 9, 11)¶
- n_activated_experts = 6¶
- n_routed_experts = 16¶
- route_scale = 4.0¶
- score_func = 'softmax'¶