Note

This page is a reference documentation. It only explains the class signature, and not how to use it. Please refer to the user guide for the big picture.

nidl.backbones.volume.MoEParams

class nidl.backbones.volume.MoEParams(dim=768, n_shared_experts=2, n_routed_experts=16, n_activated_experts=6, moe_inter_dim=384, score_func='softmax', route_scale=4.0, bias_clip=0.3, bias_update_rate=0.0001, moe_layer_indices=(1, 3, 5, 7, 9, 11))[source]

Bases: object

Hyperparameters of the sparse Mixture-of-Experts (MoE) layers used by VisionTransformer3DMoE when use_moe=True.

Parameters:
dimint, default=768

Token embedding dimension (must match the encoder’s embed_dim).

n_shared_expertsint, default=2

Number of “shared” experts, always active for every token.

n_routed_expertsint, default=16

Number of routed experts to choose from.

n_activated_expertsint, default=6

Number of routed experts activated (top-k) per token.

moe_inter_dimint, default=384

Hidden dimension of each expert’s MLP.

score_func{“softmax”, “sigmoid”}, default=”softmax”

Function used to turn router logits into routing scores.

route_scalefloat, default=4.0

Scalar applied to the routing weights.

bias_clipfloat, default=0.3

Passed to moe_bias_update as bias_clip.

bias_update_ratefloat, default=1e-4

Passed to moe_bias_update as update_rate.

moe_layer_indicestuple of int, default=(1, 3, 5, 7, 9, 11)

0-indexed block indices that use a sparse MoE instead of a dense MLP.

__init__(dim=768, n_shared_experts=2, n_routed_experts=16, n_activated_experts=6, moe_inter_dim=384, score_func='softmax', route_scale=4.0, bias_clip=0.3, bias_update_rate=0.0001, moe_layer_indices=(1, 3, 5, 7, 9, 11))
bias_clip = 0.3
bias_update_rate = 0.0001
dim = 768
moe_inter_dim = 384
moe_layer_indices = (1, 3, 5, 7, 9, 11)
n_activated_experts = 6
n_routed_experts = 16
n_shared_experts = 2
route_scale = 4.0
score_func = 'softmax'