core/nn/moe_streaming library
AirLLM-style per-expert streaming for a single MoE FFN layer,
factored out of the demo bins so the same code powers both the
synthetic-weights runner (bin/moe_streaming_demo.dart) and the
real-weights runners
(bin/deepseek_v2_moe_layer_stream.dart, bin/qwen15_moe_layer_stream.dart).
Given a ShardedSafeTensorsReader and the HF key prefix for the
layer's mlp submodule (e.g. "model.layers.5.mlp"), it can:
- Load the router weight (
gate.weight) with the right transposition forx @ router. - Run the router on an input batch and return only the union of top-K expert indices — no side effects on any module's expert-load counters or routing bias.
- Load a single routed expert's SwiGLU triplet
(
experts.{j}.{gate,up,down}_proj.weight) in fp16. - Optionally load a fused shared expert (DeepSeek-V2 style:
shared_experts.*) or a scalar-gated shared expert (Qwen2-MoE style:shared_expert.*+shared_expert_gate.weight). - Report
bytes_savedvs. loading the full routed set.
Handles both GateFunction.softmax (DeepSeek-V2, Qwen2-MoE) and
GateFunction.sigmoid (DeepSeek-V3), with optional top-K
renormalization.
Classes
- MoeRoutingDecision
- MoeStreamingLayer
- MoeStreamingReport
- Bytes-plan of a per-expert streaming forward — how many experts would be streamed and how many bytes vs. loading the full set.
- RoutedExpertWeights
-
Weights of one routed expert loaded from disk. Shapes follow the
HF layout:
wGate,wUpare[hidden, D];wDownis[D, hidden]. All fp16-preserved if the source is fp16.
Functions
-
swiGluForward(
Tensor x, RoutedExpertWeights w) → Tensor -
Utility: SwiGLU forward for a single expert given fp16-or-fp32
weights in the HF layout.
xis[T, D]; weights arewGate,wUp[hidden, D]andwDown[D, hidden]. All matmuls are done asx @ w.transpose(). -
swiGluForwardRaw(
Tensor x, Tensor gate, Tensor up, Tensor down) → Tensor - Same, but taking three raw tensors (useful when the shared expert weights come from a different HF prefix).