core/nn/moe_streaming library

AirLLM-style per-expert streaming for a single MoE FFN layer, factored out of the demo bins so the same code powers both the synthetic-weights runner (bin/moe_streaming_demo.dart) and the real-weights runners (bin/deepseek_v2_moe_layer_stream.dart, bin/qwen15_moe_layer_stream.dart).

Given a ShardedSafeTensorsReader and the HF key prefix for the layer's mlp submodule (e.g. "model.layers.5.mlp"), it can:

  • Load the router weight (gate.weight) with the right transposition for x @ router.
  • Run the router on an input batch and return only the union of top-K expert indices — no side effects on any module's expert-load counters or routing bias.
  • Load a single routed expert's SwiGLU triplet (experts.{j}.{gate,up,down}_proj.weight) in fp16.
  • Optionally load a fused shared expert (DeepSeek-V2 style: shared_experts.*) or a scalar-gated shared expert (Qwen2-MoE style: shared_expert.* + shared_expert_gate.weight).
  • Report bytes_saved vs. loading the full routed set.

Handles both GateFunction.softmax (DeepSeek-V2, Qwen2-MoE) and GateFunction.sigmoid (DeepSeek-V3), with optional top-K renormalization.

Classes

MoeRoutingDecision
MoeStreamingLayer
MoeStreamingReport
Bytes-plan of a per-expert streaming forward — how many experts would be streamed and how many bytes vs. loading the full set.
RoutedExpertWeights
Weights of one routed expert loaded from disk. Shapes follow the HF layout: wGate, wUp are [hidden, D]; wDown is [D, hidden]. All fp16-preserved if the source is fp16.

Functions

swiGluForward(Tensor x, RoutedExpertWeights w) → Tensor
Utility: SwiGLU forward for a single expert given fp16-or-fp32 weights in the HF layout. x is [T, D]; weights are wGate, wUp [hidden, D] and wDown [D, hidden]. All matmuls are done as x @ w.transpose().
swiGluForwardRaw(Tensor x, Tensor gate, Tensor up, Tensor down) → Tensor
Same, but taking three raw tensors (useful when the shared expert weights come from a different HF prefix).