core/data/t5_sp_tokenizer library

Loads a HuggingFace T5-style tokenizer.json (SentencePiece Unigram model) and provides encode/decode.

T5 uses SentencePiece with a Unigram language model, not BPE. The pipeline is:

  1. Normalization. Text is NFKC-normalized and whitespace is compressed / stripped per T5's precompiled_charsmap. We approximate this with a simple NFKC pass plus space compression.
  2. Metaspace pre-tokenizer. A leading space is prepended to the text (T5 always treats the first token as word-initial), then every space is replaced with ▁ (U+2581, "lower one-eighth block"), the SentencePiece whitespace marker.
  3. Unigram Viterbi. Dynamic-programming search for the single-best segmentation into pieces from the vocabulary, maximizing the sum of log-scores. Uses a piece trie to keep the inner loop O(text_length * max_piece_length).
  4. Vocab lookup → int ids.

Decode reverses: vocab → piece strings → concatenate → replace ▁ with space → strip leading space.

Supports the standard T5 special tokens: <pad> (id 0), </s> (id 1), <unk> (id 2), and the <extra_id_0> .. <extra_id_99> sentinel tokens used by T5's span-corruption objective.

Classes

T5SpTokenizer