core/data/t5_sp_tokenizer library
Loads a HuggingFace T5-style tokenizer.json (SentencePiece
Unigram model) and provides encode/decode.
T5 uses SentencePiece with a Unigram language model, not BPE. The pipeline is:
- Normalization. Text is NFKC-normalized and whitespace
is compressed / stripped per T5's
precompiled_charsmap. We approximate this with a simple NFKC pass plus space compression. - Metaspace pre-tokenizer. A leading space is prepended
to the text (T5 always treats the first token as
word-initial), then every space is replaced with
▁(U+2581, "lower one-eighth block"), the SentencePiece whitespace marker. - Unigram Viterbi. Dynamic-programming search for the single-best segmentation into pieces from the vocabulary, maximizing the sum of log-scores. Uses a piece trie to keep the inner loop O(text_length * max_piece_length).
- Vocab lookup → int ids.
Decode reverses: vocab → piece strings → concatenate → replace
▁ with space → strip leading space.
Supports the standard T5 special tokens: <pad> (id 0), </s>
(id 1), <unk> (id 2), and the <extra_id_0> .. <extra_id_99>
sentinel tokens used by T5's span-corruption objective.