server/fts5 library

Minimal but real FTS5 query language for the engine's MATCH operator.

Supported syntax (a subset of SQLite's FTS5 query grammar):

  • Bare token — token must appear in the document
  • "phrase here" — exact adjacent token sequence
  • prefix* — any token starting with prefix
  • a AND b / a b — implicit AND between terms
  • a OR b — disjunction
  • NOT a / - a — negation (binary form a NOT b also accepted)
  • (a OR b) c — parentheses for grouping

Tokenization splits on non-word characters and lowercases. Tokens are compared case-insensitively. Phrase matches require tokens to appear at consecutive positions in the document; prefix matches require any document token to start with the given prefix (case-insensitive).

This file is pure; it doesn't depend on any other engine type, so it can be unit-tested in isolation.

Classes

Fts5Index
Inverted index over a fixed corpus of text documents, supporting proper BM25 scoring (with IDF and length normalisation).
Fts5Node
Boolean AST node produced by parseFts5Query. Public so tests and the (future) inverted-index path can introspect compiled queries.

Functions

fts5Bm25(String document, String query, {double k1 = 1.2, double b = 0.75}) → double
BM25-style score for document vs query. With only a single document of context (no corpus), the inverse-document-frequency term collapses to a constant, so this is effectively the saturated term-frequency form of BM25:
fts5Match(String document, String query) → bool
Match query against the contents of a single document string.
fts5TermFrequency(String document, String query) → double
Per-document term-frequency score useful as a relevance signal in ORDER BY. The result is the sum over query terms of their raw occurrence counts in document (after the same tokenization used by fts5Match). Phrases count as 1 per non-overlapping occurrence; prefix terms count occurrences of every matching token. Documents that don't match the query at all score 0.
parseFts5Query(String query) → Fts5Node?
Parse query. Returns null when the query is empty or whitespace.
tokenizeFts(String text) → List<String>
Split text into lower-cased word tokens. The same tokenizer is used for both indexing and query terms.