pub fn tokenize(text: &str) -> Vec<String>Expand description
Split text into normalized, indexable terms.
Lowercase, split on anything non-alphanumeric, fold regular plurals, and drop stop words. Deliberately not a stemmer: a personal corpus is full of names, and aggressive stemming conflates them (“Rhea” and “rhe”) for very little recall. Plural folding is the one exception, because “restaurants” and “restaurant” are the same word to every user who says either.
Paraphrases with no shared term at all (“diet” against a record indexed under “dietary”) are deliberately out of scope here — that is what record aliases and the semantic fallback exist for.