A help desk sorts incoming tickets by topic. Before building a full classifier it wants the two ingredients every word-based probabilistic model needs.
examples is a list of (text, label) pairs; the words of a text are separated by single spaces. The vocabulary is the set of all distinct words in all texts, and V is its size. For a label c, let N_c be the total number of words (with repeats) in the texts labelled c, and count_c(w) the number of times the word w occurs in them.
Write word_model(examples, words) that returns a tuple (priors, likelihoods):
priors[c]is the share of the examples labelledc;likelihoods[c]is a list with one number per wordwofwords, in order:(count_c(w) + 1) / (N_c + V). The same formula is used for a word that is not in the vocabulary.
Both are dictionaries with one entry per label that occurs in examples. The setup provides support_tickets(n, seed), which returns n random tickets with labels "billing", "technical" and "delivery".
Examples
Input: examples = [("reset my password", "tech"), ("password not working", "tech"),
("refund my payment", "billing")]
words = ["password", "refund", "cat"]
Output: ({"tech": 0.6666666666666666, "billing": 0.3333333333333333},
{"tech": [0.23076923076923078, 0.07692307692307693, 0.07692307692307693],
"billing": [0.1, 0.2, 0.1]})
Explanation: the vocabulary has V = 7 words. The tech tickets have N = 6 words, 2 of them
"password": (2 + 1) / (6 + 7) = 3/13. The billing ticket has 3 words, 1 of them "refund":
(1 + 1) / (3 + 7) = 0.2. A word never seen, like "cat", gets 1 / (N + V).
Constraints
1 <= len(examples) <= 5000,0 <= len(words) <= 1000; every text has at least one word- floats are compared with a tolerance of
1e-6
Goals
- Count words per class in labelled text
- Estimate P(word | class) with add-one smoothing over the whole vocabulary
- Estimate the class priors from the share of examples