The help desk now wants every new ticket routed automatically. Train on labelled tickets and label new ones.
Write nb_classify(examples, messages), where examples is a list of (text, label) pairs and messages a list of texts, and return one label per message:
- Normalise every text (training and new): convert it to lower case, delete the characters
.,!?, and split it on whitespace into words. - The vocabulary is the set of all words of all training texts,
Vits size;N_cis the number of words in the training texts of labelcandcount_c(w)how oftenwoccurs among them. - The score of label
cfor a message islog(prior_c)plus, for every wordwof the message (with repeats) that is in the vocabulary,log((count_c(w) + 1) / (N_c + V)). Words not in the vocabulary are skipped.prior_cis the share of training examples labelledc. - The label with the highest score wins. Scores within
1e-9of the highest count as tied, and among tied labels the alphabetically first wins.
The setup provides support_tickets(n, seed) (clean text) and noisy_tickets(n, seed) (with capitals and punctuation), each returning n random (text, label) pairs with labels "billing", "technical" and "delivery".
Examples
Input: examples = [("Win cash now!", "spam"), ("free prize, win", "spam"), ("lunch at noon", "ham"),
("meeting at noon?", "ham"), ("free lunch", "ham")]
messages = ["WIN free lunch", "win cash prize", "hello"]
Output: ["ham", "spam", "ham"]
Explanation: V = 9, the spam texts have 6 words and the ham texts 8. For "WIN free lunch" spam
scores log(2/5) + log(3/15) + log(2/15) + log(1/15) = -7.249 and ham log(3/5) + log(1/17) +
log(2/17) + log(3/17) = -7.219, a narrow win for ham. "hello" is unknown, so only the priors count.
Constraints
1 <= len(examples) <= 5000,0 <= len(messages) <= 2000; every training text has at least one word- use
math.log; the tests have no near-ties between different labels other than messages with no known words and equal priors
Goals
- Train a Naive Bayes text classifier by counting words per class
- Score a message by adding log prior and log word probabilities
- Normalise text so that 'Refund!' and 'refund' are the same word
- Skip words the model has never seen and break ties by a stated rule