A learning algorithm cannot read text, but it can read lists of numbers. A simple way to turn documents into numbers:
- The vocabulary is the sorted list of all distinct words that appear in any document, after lower-casing, leaving out the words in
stopwords. - Each document becomes a vector with one entry per vocabulary word: entry
jis how many timesvocab[j]occurs in that document.
Words are separated by whitespace; compare them in lower case. stopwords is a set of lower-case words to ignore.
Write bag_of_words(docs, stopwords) that returns the pair (vocab, vectors), with one vector per document, in the order of docs.
Examples
Input: docs = ["the cat sat", "The cat saw the dog"], stopwords = {"the"}
Output: (["cat", "dog", "sat", "saw"], [[1, 0, 1, 0], [1, 1, 0, 1]])
Input: docs = ["Go go GO", "", "stop"], stopwords = set()
Output: (["go", "stop"], [[3, 0], [0, 0], [0, 1]])
Explanation: The empty document has no words, so its vector is all zeros.
Constraints
0 <= len(docs) <= 2000, each document has at most200words- Words contain letters only.
Goals
- Collect distinct items with a set comprehension and put them in order
- Map each item to its position with a dictionary comprehension over `enumerate`
- Turn every document into a fixed-length count vector