Problem 218984 · medium · Level 02 Linear Data Structures

Words as Vectors

py-comprehensions · py-enumerate · dictionaries · features

A learning algorithm cannot read text, but it can read lists of numbers. A simple way to turn documents into numbers:

  1. The vocabulary is the sorted list of all distinct words that appear in any document, after lower-casing, leaving out the words in stopwords.
  2. Each document becomes a vector with one entry per vocabulary word: entry j is how many times vocab[j] occurs in that document.

Words are separated by whitespace; compare them in lower case. stopwords is a set of lower-case words to ignore.

Write bag_of_words(docs, stopwords) that returns the pair (vocab, vectors), with one vector per document, in the order of docs.

Examples

Input:  docs = ["the cat sat", "The cat saw the dog"], stopwords = {"the"}
Output: (["cat", "dog", "sat", "saw"], [[1, 0, 1, 0], [1, 1, 0, 1]])

Input:  docs = ["Go go GO", "", "stop"], stopwords = set()
Output: (["go", "stop"], [[3, 0], [0, 0], [0, 1]])
Explanation: The empty document has no words, so its vector is all zeros.

Constraints

  • 0 <= len(docs) <= 2000, each document has at most 200 words
  • Words contain letters only.

Goals

  • Collect distinct items with a set comprehension and put them in order
  • Map each item to its position with a dictionary comprehension over `enumerate`
  • Turn every document into a fixed-length count vector
Starting Python…