A second-hand shop wants to predict which items sell within a week. Its rows mix numbers and words, for example ["red", 12, "M"] for colour, price and size. Distances and thresholds need numbers, and replacing words by arbitrary codes (red = 1, blue = 2, green = 3) would pretend that green is "further" from red than blue is. The fix is to give every category its own 0/1 column.
Write encode(train_rows, test_rows, categorical) that returns a tuple (train_encoded, test_encoded):
categoricallists the column positions that hold categories (strings); every other column holds a number and is copied unchanged;- for each categorical column, its categories are the distinct values of that column in the training rows, in sorted order;
- a categorical value becomes a block of 0s and 1s, one entry per category, with a 1 at the position of its own category; a value that never occurs in the training rows becomes a block of all 0s;
- an encoded row is the concatenation of the column results, in column order.
Examples
Input: train_rows = [["red", 3, "S"], ["blue", 5, "M"], ["red", 1, "M"]]
test_rows = [["green", 2, "S"]]
categorical = [0, 2]
Output: ([[0, 1, 3, 0, 1], [1, 0, 5, 1, 0], [0, 1, 1, 1, 0]], [[0, 0, 2, 0, 1]])
Explanation: the colours seen in training are ["blue", "red"] and the sizes ["M", "S"].
"green" never appeared in training, so its block is [0, 0].
Input: train_rows = [[7, "x"], [8, "x"]], test_rows = [], categorical = [1]
Output: ([[7, 1], [8, 1]], [])
Constraints
1 <= len(train_rows) <= 5000,0 <= len(test_rows) <= 5000; all rows have the same length (1 to 10)categoricalis a list of distinct valid positions in increasing order (it may be empty)- the input rows must not be changed
Goals
- Turn a category (a word) into numbers a distance or a rule can use
- Learn the list of categories from the training data only
- Encode test rows with the training vocabulary, including categories never seen in training