Before a model sees data, the data usually goes through a few preparation steps: pick some columns, scale them, centre them. Every step has to learn its settings from the training rows only, and then apply exactly the same settings to any new rows. Doing this by hand for every experiment is error-prone, so you will build a pipeline.
Data is a list of rows, each a list of numbers. A transformer has fit(rows) (learn settings, return self) and transform(rows) (return new rows; never change the input). A model has fit(rows, labels) (return self) and predict(rows) (return a list of labels).
Write three classes:
MinMaxScaler():fitlearns each column's minimum and maximum;transformmaps every value to(x - min) / (max - min)for its column, or0.0in a column whose training values were all equal.SelectColumns(cols):transformkeeps the columns with the indices incols, in that order (negative indices count from the end, as in lists);fitlearns nothing.Pipeline(steps): a list of steps, where every step except the last is a transformer and the last is a transformer or a model.fit(rows, labels=None)goes through the steps in order: a transformer is fitted on the current rows and then transforms them for the next step; if the last step has apredictmethod it is fitted with(rows, labels)instead. Returnsself.transform(rows)sends rows through every step withtransformonly (no fitting).predict(rows)transforms with every step but the last and returns the last step'spredict.
A pipeline is itself a transformer (or a model), so it can be a step of another pipeline. An empty pipeline transforms rows into an unchanged copy.
The setup provides steps written by someone else: Centre() (a transformer that subtracts each column's training mean), NearestCentroid() (a model), spy(step, log, tag) (wraps a step and appends (tag, method, number_of_rows) to the list log whenever fit, transform or predict is called), and the dataset flats(n, seed) returning (rows, labels) with columns of very different sizes.
Examples
Input: p = Pipeline([MinMaxScaler()]).fit([[1, 100], [2, 300], [3, 200]])
p.transform([[1, 100], [2, 300], [3, 200]]), p.transform([[5, 0]])
Output: ([[0.0, 0.0], [0.5, 1.0], [1.0, 0.5]], [[2.0, -0.5]])
Input: Pipeline([SelectColumns([-1, 0]), NearestCentroid()]).fit(
[[0, 7, 1000], [1, 7, 1000], [9, 7, 0], [10, 7, 0]], ["a", "a", "b", "b"]
).predict([[2, 7, 900], [8, 7, 50]])
Output: ['a', 'b']
Constraints
- Up to
5000rows and10columns; every row has the same number of columns. - Steps must be fitted in order, each on the output of the steps before it, and must not be fitted again by
transformorpredict.
Goals
- Build one object out of a list of steps and pass data through them in order
- Rely only on the methods a step has, so that steps from anywhere, including other pipelines, fit in
- Learn preprocessing settings on the training data only and reuse them on new data