Problem 465112 · hard · Level 04 Non-Linear Data Structures

Steps That Fit Together

py-composition · py-protocols · feature scaling · pipelines

Before a model sees data, the data usually goes through a few preparation steps: pick some columns, scale them, centre them. Every step has to learn its settings from the training rows only, and then apply exactly the same settings to any new rows. Doing this by hand for every experiment is error-prone, so you will build a pipeline.

Data is a list of rows, each a list of numbers. A transformer has fit(rows) (learn settings, return self) and transform(rows) (return new rows; never change the input). A model has fit(rows, labels) (return self) and predict(rows) (return a list of labels).

Write three classes:

  • MinMaxScaler(): fit learns each column's minimum and maximum; transform maps every value to (x - min) / (max - min) for its column, or 0.0 in a column whose training values were all equal.
  • SelectColumns(cols): transform keeps the columns with the indices in cols, in that order (negative indices count from the end, as in lists); fit learns nothing.
  • Pipeline(steps): a list of steps, where every step except the last is a transformer and the last is a transformer or a model.
    • fit(rows, labels=None) goes through the steps in order: a transformer is fitted on the current rows and then transforms them for the next step; if the last step has a predict method it is fitted with (rows, labels) instead. Returns self.
    • transform(rows) sends rows through every step with transform only (no fitting).
    • predict(rows) transforms with every step but the last and returns the last step's predict.

A pipeline is itself a transformer (or a model), so it can be a step of another pipeline. An empty pipeline transforms rows into an unchanged copy.

The setup provides steps written by someone else: Centre() (a transformer that subtracts each column's training mean), NearestCentroid() (a model), spy(step, log, tag) (wraps a step and appends (tag, method, number_of_rows) to the list log whenever fit, transform or predict is called), and the dataset flats(n, seed) returning (rows, labels) with columns of very different sizes.

Examples

Input:  p = Pipeline([MinMaxScaler()]).fit([[1, 100], [2, 300], [3, 200]])
        p.transform([[1, 100], [2, 300], [3, 200]]), p.transform([[5, 0]])
Output: ([[0.0, 0.0], [0.5, 1.0], [1.0, 0.5]], [[2.0, -0.5]])

Input:  Pipeline([SelectColumns([-1, 0]), NearestCentroid()]).fit(
            [[0, 7, 1000], [1, 7, 1000], [9, 7, 0], [10, 7, 0]], ["a", "a", "b", "b"]
        ).predict([[2, 7, 900], [8, 7, 50]])
Output: ['a', 'b']

Constraints

  • Up to 5000 rows and 10 columns; every row has the same number of columns.
  • Steps must be fitted in order, each on the output of the steps before it, and must not be fitted again by transform or predict.

Goals

  • Build one object out of a list of steps and pass data through them in order
  • Rely only on the methods a step has, so that steps from anywhere, including other pipelines, fit in
  • Learn preprocessing settings on the training data only and reuse them on new data
Starting Python…