A garden centre asked its customers to fill in a survey. Every answer arrived as a dictionary (a record), for example {"age": 34, "visits": 5, "garden_m2": 120, "bought_tools": "yes"}. Before any model can learn from them, the records must become a feature table X (a list of rows) and a label list y.
Write to_xy(records, features, label) that returns the tuple (X, y):
X[i]is the list of the values of the keys infeatures, in the order given byfeatures;y[i]is the value oflabelfor the same record;- a record is skipped if any of those keys (a feature or the label) is missing from it or has the value
None. Other keys in a record are ignored.
The records keep their original order, and records must not be changed.
Examples
Input: records = [{"age": 34, "visits": 5, "bought": "yes"},
{"age": 51, "visits": None, "bought": "no"},
{"visits": 2, "age": 19, "bought": "no", "city": "Leeds"}]
features = ["visits", "age"], label = "bought"
Output: ([[5, 34], [2, 19]], ["yes", "no"])
Explanation: the second record has no value for "visits", so it is skipped;
the columns follow the order of `features`, not of the record.
Input: records = [{"a": 1}, {"a": 2, "y": 0}], features = ["a"], label = "y"
Output: ([[2]], [0])
Constraints
0 <= len(records) <= 10**4,1 <= len(features) <= 20labelis not one offeatures- values are numbers or strings;
0,""andFalseare real values, only a missing key orNonemakes a record incomplete
Goals
- Turn a list of records into a feature table X and a label list y
- Keep X and y in step: row i of X always belongs to label i of y
- Drop incomplete examples instead of guessing values