Problem 270109 · easy · Level 02 Linear Data Structures

The Bar Every Model Must Clear

majority baseline · accuracy · counting · training and test data

Before anyone trains a clever model, the team agrees on the bar it must clear: the "model" that ignores the features and always predicts the most common label of the training set. Its accuracy on the test set is the number to beat.

Write majority_bar(y_train, y_test) that returns a tuple (label, accuracy, need):

  • label is the most common label in y_train; if several labels are equally common, the one that appears first in y_train;
  • accuracy is the fraction of y_test equal to label (a float);
  • need is the smallest number of correct test predictions a model must make to be strictly more accurate than the baseline. It may be larger than len(y_test) when the baseline is already perfect: then no model can beat it on this test set.

Examples

Input:  y_train = ["ok", "late", "ok", "ok", "late"], y_test = ["late", "ok", "ok", "late"]
Output: ("ok", 0.5, 3)
Explanation: "ok" wins 3 to 2 in training. It is right on 2 of the 4 test cases,
so a model needs 3 correct test predictions to beat it.

Input:  y_train = [1, 0, 0, 1], y_test = [0, 0, 0]
Output: (1, 0.0, 1)
Explanation: 1 and 0 are tied; 1 appears first in the training labels.

Constraints

  • 1 <= len(y_train), len(y_test) <= 10**5
  • labels are strings or whole numbers

Goals

  • Learn the majority-class baseline from the training labels only
  • Measure the baseline's accuracy on the test labels
  • Turn the baseline into a concrete target for a real model
Starting Python…