Problem 525919 · easy · Level 05 Advanced Algorithms & Graphs

Does the Commute Depend on the District?

chi-square test · two-way table · independence · expected counts · p-value

A city surveyed how people get to work. table[i][j] is the number of people from district i who use transport mode j (for example car, bus, bicycle). If the mode did not depend on the district, each cell would be close to its expected count row total × column total / grand total. The chi-square statistic measures how far the table is from that:

X² = sum over all cells of (observed - expected)² / expected

Write chi_square_table(table) that returns a dict:

  • "x2": the statistic X²,
  • "df": the degrees of freedom (rows - 1) · (columns - 1),
  • "p": the p-value P(χ² >= X²) when it has a closed form: math.erfc(math.sqrt(X² / 2)) for 1 degree of freedom, math.exp(-X² / 2) for 2 degrees of freedom, and None otherwise,
  • "smallest_expected": the smallest expected count (the usual rule of thumb trusts the chi-square p-value only when it is at least 5).

Examples

Input:  table = [[48, 22], [31, 39]]
Output: {"x2": 8.395932766134052, "df": 1, "p": 0.003760614893667963, "smallest_expected": 30.5}
Explanation: the row totals are 70 and 70, the column totals 79 and 61, so every cell of
the first column expects 39.5 and of the second 30.5. The four squared gaps, each
divided by its expected count, add up to 8.40.

Input:  table = [[34, 21, 15], [22, 30, 28]]
Output: {"x2": 7.456369175577208, "df": 2, "p": 0.024036432284661666, "smallest_expected": 20.066666666666666}

Constraints

  • the table has at least 2 rows and 2 columns, at most 20 of each
  • the counts are whole numbers >= 0, and every row and column total is positive
  • floats are compared with a tolerance of 1e-6

Goals

  • Compute the expected counts of a two-way table under independence
  • Add up the chi-square statistic and find its degrees of freedom
  • Use the closed-form tail probabilities that exist for 1 and 2 degrees of freedom
Starting Python…