A museum's visitor report is produced by this function, which receives the lines of an exported file (often straight from the file, so possibly a one-pass iterator):
def busiest_days(lines, k=3):
rows = (line.strip().split(",") for line in lines if line.strip())
total = sum(int(v) for _, v in rows)
ranked = sorted(rows, key=lambda r: r[1], reverse=True)
return total, [date for date, _ in ranked[:k]]
The intended behaviour: the first non-empty line is a header (date,visitors) and is skipped; every other non-empty line is YYYY-MM-DD,count. Return (total visitors, the dates of the k busiest days), busiest first, and among days with the same count the earlier date first. Blank lines (possibly with spaces) are ignored, and a file with only a header, or nothing at all, gives (0, []). Four reports came in:
- "With the real export it crashes:
ValueError: invalid literal for int() with base 10: 'visitors'." - "After working around that, the list of days is always empty, though the total is right."
- "A day with 900 visitors is ranked above a day with 1200."
- "Two days with the same count come out in file order; the earlier date must come first."
Write the corrected busiest_days(lines, k=3). The setup helper export(n, seed) generates a file's lines one at a time (a generator, header first), as the real file reader does.
Examples
Input: busiest_days(iter(["date,visitors", "2024-05-02,900", "", "2024-05-01,1200", "2024-05-03,900"]), k=2)
Output: (3000, ["2024-05-01", "2024-05-02"])
Input: busiest_days(["date,visitors", " "])
Output: (0, [])
Constraints
- Up to
10**5lines; counts are whole numbers from0to10**6; each date appears at most once. 0 <= k; if there are fewer thankdays, all are returned.
Goals
- Reproduce each bug report with a small input before changing anything
- Recognise an iterator that is used up by a first pass
- Fix sort keys that compare text instead of numbers, and tie-breaks that follow the wrong order