A device sends a text log over a slow link. The text arrives as a stream of chunks of arbitrary sizes: a chunk may hold many lines, part of one line, or nothing at all, and a line ending can be split between two chunks.
Write a generator lines_from(chunks) that yields every line of the text, without its line ending, in order. A line ends with \n, with \r\n or with a lone \r; no other character ends a line (a form feed \x0c, for example, is ordinary text). Empty lines are yielded as "". If the text does not end with a line ending, the last piece is yielded when the stream ends (but an empty last piece is not).
The link is interactive, so each line must be yielded as soon as its line ending has arrived, without reading another chunk first. The stream may be endless.
Helpers available with Run: stream(items), first(gen, n), packets(seed) (an endless stream of chunks), and chunks_read(lines_from, chunks, k), which says how many chunks were pulled from the list chunks by the time the first k lines were taken.
Examples
Input: first(lines_from(stream(["ab\ncd", "\r\nef\r", "\ngh"])), 10)
Output: ["ab", "cd", "ef", "gh"]
Explanation: "\r" ends "ef" at once; the "\n" at the start of the next chunk belongs to it.
Input: first(lines_from(stream(["a\r", "\r\n", "b"])), 10)
Output: ["a", "", "b"]
Input: chunks_read(lines_from, ["one\r", "\ntwo\r", "\nthree"], 1)
Output: 1
Constraints
- The text may be endless; up to
2 * 10**5characters are read in a test, sometimes one character per chunk, and a single line can be very long. - Aim for time proportional to the amount of text: do not search the whole unfinished line again for every new chunk.
Goals
- Rebuild lines from text that arrives in pieces of arbitrary size
- Keep a small buffer between chunks and yield each line as soon as it is complete
- Handle `\r\n` split across two chunks without waiting for data you do not need