A birdwatching club receives reports of a bird seen at the lake. Local records say how common each species is: priors is a dictionary from species to the probability that a bird seen there belongs to it (the values add up to 1).
Each report is a pair (species, accuracy) from one observer. In words: an observer with accuracy 0.8 names the right species 80% of the time; when they are wrong, they name each of the other species in priors equally often. Given the true species, different observers make their mistakes independently.
Write bird_posterior(priors, reports) that returns a tuple (chances, likeliest): chances is a dictionary from every species to its probability after all the reports, and likeliest is the species with the largest probability (on a tie within 1e-12, the one that comes first in priors).
Examples
Input: priors = {"heron": 0.6, "egret": 0.39, "hoopoe": 0.01}, reports = [("hoopoe", 0.9)]
Output: ({"heron": 0.5128205128205128, "egret": 0.3333333333333333,
"hoopoe": 0.15384615384615388}, "heron")
Explanation: a very good observer says "hoopoe". But hoopoes are rare: per 10 000 birds,
100 hoopoes give 90 correct reports, while 6000 herons give 6000 · 0.05 = 300 and 3900 egrets
195 wrong "hoopoe" reports. Only 90 of 585 "hoopoe" reports are right.
Input: the same priors, reports = [("hoopoe", 0.9), ("hoopoe", 0.8), ("hoopoe", 0.9)]
Output: hoopoe becomes the likeliest, with probability about 0.96
Constraints
- 2 to 30 species; every prior is positive
0 <= len(reports) <= 200; every accuracy is strictly between 0 and 1; every reported species is inpriors- floats are compared with a tolerance of
1e-6
Goals
- Translate a verbal description of reliability into likelihoods P(report | truth)
- Combine a prior over several possibilities with several independent reports
- See how a rare possibility needs strong evidence before it becomes the likeliest