It matters which number you are asking about, because the site is much better at one than the other. Keeping these apart is the whole point of this page — a single headline accuracy figure here would be misleading.
Every error on this page is our error. The reference is treated as correct and we are scored against it. When this page quotes a number, it does not mean the published figure is that far off — it means our estimate sits that many percentage points away from theirs.
"Percentage points" is the gap between two percentages. If we say a river took 43% of the run and the published figure says 30%, we are 13 points off — not 13% off.
There are two references, they answer different questions, and their scores are not interchangeable:
How good is the ruler itself? The IDFG hatchery values were read off printed bar charts, accurate to roughly ±500 fish. We tested how much of the error that could explain: about 3%. The ruler is blunt, but the gap is ours.
The humbling result first: simply using last year's published composition is the most accurate option we have. River-to-river composition changes slowly, so a rule that uses none of our tag data beats our whole pipeline. Our uncorrected PIT estimate is the worst thing on the chart.
That is not an argument for deleting the model. Published escapement figures arrive years late, so for the current season there is nothing to copy — and that is exactly the job the model exists to do. The proposed design anchors on the published composition and uses our tags only for the change since then.
This is the chart that made the fix possible. We undercount the Grande Ronde in every single year, and overcount the Imnaha in every single year. An error with a consistent direction is a bias — something is systematically wrong and can be measured and removed. An error that flips sign is noise, and nothing can be done with it.
The Imnaha is a short river with antennas nearly everywhere, so its fish are easy to catch and it looks bigger than it is. The Grande Ronde is large and thinly instrumented, so its fish slip through and it looks smaller. Correcting for how well each basin is watched is not a fudge factor — it is a claim about the antenna network that can be checked.
And it checks out. The diamonds are measured antenna efficiency, calculated from completely separate data. The Grande Ronde comes last on both measures. Two independent methods agreeing is the strongest evidence on this page.
Not all error is the same. An error that leans the same way every year is a bias — something is systematically wrong, and once you find out what, you can measure it and take it out. An error that swings either side of zero and averages to nearly nothing is noise, and no correction removes noise; subtracting a constant from it just makes things worse.
That distinction is what made the August fix possible. The river-to-river error leaned one way in every single year, which said a mechanism was responsible rather than bad luck. Two turned out to be: some hatcheries tag a far larger share of their fish than others, and some rivers are watched by far more antennas than others. Both make the same rivers look bigger than they are. Dividing both back out moved this site's biggest chart from worse than a fifteen-year average to better than anything else we could do.