@cthos @xgranade @HeavenlyPossum we used monte carlo bootstrapping to get our confidence intervals because we really had no way of concretely knowing what our actual distribution was. Anyway, we used to use a very straightforward method of obtaining the metric we were interested in and would calculate it directly from the data, then run monte carlo bootstrapping over it to get the average of the metric and error bars at the 95% CI. Standard stuff.
The problem arose when we began to implement another method of obtaining the metric we were interested in using a markov model. We built a rate matrix based on observations from the simulation and then get the eigenvectors/eigenvalues out, and that was our desired metric. We'd then use that metric as the starting point for the bootstrap...
You might be able to spot the error here, or maybe not, but a bootstrap uses a statistical estimator to pull from a distribution and the metric we were interested in was an average quantity. In the first scenario, the average was calculated as part of the bootstrapping procedure (which is correct). Unfortunately, in the second, building up a rate matrix is *itself* a type of averaging procedure, so when we naively ran the bootstrapping procedure on the eigenvalues from that our error bars were really, really, REALLY small.
In the first scenario, we were calculating the average and stddev/CI from a distribution of kinetic observables. In the second scenario, we were calculating the average and stddev/CI from a distribution of *averaged* kinetic observables, so instead of saying, "95 times out of a hundred that we run this simulation, our measurement of this observable is going to be between these two values", we were saying "95 times out of a hundred that we run this *averaging* sequence based on data from our observable, our measurement of this *average* is going to be between these two values" and those are not at all the same thing.
Anyway, I never took a stats class, but tracking down this error required me to learn an ass ton of stats on the fly and I had to spend a long time implementing and integrating a rate matrix estimator that would create a rate matrix from a random time sampling of our observables and then get the eigenvector/eigenvalue out of... that could be done 1,000 fucking times, at least, without taking up *forever*.
So I've been a bit of a real cranky asshole about misuse of stats ever since then and the shit they pull in for-profit machine learning fills my little "I earned $20,000 a year and it was the most money and food I ever had in my life and I worked 10 hours a day" grad student heart with rage.