As I was preparing for an R intro course I came up with the idea of creating a fake data set that is stuffed full of all the conceivable errors one can imagine. Just in case my imagination falls short, I’d appreciate all the suggestions in the comments so that I can incorporate more errors.
There is a Hungarian saying about the veterinarian’s horse to describe
a case that exhibits all the possible conditions a subject can suffer from
(read more of the etymology here).
I would like to create a data set that shows all the
possible errors a data set can exhibit. This data would be then used in
the aforementioned course to make participants’
life miserable experience more diverse.
So far I have been able to come up with the following issues:
"1,234,567.0058654"
(needs to clear commas, turn it into numeric, digits are irrelevant but eating up memory)"W-123"
vs. "w-123"
"W-123"
vs. "W-123 "
0-3
works fine, but 3-5
becomes 05-Mar
)I don’t imagine that this list can ever be complete, but right now it is far from complete. If you have struggled with a problem in the past and would like others to learn from it, please leave a comment and I will expand the list accordingly.
I moved to Canada in 2008 to start a postdoctoral fellowship with Prof. Subhash Lele at the stats department of the University of Alberta. Subhash at the time just published a paper about a statistical technique called data cloning. Data cloning is a way to use Bayesian MCMC algorithms to do frequentist inference. Yes, you read that right.
ABMI (7) ARU (1) Alberta (1) BAM (1) C (1) CRAN (1) Hungary (2) JOSM (2) MCMC (1) PVA (2) PVAClone (1) QPAD (3) R (20) R packages (1) abundance (1) bioacoustics (1) biodiversity (1) birds (2) course (2) data (1) data cloning (4) datacloning (1) dclone (3) density (1) dependencies (1) detect (3) detectability (3) footprint (3) forecasting (1) functions (3) intrval (4) lhreg (1) mefa4 (1) monitoring (2) pbapply (5) phylogeny (1) plyr (1) poster (2) processing time (2) progress bar (4) publications (2) report (1) sector effects (1) shiny (1) single visit (1) site (1) slider (1) slides (2) special (3) species (1) trend (1) tutorials (2) video (4) workshop (1)