To me this seems less understandable than the reproducibility crisis in psychology research, for at least in psychology research there is always the possibility that different samples can belegitimately different, and thus legitimately provide different results, but in machine learning, there are standardized training and testing samples, and the methods can be entirely encapsulated into immutable computer code. In this respect, the lack of reproducibility in machine learning papers suggests to me either distorting incentives or rampant unreported tweaking of parameters and/or falsification, or probably, I suspect, intentional obsfuscation of methods to retain value as IP in the private market. I'm not sure which or which mixture, but publishing code, and datasets if funded with public money should be made more manditory.
Interestingly --- and in my impression unlike psychology research --- I don't think this paper ever attempted to replicate and got contradictory results. Instead, I think the complaint is that for the majority of papers they were unable to begin the replication for lack of working code or available data sets. When they had these, they appear to have obtained the same results.