A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection
We review accuracy estimation methods and compare the two most common methods: crossvalidation and bootstrap. Recent experimental results on arti cial data and theoretical results in restricted settings have shown that for selecting a good classi er from a set of classiers (model selection), ten-fold cross-validation may be better than the more expensive leaveone-out cross-validation. We report on a largescale experiment| over half a million runs of C4. 5 and a Naive-Bayes algorithm| to estimate the e ects of di erent parameters on these algorithms on real-world datasets. For crossvalidation, we vary the number of folds and whether the folds are strati ed or not; for bootstrap, we vary the number of bootstrap samples. Our results indicate that for real-word datasets similar to ours, the best method to use for model selection is ten-fold strati ed cross validation, even if computation power allows using more folds.
What this paper cites, inside the corpus
| Paper | Year | Cited |
|---|---|---|
| Classification and Regression Trees. | 1986 | 21,044 |
| Stacked generalization | 1992 | 7,777 |
What cites it, inside the corpus
Links
Topics
| Algorithms and Data Compression | Computer Science |
| Data Mining Algorithms and Applications | Computer Science |
| Machine Learning and Data Classification | Computer Science |
Is this record sound?
complete
Nothing in this record contradicts itself and no field we check is missing.
- supports1 author record(s) attached.
- supports14 reference(s) recorded.
- neutralThe DOI carries no year to check against.
- supportsA title is present.
Provenance
sha256 7e3d99a592f7f61f…