Treating Missing Data in Classification Trees

Treating Missing Data in Classification Trees

190 pages· 2008· ISBN 9780549947981
About
Classification trees are a type of supervised learning method used in classification problems where the response variable is categorical, with each category representing one target class. Classification tree algorithms build models (classification trees) on training data and apply the models to testing data. Ideally, all of the data points in the training data and all of the independent variables in the testing data are assumed to be observed. However, in reality, those data values can be missing (unobserved) and in fact, missing data is a fairly common problem. There are many different methods used by classification tree algorithms when missing data occur in the predictors, but few studies have been done comparing their appropriateness and performance. This research provides both analytic and Monte Carlo evidence regarding the effectiveness of six popular missing data methods for classification trees applied to binary response data. We make recommendations as to the best method to use in various situations when clear differences occur. We also show that in the context of classification trees (with extension to logistic regression), the relationship between the missingness and the dependent variable, rather than the standard missingness classification approach of Rubin (1976) and Little and Rubin (2002) (missing completely at random (MCAR), missing at random (MAR) and not missing at random (NMAR)), is the most helpful criterion to distinguish between different missing data methods.

Discuss Treating Missing Data in Classification Trees with other readers

Join or start a book club for Treating Missing Data in Classification Trees on Readfeed. Live chat, shared reading progress, and AI discussion questions — free to get started.

Frequently asked questions

How do I join a book club for Treating Missing Data in Classification Trees?

Sign up free on Readfeed, then browse public clubs or start your own club with Treating Missing Data in Classification Trees as the current read. Invite friends with a share link and discuss together with live chat and AI discussion questions.

Can I discuss Treating Missing Data in Classification Trees with other readers online?

Yes. Readfeed book clubs let you chat live, share progress, and join discussions about Treating Missing Data in Classification Trees with readers worldwide — whether your club is virtual, in-person, or hybrid.

Is Readfeed free?

Yes. Creating an account and joining book clubs is free. Sign up to find readers who love the same books and start discussing today.