{"id":9819,"date":"2020-01-28T11:15:18","date_gmt":"2020-01-28T10:15:18","guid":{"rendered":"https:\/\/www.makingscience.com\/blog\/trees-and-forests-in-machine-learning\/"},"modified":"2020-01-28T11:15:18","modified_gmt":"2020-01-28T10:15:18","slug":"trees-and-forests-in-machine-learning","status":"publish","type":"post","link":"https:\/\/www.makingscience.com\/en\/blog\/trees-and-forests-in-machine-learning\/","title":{"rendered":"Trees and Forests in Machine Learning"},"content":{"rendered":"<div>Trees and Forests in Machine Learning<\/div>\n<div><span style=\"font-weight: 400;\">Tree based learning algorithms are among the most popular and most frequently used supervised machine learning methods. They are versatile, can be applied on a large variety of problems, they\u2019re interpretable, and easy to use.\u00a0<\/span>  <span style=\"font-weight: 400;\">First, let\u2019s see what those phrases in the previous sentence mean. <\/span><b>What are tree based learning algorithms?<\/b><span style=\"font-weight: 400;\"> The simplest of these is the decision tree, which exactly is as it sounds, as you can see in the image below. Starting from a root node, the algorithm makes decisions based on some predefined criteria about how to split the data, creating branches until it reaches a conclusion (leaf).<\/span>  <img fetchpriority=\"high\" decoding=\"async\" class=\"size-full wp-image-3254 alignleft\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2020-01-27-a-las-16.38.31.png\" alt=\"\" width=\"331\" height=\"336\" \/>  <i><span style=\"font-weight: 400;\">Source: By Stephen Milborrow &#8211; Own work, CC BY-SA 3.0, <\/span><\/i><a href=\"https:\/\/commons.wikimedia.org\/w\/index.php?curid=14143467\" target=\"_blank\" rel=\"noopener\"><i><span style=\"font-weight: 400;\">https:\/\/commons.wikimedia.org \/w\/index.php?curid=14143467<\/span><\/i><\/a>  &nbsp;  <i><span style=\"font-weight: 400;\">\u201cA tree showing survival of passengers on the<\/span><\/i><a href=\"https:\/\/en.wikipedia.org\/wiki\/Titanic\" target=\"_blank\" rel=\"noopener\"> <i><span style=\"font-weight: 400;\">Titanic<\/span><\/i><\/a><i><span style=\"font-weight: 400;\"> (&#8220;sibsp&#8221; is the number of spouses or siblings aboard). The figures under the leaves show the probability of survival and the percentage of observations in the leaf. Summarizing: Your chances of survival were good if you were (i) a female or (ii) a male younger than 9.5 years with less than 2.5 siblings.\u201d<\/span><\/i>  &nbsp;  <span style=\"font-weight: 400;\">Tree based algorithms are <\/span><b>supervised<\/b><span style=\"font-weight: 400;\"> methods, meaning that we need labeled data to train the model, to figure out what decisions to make and in which order. Only after training can we use the model to make predictions on unlabeled data.<\/span>  <span style=\"font-weight: 400;\">Tree based algorithms can be used on both <\/span><b>classification<\/b><span style=\"font-weight: 400;\"> and <\/span><b>regression<\/b><span style=\"font-weight: 400;\"> problems. Classification problems are those where the variable we wish to predict has only a small number of discrete outcomes. For example in the image above, there are two outcomes: the person survived or not. Regression problems predict a continuous variable, for example, how much a customer will spend in an online shop.<\/span>  <span style=\"font-weight: 400;\">An important aspect of tree based algorithms is that they are easily <\/span><b>interpretable<\/b><span style=\"font-weight: 400;\">. As you can see in the image, it\u2019s quite simple to understand how the algorithm predicts what it predicts, what are the decisions it makes. In machine learning we often use algorithms that are practically \u201cblack boxes\u201d: they seem to work well, give accurate predictions, but we cannot tell how those conclusions were reached.\u00a0<\/span>  <span style=\"font-weight: 400;\">In business settings \u201cblack box\u201d algorithms are often unacceptable no matter how good their scores are &#8211; which is understandable. If an algorithm refuses your mortgage application or predicts that you have a high chance to get a certain kind of disease, wouldn\u2019t you want to know why it reached that conclusion?<\/span>  <span style=\"font-weight: 400;\">Of course, most problems are more complicated than the one in the above example and a simple decision tree is not sufficient to solve them. But ensemble methods provide enough flexibility to deal with very complex problems. Ensemble methods are constructed of several individually trained models, which are then combined. There are several such algorithms, let\u2019s take a quick look at the most popular ones.<\/span>  &nbsp; <\/p>\n<h2><b>Random forest<\/b><\/h2>\n<p> <span style=\"font-weight: 400;\">Until a few years ago the Random Forest was considered as one of the most powerful machine learning algorithms. As you can guess from its name, a random forest model contains many decision trees. In order to train these trees, the algorithm samples the data randomly with replacement creating several subsamples and trains one tree on each of the subsamples. Then it combines the prediction of the trees, for example by using majority vote for a classification problem, and averaging the predictions for a regression problem.<\/span>  &nbsp; <\/p>\n<h2><b>Boosting<\/b><\/h2>\n<p> <span style=\"font-weight: 400;\">Boosting algorithms are newer and even more powerful tools than Random Forest. While Random Forest builds trees in parallel,\u00a0 boosting algorithms build the trees one after the other, taking into account the weakest points of the previous models, and creating a model that strengthens those points in the next step, converting weak learners into a strong one as a result.<\/span>  <span style=\"font-weight: 400;\">The two most popular methods both use gradient boosting. LightGBM\u00a0 was developed by Microsoft (<\/span><a href=\"https:\/\/github.com\/microsoft\/LightGBM\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">https:\/\/github.com\/microsoft\/LightGBM<\/span><\/a><span style=\"font-weight: 400;\">), while XGBoost (<\/span><a href=\"https:\/\/github.com\/dmlc\/xgboost\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">https:\/\/github.com\/dmlc\/xgboost<\/span><\/a><span style=\"font-weight: 400;\">)\u00a0 was developed as an open source project. Both algorithms are fast, implement parallel processing, and can be used for a great variety of problems.<\/span>  &nbsp; <\/p>\n<h2><b>Interpreting ensemble methods<\/b><\/h2>\n<p> <span style=\"font-weight: 400;\">As mentioned above, a huge advantage of tree-based methods is their interpretability, that it\u2019s possible to understand why the algorithm predicts what it predicts.\u00a0<\/span>  <span style=\"font-weight: 400;\">All the algorithms we discussed in this post\u00a0 provide tools to examine the importance of each of the input features. There are also several external tools that help to interpret your model, such as xgboostExplainer for the R language (<\/span><a href=\"https:\/\/github.com\/gameofdimension\/xgboost_explainer\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">https:\/\/github.com\/gameofdimension\/xgboost_explainer<\/span><\/a><span style=\"font-weight: 400;\">) and LIME (<\/span><a href=\"https:\/\/github.com\/marcotcr\/lime\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">https:\/\/github.com\/marcotcr\/lime<\/span><\/a><span style=\"font-weight: 400;\">) or SHAP (<\/span><a href=\"https:\/\/github.com\/slundberg\/shap\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">https:\/\/github.com\/slundberg\/shap<\/span><\/a><span style=\"font-weight: 400;\">) for Python. These tools can help you discover deeper connections between the features, and refine your\u00a0 models; you can even use them to create illustrations that help explain how your model works to non-experts.<\/span>  <span style=\"font-weight: 400;\">The image below shows a figure created with SHAP using a model for predicting the survival of passengers of the Titanic (<\/span><a href=\"https:\/\/meichenlu.com\/2018-11-10-SHAP-explainable-machine-learning\/\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">https:\/\/meichenlu.com\/2018-11-10-SHAP-explainable-machine-learning\/<\/span><\/a><span style=\"font-weight: 400;\">). It shows that females, first and second class passengers, and children had the best chance of survival.<\/span>  <img decoding=\"async\" class=\"alignnone size-full wp-image-3255\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2020-01-27-a-las-16.38.43.png\" alt=\"\" width=\"611\" height=\"331\" \/><\/div>\n","protected":false},"excerpt":{"rendered":"<p>Trees and Forests in Machine Learning Tree based learning algorithms are among the most popular and most frequently used supervised machine learning methods. They are versatile, can be applied on a large variety of problems, they\u2019re interpretable, and easy to use.\u00a0 First, let\u2019s see what those phrases in the previous sentence mean. What are tree [&hellip;]<\/p>\n","protected":false},"author":21,"featured_media":9822,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[913],"tags":[],"class_list":["post-9819","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technology-ai-en"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/posts\/9819","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/users\/21"}],"replies":[{"embeddable":true,"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/comments?post=9819"}],"version-history":[{"count":0,"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/posts\/9819\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/media\/9822"}],"wp:attachment":[{"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/media?parent=9819"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/categories?post=9819"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.makingscience.com\/en\/wp-json\/wp\/v2\/tags?post=9819"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}