{"id":9662,"date":"2021-02-16T10:45:46","date_gmt":"2021-02-16T09:45:46","guid":{"rendered":"https:\/\/www.makingscience.com\/blog\/clustering-mas-alla-de-k-means\/"},"modified":"2021-02-16T10:45:46","modified_gmt":"2021-02-16T09:45:46","slug":"clustering-mas-alla-de-k-means","status":"publish","type":"post","link":"https:\/\/www.makingscience.com\/es\/blog\/clustering-mas-alla-de-k-means\/","title":{"rendered":"Clustering: m\u00e1s all\u00e1 de k-means"},"content":{"rendered":"<div>Clustering: m\u00e1s all\u00e1 de k-means<\/div>\n<div><span style=\"font-weight: 400;\">Clustering hace referencia a un conjunto de t\u00e9cnicas y algoritmos no supervisados que tienen como objetivo segmentar datos seg\u00fan grupos homog\u00e9neos, tambi\u00e9n llamados clusters. Dado un conjunto de puntos en un dataset, podemos aplicar estos algoritmos para clasificar cada punto en un grupo espec\u00edfico. Idealmente los puntos de un mismo cluster tienen caracter\u00edsticas similares entre s\u00ed y, al mismo tiempo, los puntos pertenecientes a clusters distintos tendr\u00e1n caracter\u00edsticas distintas. Esto, desde un punto de vista de negocio, permite sacar ventaja en una multitud de situaciones. Por ejemplo, clusterizar a tus clientes seg\u00fan sus compras y variables demogr\u00e1ficas para poder personalizar comunicaciones en base a las caracter\u00edsticas de cada cluster, o bien, ver qu\u00e9 grupos existen seg\u00fan el comportamiento que tienen los usuarios en tu web o app.<\/span>  <span style=\"font-weight: 400;\">Es muy probable que lo primero que se nos venga a la cabeza cuando queremos crear clusters con nuestros datos sea la t\u00e9cnica <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means. Sin embargo, esta no es siempre la mejor opci\u00f3n. Un buen an\u00e1lisis previo junto con un buen entendimiento de los datos son clave para tomar la decisi\u00f3n acertada.<\/span>  <span style=\"font-weight: 400;\">El algoritmo de clustering m\u00e1s conocido es <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means. Este algoritmo goza de una gran y bien merecida popularidad gracias a su sencillez y amplio soporte. Pero, como es razonable, es imposible que una sola t\u00e9cnica sea capaz de proporcionar una buena soluci\u00f3n para todo tipo de problemas. En este art\u00edculo vamos a introducir este y otros algoritmos.<\/span>  &nbsp; <\/p>\n<h2><strong>Cuando <i>k<\/i>-means es la soluci\u00f3n<\/strong><\/h2>\n<p> <span style=\"font-weight: 400;\">Supongamos que tenemos el siguiente dataset con 2 dimensiones. Por ejemplo, imaginemos que estas variables son el peso y la altura normalizadas (de 0 a 1) de un conjunto de personas:<\/span>  <img fetchpriority=\"high\" decoding=\"async\" class=\"alignnone  wp-image-4124\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.24.24.png\" alt=\"\" width=\"506\" height=\"397\" \/>  <span style=\"font-weight: 400;\">A simple vista se pueden diferenciar 4 grupos claramente. Esta informaci\u00f3n servir\u00e1 de entrada para <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means: <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">=4. Una vez el algoritmo recibe el n\u00famero de grupos deseado los pasos que sigue son sencillos:<\/span> <\/p>\n<ol>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Se inicializan <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\"> centroides aleatoriamente, en este caso <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">=4 centroides.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Se clasifica cada punto del dataset calculando la distancia de cada punto a cada centroide. Cada punto pertenece al grupo cuyo centroide se encuentra m\u00e1s cercano.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Con los grupos creados en el paso 2, se recalculan los centroides como la media de todos los puntos del grupo. De ah\u00ed el nombre <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Se repiten los pasos 2 y 3 hasta que los centroides ya no cambian demasiado por cada iteraci\u00f3n.<\/span><\/li>\n<\/ol>\n<p> <span style=\"font-weight: 400;\">Para m\u00e1s detalle as\u00ed como una implementaci\u00f3n simple en python, puedes consultar <\/span><a href=\"https:\/\/www.kaggle.com\/shrutimechlearn\/step-by-step-kmeans-explained-in-detail\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">este art\u00edculo<\/span><\/a><span style=\"font-weight: 400;\">.\u00a0\u00a0<\/span>  <span style=\"font-weight: 400;\">Como vemos en la siguiente figura, <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means consigue segmentar el dataset tal y como lo esper\u00e1bamos:<\/span>  <img decoding=\"async\" class=\"alignnone  wp-image-4125\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.24.30.png\" alt=\"\" width=\"472\" height=\"297\" \/>  &nbsp; <\/p>\n<h2><strong>\u00bfEs siempre <i>k<\/i>-means la opci\u00f3n m\u00e1s recomendable?<\/strong><\/h2>\n<p> <span style=\"font-weight: 400;\">La respuesta es no. Hay tipos de datos en los que <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means, por su propia naturaleza, no funciona de la forma deseada. Como hemos visto, para que <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means tenga un funcionamiento adecuado los centroides, <\/span><i><span style=\"font-weight: 400;\">i.e.<\/span><\/i><span style=\"font-weight: 400;\"> las medias de cada cluster, tienen que estar suficientemente alejados. Adem\u00e1s, en el punto 3 del algoritmo <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means se recalculan los centroides calculando la media de cada grupo. Esto \u201cfuerza\u201d que los clusters tiendan a ser circulares por el hecho de que el centroide est\u00e1 justo en el centro del cluster. Pero, \u00bfqu\u00e9 pasar\u00eda si nuestros datos no son circulares, o si los centroides no est\u00e1n suficientemente alejados?<\/span>  &nbsp; <\/p>\n<h3><strong>Gaussian Mixture Model (GMM)<\/strong><\/h3>\n<p> <span style=\"font-weight: 400;\">Si los grupos que vemos a simple vista no son \u201ccirculares\u201d, puede que <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means no sea la opci\u00f3n a utilizar. En el siguiente dataset vemos 3 grupos claramente diferenciables a simple vista. Estos tienen forma el\u00edptica, no circular:<\/span>  <img decoding=\"async\" class=\"alignnone  wp-image-4126\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.24.34.png\" alt=\"\" width=\"443\" height=\"348\" \/>  <span style=\"font-weight: 400;\">Para este tipo de datos, una buena opci\u00f3n podr\u00eda ser Gaussian Mixture Models (GMM). Con este algoritmo asumimos que los puntos siguen una distribuci\u00f3n Gaussiana. Esta asunci\u00f3n es menos restrictiva que suponer que son circulares como en el caso de <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means. De hecho, <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means es un caso especial de GMM donde la covarianza de cada cluster en todas las dimensiones tiende a 0. Con GMM, los clusters tienen 2 par\u00e1metros a optimizar: la media y la desviaci\u00f3n t\u00edpica. Con estos par\u00e1metros, en este ejemplo con 2 dimensiones, los clusters pueden adoptar cualquier forma el\u00edptica. En <\/span><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2019\/10\/gaussian-mixture-models-clustering\/\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">este enlace<\/span><\/a><span style=\"font-weight: 400;\"> tienes m\u00e1s detalle, as\u00ed como una implementaci\u00f3n simple en python.<\/span>  <img loading=\"lazy\" decoding=\"async\" class=\"alignnone  wp-image-4127\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.24.46.png\" alt=\"\" width=\"489\" height=\"160\" \/>  <span style=\"font-weight: 400;\">Como podemos ver en la figura de arriba, <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means (izquierda) en este caso no ha sido capaz de diferenciar bien los 3 grupos. En cambio, GMM (derecha) s\u00ed ha sido capaz de encontrarlos, gracias al par\u00e1metro de la desviaci\u00f3n t\u00edpica.<\/span>  &nbsp; <\/p>\n<h2><strong>Density-Based Spatial Clustering of Applications with Noise (DBSCAN)<\/strong><\/h2>\n<p> <span style=\"font-weight: 400;\">En un ejercicio de imaginaci\u00f3n, supongamos ahora que nuestros datos con las variables altura y peso tienen esta pinta:<\/span>  <img loading=\"lazy\" decoding=\"async\" class=\"alignnone  wp-image-4128\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.24.52.png\" alt=\"\" width=\"471\" height=\"354\" \/>  <span style=\"font-weight: 400;\">A simple vista es f\u00e1cil distinguir 3 grupos, que se corresponden con los 3 anillos conc\u00e9ntricos que se observan. En este caso, ni <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means ni GMM van ser capaces de separar los 3 grupos dado que la media ideal de cada grupo se encontrar\u00eda en el mismo punto: en el centro.<\/span>  <span style=\"font-weight: 400;\">En la siguiente figura, vemos precisamente c\u00f3mo los clusters que devuelven <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means (izquierda) y GMM (derecha) no son v\u00e1lidos para estos datos:<\/span>  <img loading=\"lazy\" decoding=\"async\" class=\"alignnone  wp-image-4129\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.24.59.png\" alt=\"\" width=\"618\" height=\"183\" \/>  <span style=\"font-weight: 400;\">En este caso, utilizar el algoritmo Density-Based Spatial Clustering of Applications with Noise (DBSCAN) puede ser la opci\u00f3n correcta. Una caracter\u00edstica de este algoritmo es que no necesita recibir el n\u00famero de clusters como par\u00e1metro, sino que el propio algoritmo encuentra el n\u00famero \u00f3ptimo de clusters. El funcionamiento es sencillo:<\/span> <\/p>\n<ol>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Se recibe como par\u00e1metro  , que es la distancia con la que vamos a trabajar y <\/span><i><span style=\"font-weight: 400;\">min_points.<\/span><\/i><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Se empieza con un punto cualquiera del dataset y se miran los puntos que est\u00e1n a distancia   de este punto. Si el n\u00famero de puntos encontrados es mayor que <\/span><i><span style=\"font-weight: 400;\">min_points, <\/span><\/i><span style=\"font-weight: 400;\">entonces el algoritmo empieza y consideramos al punto actual como el primer punto del nuevo cluster, si no, este punto se considera <\/span><i><span style=\"font-weight: 400;\">ruido<\/span><\/i><span style=\"font-weight: 400;\">.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Todos los puntos que est\u00e1n a distancia   del punto actual se unir\u00e1n al cluster. Este proceso se repite para todos los puntos nuevos que se han a\u00f1adido al cluster.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Los pasos 2 y 3 se repiten hasta que todos los puntos del cluster han sido etiquetados.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Cuando terminamos con el primer cluster, se elige aleatoriamente otro punto no visitado anteriormente y se repite el proceso hasta que todo el dataset ha sido etiquetado, ya sea como perteneciente a un cluster o como ruido.<\/span><\/li>\n<\/ol>\n<p> <span style=\"font-weight: 400;\">Para m\u00e1s detalle con este algoritmo puedes consultar el <\/span><a href=\"https:\/\/towardsdatascience.com\/dbscan-clustering-explained-97556a2ad556\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">siguiente art\u00edculo<\/span><\/a><span style=\"font-weight: 400;\"> donde tambi\u00e9n se incluye una implementaci\u00f3n sencilla en python.<\/span>  <span style=\"font-weight: 400;\">En la figura de abajo vemos c\u00f3mo DBSCAN ha sido capaz de segmentar correctamente todos los puntos del dataset.<\/span>  <img loading=\"lazy\" decoding=\"async\" class=\"alignnone  wp-image-4130\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.25.05.png\" alt=\"\" width=\"454\" height=\"278\" \/>  &nbsp;  &nbsp; <\/p>\n<h2><strong>Cuando los datos son <i>sparse<\/i><\/strong><\/h2>\n<p> <span style=\"font-weight: 400;\">Finalmente, vamos a ver un tipo de datos diferente a los anteriores. Supongamos que tenemos un e-commerce donde vendemos productos de 30 categor\u00edas diferentes: moda, electr\u00f3nica, libros, hogar, etc. Tenemos datos de nuestros usuarios (20k usuarios) que nos indican el n\u00famero de compras ponderados\u00a0 que han realizado de cada categor\u00eda. Aqu\u00ed podemos ver un extracto de estos datos:<\/span>  <img loading=\"lazy\" decoding=\"async\" class=\"alignnone  wp-image-4131\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.25.11.png\" alt=\"\" width=\"725\" height=\"261\" \/>  <span style=\"font-weight: 400;\">Como vemos, los datos son <\/span><i><span style=\"font-weight: 400;\">sparse<\/span><\/i><span style=\"font-weight: 400;\">, es decir, hay muchos ceros. Esto es debido a que, por lo general, los usuarios realizan compras de unas pocas categor\u00edas, pero no de todas. El objetivo de aplicar clustering en estos datos es conseguir obtener grupos de usuarios en base a sus compras. Idealmente, obtendremos grupos de usuarios que compran varias categor\u00edas de productos y por lo tanto tienen un comportamiento similar en nuestra web.<\/span>  <span style=\"font-weight: 400;\">Si utilizamos <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means para clusterizar a los usuarios ocurre lo siguiente:<\/span>  <img loading=\"lazy\" decoding=\"async\" class=\"alignnone  wp-image-4132\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.25.17.png\" alt=\"\" width=\"692\" height=\"146\" \/>  <img loading=\"lazy\" decoding=\"async\" class=\"alignnone  wp-image-4133\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.25.36.png\" alt=\"\" width=\"713\" height=\"503\" \/>  <span style=\"font-weight: 400;\">En la figura de arriba vemos por un lado el tama\u00f1o de cada cluster, y debajo, vemos c\u00f3mo son los centroides. Cuanto m\u00e1s intenso es el color, m\u00e1s importancia tiene esa categor\u00eda en ese cluster. Lo que ha identificado <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means es que hay un grupo muy numeroso (cluster 0) que casi no compra en comparaci\u00f3n con los otros. Los dem\u00e1s, aunque est\u00e9n formados por muy pocos usuarios s\u00ed tienen bastantes compras y se han agrupado en base a las diferentes categor\u00edas. Esto de cara al negocio puede que no sea muy \u00fatil ya que el objetivo es encontrar patrones de comportamiento en cuanto a las categor\u00edas para que los grupos se puedan activar de alguna manera y con <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means estamos descartando a 14k usuarios al meterlos en el grupo de usuarios que compran poco en vez de agruparlos tambi\u00e9n seg\u00fan las categor\u00edas que compran.<\/span>  &nbsp; <\/p>\n<h2><strong>Non-negative Matrix Factorization (NMF)<\/strong><\/h2>\n<p> <span style=\"font-weight: 400;\">Non-negative Matrix Factorization (NMF) es una t\u00e9cnica que tiene una propiedad inherente para clusterizar y es utilizada ampliamente para ello, especialmente cuando los datos de entrada son <\/span><i><span style=\"font-weight: 400;\">sparse<\/span><\/i><span style=\"font-weight: 400;\">. La idea de esta t\u00e9cnica es tratar de aproximar una matriz <\/span><span style=\"font-weight: 400;\">V<\/span><span style=\"font-weight: 400;\"> mediante la multiplicaci\u00f3n de dos matrices no negativas, <\/span><span style=\"font-weight: 400;\">W<\/span><span style=\"font-weight: 400;\"> y <\/span><span style=\"font-weight: 400;\">H<\/span><span style=\"font-weight: 400;\">:<\/span> <\/p>\n<p style=\"text-align: center;\"><b>V=<\/b><span style=\"font-weight: 400;\">WH<\/span><\/p>\n<p> <span style=\"font-weight: 400;\">Donde <\/span><span style=\"font-weight: 400;\">V<\/span><span style=\"font-weight: 400;\"> es nuestro dataset transpuesto, que tiene dimensiones <\/span><em><span style=\"font-weight: 400;\">f x <\/span><span style=\"font-weight: 400;\">n<\/span><\/em><span style=\"font-weight: 400;\">, donde <\/span><span style=\"font-weight: 400;\">f<\/span><span style=\"font-weight: 400;\"> es el n\u00famero de categor\u00edas y <\/span><span style=\"font-weight: 400;\">n<\/span><span style=\"font-weight: 400;\"> es el n\u00famero de usuarios. Por lo tanto, la matriz resultante <\/span><span style=\"font-weight: 400;\">W<\/span><span style=\"font-weight: 400;\"> tiene dimensiones <\/span><em><span style=\"font-weight: 400;\">f x <\/span><span style=\"font-weight: 400;\">t<\/span><\/em><span style=\"font-weight: 400;\">, y la matriz resultante <\/span><span style=\"font-weight: 400;\">H<\/span><span style=\"font-weight: 400;\">, <\/span><em><span style=\"font-weight: 400;\">t x <\/span><span style=\"font-weight: 400;\">n<\/span><\/em><span style=\"font-weight: 400;\">, siendo <\/span><em><span style=\"font-weight: 400;\">t<\/span><\/em><span style=\"font-weight: 400;\"> el n\u00famero de clusters recibido como entrada del algoritmo. Las columnas de la matriz resultante <\/span><em><span style=\"font-weight: 400;\">WH <\/span><\/em><span style=\"font-weight: 400;\">son una combinaci\u00f3n lineal de los <\/span><span style=\"font-weight: 400;\">t<\/span><span style=\"font-weight: 400;\"> clusters representados como las <\/span><span style=\"font-weight: 400;\">t<\/span><span style=\"font-weight: 400;\"> columnas de <\/span><span style=\"font-weight: 400;\">W<\/span><span style=\"font-weight: 400;\">. Esto es importante ya que es la base del NMF: se asume que cada columna en <\/span><span style=\"font-weight: 400;\">V <\/span><span style=\"font-weight: 400;\">est\u00e1 construida a partir de un n\u00famero peque\u00f1o <em>(<\/em><\/span><em><span style=\"font-weight: 400;\">t<\/span><\/em><span style=\"font-weight: 400;\"><em>)<\/em> de <\/span><i><span style=\"font-weight: 400;\">features<\/span><\/i><span style=\"font-weight: 400;\"> desconocidas y lo que se consigue con NMF es obtener esas <\/span><i><span style=\"font-weight: 400;\">features<\/span><\/i><span style=\"font-weight: 400;\">.<\/span>  <span style=\"font-weight: 400;\">Se puede pensar en las <\/span><em><span style=\"font-weight: 400;\">t<\/span><\/em><span style=\"font-weight: 400;\"> columnas de la matriz <\/span><span style=\"font-weight: 400;\">W <\/span><span style=\"font-weight: 400;\">como cada uno de los <\/span><em><span style=\"font-weight: 400;\">t<\/span><\/em><span style=\"font-weight: 400;\"> clusters de usuarios que contienen valores para cada una de las <\/span><em><span style=\"font-weight: 400;\">f<\/span><\/em><span style=\"font-weight: 400;\"> categor\u00edas. Cuanto m\u00e1s alto sea el valor de <em>wij<\/em><\/span><em><span style=\"font-weight: 400;\">\u00a0<\/span><\/em><span style=\"font-weight: 400;\">la categor\u00eda <\/span><em><span style=\"font-weight: 400;\">i<\/span><\/em><span style=\"font-weight: 400;\"> tendr\u00e1 m\u00e1s importancia en el cluster <\/span><em><span style=\"font-weight: 400;\">j<\/span><\/em><span style=\"font-weight: 400;\">. De forma equivalente, cada columna de la matriz <\/span><em><span style=\"font-weight: 400;\">H<\/span><\/em><span style=\"font-weight: 400;\"> representa el peso que tiene cada cluster para cada usuario.<\/span>  <span style=\"font-weight: 400;\">Siguiendo la explicaci\u00f3n anterior, en la matriz <\/span><em><span style=\"font-weight: 400;\">H<\/span><\/em><span style=\"font-weight: 400;\"> tenemos la informaci\u00f3n de a qu\u00e9 cluster pertenecen los usuarios. Solo tenemos que quedarnos con el cluster que tenga m\u00e1s peso para cada usuario:<\/span>  <img loading=\"lazy\" decoding=\"async\" class=\"wp-image-4134 aligncenter\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.38.04.png\" alt=\"\" width=\"132\" height=\"34\" \/>  <span style=\"font-weight: 400;\">Es decir, el usuario <\/span><em><span style=\"font-weight: 400;\">j<\/span><\/em><span style=\"font-weight: 400;\"> pertenece al cluster <\/span><em><span style=\"font-weight: 400;\">k<\/span><\/em><span style=\"font-weight: 400;\"> que maximice la expresi\u00f3n anterior. Los centroides de cada cluster se corresponden con las columnas de <\/span><em><span style=\"font-weight: 400;\">W<\/span><\/em><span style=\"font-weight: 400;\">.<\/span>  <span style=\"font-weight: 400;\">Para m\u00e1s informaci\u00f3n sobre este algoritmo, puedes consultar el <\/span><a href=\"https:\/\/iksinc.online\/2016\/03\/21\/what-is-nmf-and-what-can-you-do-with-it\/\" target=\"_blank\" rel=\"noopener\"><span style=\"font-weight: 400;\">siguiente art\u00edculo<\/span><\/a><span style=\"font-weight: 400;\">.<\/span>  <span style=\"font-weight: 400;\">Tras aplicar el algoritmo a nuestro dataset, obtenemos los siguientes resultados:<\/span>  <img loading=\"lazy\" decoding=\"async\" class=\"alignnone  wp-image-4137\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.25.50.png\" alt=\"\" width=\"912\" height=\"208\" \/> <img loading=\"lazy\" decoding=\"async\" class=\"alignnone  wp-image-4138\" src=\"https:\/\/www.makingscience.com\/wp-content\/uploads\/2021\/04\/Captura-de-pantalla-2021-02-16-a-las-11.26.10.png\" alt=\"\" width=\"892\" height=\"616\" \/>  <span style=\"font-weight: 400;\">Como vemos, los tama\u00f1os de los clusters son m\u00e1s uniformes y todos los clusters tienen un comportamiento asociado en cuanto a las categor\u00edas que compran, aunque hayan comprado poco. Con este resultado podr\u00edamos describir a nuestros clusters como:<\/span> <\/p>\n<ul>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Cluster 0: compran en las categor\u00edas 2, 9, 12, 17 y 19, pero con poca intensidad.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Cluster 1: tambi\u00e9n con poca intensidad pero sobre todo las categor\u00edas 13, 1 y 2.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Cluster 2: en este cluster se encuentran los usuarios que compran con m\u00e1s intensidad y sobre todo las categor\u00edas 28, 27, 13, 18, 8 y 29.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Cluster 3: compran en las categor\u00edas 19, 4 y 12.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Cluster 4: compran en las categor\u00edas 13, 27, 8 y 11.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Cluster 5: compran sobre todo en las categor\u00edas 11 y 9 aunque con poca intensidad.<\/span><\/li>\n<\/ul>\n<p> &nbsp;  <span style=\"font-weight: 400;\">En conclusi\u00f3n, aunque <\/span><i><span style=\"font-weight: 400;\">k<\/span><\/i><span style=\"font-weight: 400;\">-means puede funcionar bien en una gran cantidad de problemas, hay que realizar un estudio previo y entender bien de d\u00f3nde viene y c\u00f3mo calculamos nuestro dataset. Este an\u00e1lisis puede darnos grandes pistas sobre qu\u00e9 algoritmo es el m\u00e1s apropiado para realizar una buena segmentaci\u00f3n. Aparte de esto, no hay que perder el punto de vista de negocio que va a ser el que finalmente nos haga afinar el algoritmo que elijamos para que los grupos sean lo m\u00e1s activables posible.<\/span>  Si te ha gustado este art\u00edculo y quieres saber m\u00e1s sobre <em>K-means<\/em> aplicado a tu negocio, escr\u00edbenos a info@makingscience.com<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Clustering: m\u00e1s all\u00e1 de k-means Clustering hace referencia a un conjunto de t\u00e9cnicas y algoritmos no supervisados que tienen como objetivo segmentar datos seg\u00fan grupos homog\u00e9neos, tambi\u00e9n llamados clusters. Dado un conjunto de puntos en un dataset, podemos aplicar estos algoritmos para clasificar cada punto en un grupo espec\u00edfico. Idealmente los puntos de un mismo [&hellip;]<\/p>\n","protected":false},"author":21,"featured_media":9676,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[34],"tags":[57,74,62],"class_list":["post-9662","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technology-ai","tag-data-analytics","tag-data-science","tag-gauss-ai"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/posts\/9662","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/users\/21"}],"replies":[{"embeddable":true,"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/comments?post=9662"}],"version-history":[{"count":0,"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/posts\/9662\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/media\/9676"}],"wp:attachment":[{"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/media?parent=9662"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/categories?post=9662"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.makingscience.com\/es\/wp-json\/wp\/v2\/tags?post=9662"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}