
K-Means explained: I asked Chile's weather for four kinds of day and it found Santiago
What K-Means is, how Lloyd's algorithm works and when it misleads, measured on daily weather from seven Chilean cities between 1984 and 2026. Unscaled, humidity accounts for 76% of the separation; with k = 8, only 18% of random starts end within 0.1% of the best inertia found; and no criterion agrees on how many groups there are.
I asked K-Means to split 74,145 days of weather into four groups, without telling it which city each day came from. One of the groups turned out to be almost only Santiago: 92% of Santiago’s days landed there. It looks like a discovery, but it is worth checking what went into the distance: in this data, Santiago’s pressure is much lower than the other cities’.
This is the third part of the series, with the same NASA POWER daily data I used for Random Forest and linear regression: seven cities from La Serena to Punta Arenas, 1984 to 2026, measured inside a container with 8 CPUs. The difference is that here there is nothing to predict. K-Means does not learn from correct answers; it only groups what looks alike.
What is K-Means?
K-Means splits the rows of a table into k groups. Each group has a centroid, a point that works as its prototype, and each row belongs to the nearest centroid. The goal is to make the sum of squared distances between each row and its centroid, called inertia, as small as possible.
Here each row is one day in one city, described by eight variables: max and min temperature, rain, max wind, radiation, humidity, pressure and dew point. I do not use the city, the latitude or the date. Whatever comes out comes only from the weather.
How it works: Lloyd’s algorithm
The classic way to compute it is the algorithm Stuart Lloyd published in 1982, and it is what scikit-learn uses by default. It has two steps that repeat:
- Assign: each row goes to the nearest centroid.
- Move: each centroid moves to the mean of the rows it got.
Each step can only lower inertia or keep it the same, so the algorithm ends; in practice it stops when the centroids barely move or when it reaches a maximum number of iterations. It ends, but not necessarily at the best possible result.
From the poor start (all four centroids taken from the warmest quarter of days), inertia fell from 202,896 to 27,909 in 25 iterations. From k-means++ it started at 41,781 and reached the same 27,909, also in 25 iterations: with two variables and four groups both starts end at practically the same solution.
To draw it I used only two variables, max temperature and humidity, and four centroids. For the poor start I deliberately took them from the warmest quarter of days. Inertia starts at 202,896; after four iterations it is already down to 28,063 and at 25 it stops at 27,909. From k-means++ it starts at 41,781 and reaches practically the same solution, also in 25 iterations. With two variables and four groups, in this case, the start barely changed the end. Further down you will see when it does.
The scale redraws the map
K-Means does not know what a degree or a millimeter is. It only subtracts numbers. If a variable moves over a large range, it dominates the distance even if it is not the most important one.
share of the separation between centroids
share of each city’s days in each group · agreement with the city (ARI) 0.17
| rainy | mild | semi-dry | dry | |
|---|---|---|---|---|
| La Serena | 4 % | 14 % | 75 % | 7 % |
| Valparaíso | 31 % | 68 % | 0 % | 0 % |
| Santiago | 7 % | 1 % | 26 % | 66 % |
| Concepción | 37 % | 37 % | 25 % | 1 % |
| Temuco | 46 % | 43 % | 11 % | 0 % |
| Puerto Montt | 62 % | 38 % | 0 % | 0 % |
| Punta Arenas | 71 % | 29 % | 0 % | 0 % |
share of the separation between centroids
share of each city’s days in each group · agreement with the city (ARI) 0.26
| rainy | cold | mild | dry | |
|---|---|---|---|---|
| La Serena | 9 % | 0 % | 77 % | 14 % |
| Valparaíso | 23 % | 6 % | 71 % | 0 % |
| Santiago | 7 % | 1 % | 0 % | 92 % |
| Concepción | 48 % | 1 % | 51 % | 0 % |
| Temuco | 56 % | 4 % | 41 % | 0 % |
| Puerto Montt | 74 % | 0 % | 26 % | 0 % |
| Punta Arenas | 13 % | 84 % | 3 % | 0 % |
share of the separation between centroids
Not computed for this version.
share of each city’s days in each group · agreement with the city (ARI) 0.18
| rainy | cold | mild | dry | |
|---|---|---|---|---|
| La Serena | 1 % | 5 % | 68 % | 26 % |
| Valparaíso | 2 % | 9 % | 88 % | 0 % |
| Santiago | 2 % | 18 % | 3 % | 77 % |
| Concepción | 10 % | 28 % | 56 % | 6 % |
| Temuco | 10 % | 42 % | 46 % | 2 % |
| Puerto Montt | 15 % | 43 % | 42 % | 0 % |
| Punta Arenas | 1 % | 95 % | 4 % | 0 % |
Raw, humidity (0 to 100) accounts for 76 % of the separation and wind for 0.04 %. Standardized, no variable passes 17.3 %, and the raw and standardized partitions barely agree (ARI 0.37). Santiago, whose reanalysis cell is the highest (mean pressure 88.6 kPa against 96.1 or more elsewhere), puts 92 % of its days in a single standardized group; without pressure, still 77 %.
Humidity goes from 0 to 100 and its standard deviation in this data is 18.8; wind’s is 2.7. With raw variables, humidity accounted for 75.9% of the separation between the four centroids and wind for 0.04%. Standardized, with mean 0 and deviation 1, no variable passed 17.3%. The two versions split the days quite differently: their agreement measured with the adjusted Rand index (ARI), which is 1 if two partitions are identical and close to 0 if they agree as if by chance, is 0.37.
With standardized variables the Santiago group appears. Its centroid has 23.7 °C max, 38.5% humidity and 89.6 kPa pressure, and 92% of Santiago’s days fall in it. Santiago’s mean pressure in this data is 88.6 kPa; the other cities range from 96.1 to 101.1. It is lower than the city itself would read: NASA POWER returns the value of a reanalysis cell of half a degree of latitude by ⅝ of a degree of longitude, and Santiago’s probably includes Andean terrain higher than the city. For K-Means, that is a huge difference in one of the eight columns.
To see what happened without pressure, I repeated the fit without it. Santiago stayed concentrated: 77% of its days landed in one group, a dry one with 38% humidity. This experiment does not isolate how much each variable contributes to the group; it only shows that without pressure Santiago stays concentrated. Without pressure the rest changed too: the moderate rainy regime gave way to a heavy rain one averaging 18.4 mm.
None of this is a K-Means error. It is the distance I chose. Scaling is a decision, and so is which columns go in.
How many groups? Three criteria, three answers
K-Means needs to be told k. The temptation is to look for the “elbow”: the point where adding groups stops lowering inertia much.
group sizes with k = 2
group sizes with k = 3
group sizes with k = 4
group sizes with k = 5
group sizes with k = 6
group sizes with k = 7
group sizes with k = 8
group sizes with k = 9
group sizes with k = 10
group sizes with k = 11
group sizes with k = 12
Inertia keeps falling, from 434,884 with 2 groups to 140,405 with 12, without an obvious bend. On one sample the silhouette peaks at k = 3, but repeating it on five samples of 10,000 days, k = 3, 5 and 6 land within the sampling noise (3: 0.2771, 4: 0.2724, 5: 0.2778, 6: 0.2775; standard deviation up to 0.0042); Davies-Bouldin prefers k = 5. With silhouettes between 0.25 and 0.28, separation is weak and no criterion tested points to a clear winner.
Inertia always fell, from 434,884 with two groups to 140,405 with twelve, with no clear bend. The silhouette, which compares how much each day resembles its group against the neighboring group, peaked at k = 3 on a sample of 10,000 days. Repeating it on five different samples, k = 3, 5 and 6 landed at 0.2771, 0.2778 and 0.2775, with standard deviations of 0.003 to 0.004 between samples: differences smaller than the sampling noise. The Davies-Bouldin index preferred k = 5.
Every silhouette stayed between 0.25 and 0.28. On a scale from −1 to 1, that means poorly separated groups with these eight variables, and none of the criteria tested points to a clear winner. I chose k = 4 for the rest of the post because four regimes can be named and read, not because the data asks for it. In practice, that is the honest way to choose k: by its use, checking that the groups make sense.
The start is a lottery
Lloyd always converges, but to a local minimum that depends on where the centroids started. To measure how much it matters, I ran K-Means with a single start and 100 different seeds, random and with k-means++, which picks initial centroids far from each other.
With k = 4, 92% of random starts and 96% of k-means++ starts land within 0.1% of the best, and the worst run of both methods is 7.8% worse. With k = 8 the lottery gets harder: 18% against 38%, and the worst random run ends 26.1% above the best against 6.8% for k-means++. That is why it pays to set n_init above 1 and keep the best fit: with k-means++, scikit-learn's default runs only one.
With k = 4, 92% of random runs and 96% of k-means++ runs got within 0.1% of the best inertia found across the 200 runs. The worst run of both methods ended 7.8% above: a bad local minimum exists and both find it sometimes. With k = 8 the difference shows: only 18% of random runs got close to the best, against 38% for k-means++, and the worst random run ended 26.1% above, against 6.8% for k-means++.
That is why scikit-learn uses k-means++ by default and lets you repeat the fit several times with n_init to keep the one with the lowest inertia. Watch the default: with k-means++, n_init='auto' runs a single fit. If runs were independent, with 38% success each, ten starts would all fail with a probability close to 0.8%.
Four weather regimes, before and after
With k = 4 on standardized variables, the centroids read as kinds of day. I named them by their centroid in real units: a rainy regime (13.2 °C max, 4.6 mm, 86% humidity), a cold and windy one (8.4 °C, 10.3 m/s max wind), a mild one (20.8 °C, almost no rain) and Santiago’s dry inland regime.
- rainy13.2 / 6.5 °C · 86 % · 4.6 mm · 4.4 m/s · 99.7 kPadays with ≥1 mm: 52 % · 32.8 % of all days
- cold and windy8.4 / 3.9 °C · 82 % · 1.8 mm · 10.3 m/s · 99.6 kPadays with ≥1 mm: 43 % · 13.7 % of all days
- dry inland23.7 / 9.6 °C · 39 % · 0.3 mm · 4.8 m/s · 89.6 kPadays with ≥1 mm: 7 % · 15.1 % of all days
- mild20.8 / 11.9 °C · 72 % · 0.3 mm · 5.6 m/s · 99.0 kPadays with ≥1 mm: 7 % · 38.4 % of all days
share of each month’s days, training | 2016-2026
share of each city’s days, training | 2016-2026
- La Serena
- Valparaíso
- Santiago
- Concepción
- Temuco
- Puerto Montt
- Punta Arenas
The rainy regime covers 32.8 % of training days and 30.5 % of 2016-2026 days. The dry inland regime is almost all of Santiago and grows in La Serena from 13.7 % to 18.0 %. These are shares of assigned days, not a climate trend test: the source switches products in recent months and there is no uncertainty estimate here.
The cold and windy regime is mostly Punta Arenas: 84% of its days. In 2016-2026, the mild regime holds more than 60% of the days from November to March and the rainy one reaches 57% of July days. It rained at least 1 mm on 52% of rainy-regime days and on less than 8% of mild or dry days.
Then I assigned the 2016-2026 days to those same centroids, without retraining. The rainy regime went from 32.8% to 30.5% of days, and the dry inland regime rose in La Serena from 13.7% to 18.0%. These are shares of days assigned to fixed centroids, not a test of climate change: I did not compute uncertainty and the source switches products in the most recent months, as I explained in the Random Forest post.
The days that do not fit
A centroid is a prototype, and the distance to it says how typical a day is. That turns K-Means into a cheap anomaly detector.
- 1984-2012
- 2016-2026
farthest 2016-2026 days
Temuco · July 9, 2020 · 104.2 mm distance 19.4
10.9 / 2.5 °C · 95.9 % · 9.2 m/s · 0.8 MJ/m² · 98.1 kPa · 6.9 °C
Puerto Montt · April 26, 2020 · 96.1 mm distance 17.7
17.7 / 10.6 °C · 89.4 % · 4.5 m/s · 4.5 MJ/m² · 100.6 kPa · 11.3 °C
Concepción · August 1, 2026 · 89.4 mm distance 16.4
12.1 / 9.2 °C · 95.0 % · 7.0 m/s · 1.4 MJ/m² · 99.3 kPa · 10.3 °C
Concepción · August 25, 2026 · 78.3 mm distance 14.3
11.8 / 10.0 °C · 94.7 % · 5.0 m/s · 0.8 MJ/m² · 99.4 kPa · 10.1 °C
Temuco · June 16, 2017 · 77.6 mm distance 14.3
12.5 / -0.1 °C · 96.5 % · 7.2 m/s · 0.8 MJ/m² · 97.2 kPa · 4.8 °C
La Serena · May 11, 2017 · 70.4 mm distance 13.0
15.5 / 12.8 °C · 84.2 % · 10.0 m/s · 4.6 MJ/m² · 96.0 kPa · 11.7 °C
0.051% of 2016-2026 days pass the training threshold, about half of the 0.1% it marks in training. The farthest are heavy rain days, above 70 mm: no centroid represents them because K-Means pulls centroids toward where most days are. Distance flags a day as unusual for this model; it does not say whether the data are wrong or the event was real.
I took the 99.9th percentile of the training distance as the threshold, 9.45 in standardized units. In 2016-2026, 0.051% of days passed it. The farthest from any centroid are heavy rain days: Temuco on July 9, 2020 with 104.2 mm, Puerto Montt on April 26, 2020 with 96.1 mm and Concepción on August 1, 2026 with 89.4 mm. No centroid represents them: the rainy regime’s centroid averages 4.6 mm, because days of moderate rain are far more common.
Distance does not say whether the data are wrong or the storm was real. It only says this model represents it poorly. For a monitor, that is already useful: it is the alarm someone has to review.
The rainy regime also has the lowest silhouette: 0.241 on a sample of 10,000 days from 2016-2026, against 0.283 to 0.301 for the other three.
How fast it is
K-Means is cheap per row: assigning a new day means comparing eight numbers against four centroids. With scikit-learn that took 0.042 ms per row; with a direct NumPy operation, 0.0024 ms. Training is another story when k grows, because each iteration compares every row against every centroid.
At twice real time.
Eight times slower than real time.
With k = 256, K-Means ran 92 iterations: 34 ms each on 1 thread and 7.9 ms on 8. With k = 64 it hit the cap of 100 iterations without converging, so that time is a lower bound. MiniBatchKMeans ended with 3.7% more inertia at k = 256. With k = 4 everything takes less than 26 ms.
With 8 threads, K-Means with k = 256 was 4.3 times faster than with one. MiniBatchKMeans, which updates centroids with batches of 4,096 rows instead of scanning the whole table, was 5.6 times faster than K-Means on 8 threads, in exchange for 3.7% more inertia. With k = 4 the difference does not matter: every fit took less than 26 milliseconds.
Where K-Means lives in a real system
the same grouping, elsewhere
- image paletteeach pixel takes its centroid color
- customer segmentsa centroid per profile, reviewed by people
- distance monitoralert when a new row is far from every centroid
Measured on the 101,401 rows of this table with k = 256: K-Means takes 3.14 s with 1 thread and 0.73 s with 8; MiniBatchKMeans takes 0.13 s with 8 threads and ends with 3.7% more inertia. Large-scale indexes usually train on far more vectors and dimensions; these times only show the shape of the trade-off.
- Vector indexes. IVF-type indexes, such as FAISS’s
IndexIVFFlat, split the space into cells with K-Means. Each stored vector lands in its centroid’s list, and a query only searches inside thenprobenearest cells. It is the same idea as the anomaly detector, used to avoid comparing against everything. - Compression and palettes. Reducing an image to 16 colors is K-Means over the pixels: each pixel takes its centroid’s color.
- Segmentation. Grouping customers, sessions or sensors into profiles a person can review and name, as I did with the weather regimes.
- Monitoring. Distance to the nearest centroid as an alarm for rows that look like nothing known, just like the storms above.
When I would choose it: when I want a few understandable prototypes, the variables are numeric and comparable after scaling, and the groups can be roughly round. When I would not: with elongated groups or very different densities, with many categorical variables, or when I cannot justify k. And never without checking which columns dominate the distance.
Sources
- Lloyd, S. P. (1982). “Least squares quantization in PCM”. IEEE Transactions on Information Theory, 28(2), 129–137. DOI 10.1109/TIT.1982.1056489.
- Arthur, D. and Vassilvitskii, S. (2007). “k-means++: The advantages of careful seeding”. Proceedings of the Eighteenth Annual ACM-SIAM Symposium on Discrete Algorithms. Society for Industrial and Applied Mathematics.
- Sculley, D. (2010). “Web-scale k-means clustering”. Proceedings of the 19th International Conference on World Wide Web, 1177–1178. DOI 10.1145/1772690.1772862.
- Rousseeuw, P. J. (1987). “Silhouettes: A graphical aid to the interpretation and validation of cluster analysis”. Journal of Computational and Applied Mathematics, 20, 53–65. DOI 10.1016/0377-0427(87)90125-7.
- Davies, D. L. and Bouldin, D. W. (1979). “A Cluster Separation Measure”. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-1(2), 224–227. DOI 10.1109/TPAMI.1979.4766909.
- scikit-learn,
KMeans(Lloyd by default, k-means++ andn_init) and the clustering guide. - FAISS, index overview, section on
IndexIVFFlatandnprobe. - NASA POWER, Daily API and data sources.
Comments
No comments yet. The first one is yours.