K-Means Kümeleme Algoritması: Mantık, K Değeri ve Python Uygulaması

Kümeleme mantığından elbow method'a, K değerinin seçiminden gerçek bir müşteri segmentasyonu uygulamasına kadar K-Means'i adım adım inceliyorum.

  • Makine Öğrenmesi
  • Araçlar

Her veri setinde elimizde önceden belirlenmiş sınıflar olmayabilir. Bazen amacımız “Bu kayıt hangi sınıfa ait?” sorusunu cevaplamak değil, “Verinin içinde doğal olarak hangi gruplar oluşuyor?” sorusunu keşfetmektir.

Kümeleme algoritmaları tam olarak bu problemi çözmeye çalışır. K-Means ise bu alanda en bilinen ve başlangıç seviyesinde öğrenilmesi en kolay yöntemlerden biri. Müşteri segmentasyonu, ürün gruplama, davranış analizi ve benzer gözlemleri bir araya getirme gibi birçok problemde kullanılabilir.

“K-Means'in amacı veriyi etiketlemek değil; verinin içindeki benzerlikleri kullanarak doğal grupları ortaya çıkarmaktır.”

K-Means Nedir?

K-Means, gözetimsiz öğrenme (unsupervised learning) kapsamında yer alan bir kümeleme algoritmasıdır. Temel fikir, gözlemleri birbirine benzer olacak şekilde K adet kümeye ayırmak ve her kümenin merkezini, yani centroid'ini iteratif olarak güncellemektir.

K-Means'in temel mantığı
  1. Veriyi
    Al
  2. K
    Seç
  3. Merkezleri
    Başlat
  4. Kümeleri
    Ata
  5. Merkezleri
    Güncelle
  6. Tekrarla

1. Kümeleme Mantığını Basit Bir Örnekle Anlamak

Bir mağazanın müşterilerini yıllık harcama ve alışveriş sıklığı değişkenlerine göre gruplamak istediğimizi düşünelim. Veri setinde herhangi bir “VIP”, “standart” veya “potansiyel” etiketi olmadığını varsayalım.

K-Means, müşterilerin bu iki özellik üzerindeki benzerliklerini kullanarak doğal kümeler oluşturmaya çalışır.

Küme 1

Yüksek değerli

Yüksek harcama ve yüksek alışveriş sıklığı gösteren müşteriler.

Küme 2

Potansiyel

Yüksek sıklık ancak daha düşük ortalama harcama gösteren müşteriler.

Küme 3

Düşük etkileşim

Hem harcaması hem de alışveriş sıklığı düşük olan müşteriler.

Kümeleme

Etiketler sonradan gelir

Küme isimlerini algoritma vermez; analist oluşan grupları yorumlar.

2. K Değeri Nedir?

K-Means'in en kritik parametresi K. K, verinin kaç kümeye ayrılacağını ifade eder.

KAnlamı
2Veriyi 2 kümeye ayır
3Veriyi 3 kümeye ayır
4Veriyi 4 kümeye ayır
5Veriyi 5 kümeye ayır

Buradaki asıl soru şu: K değerini nereden bileceğiz? İşte bu noktada elbow method gibi yöntemlerden yararlanabiliriz.

3. K-Means Algoritması Nasıl Çalışır?

Adım 1 — Başlangıç merkezlerini seçmek

Algoritma K adet başlangıç merkezi belirler. Bu merkezler ilk aşamada rastgele veya uygun bir başlatma yöntemiyle seçilebilir.

Adım 2 — Her noktayı en yakın merkeze atamak

Her gözlemin centroid'lere olan mesafesi hesaplanır. Gözlem, kendisine en yakın merkezin bulunduğu kümeye atanır.

Adım 3 — Centroid'leri güncellemek

Her kümedeki gözlemlerin ortalaması alınarak yeni merkez hesaplanır.

Adım 4 — Atamaları tekrar yapmak

Yeni merkezler kullanılarak gözlemler yeniden kümelere atanır.

Adım 5 — Yakınsama sağlanana kadar devam etmek

Merkezler ve küme atamaları artık önemli ölçüde değişmiyorsa algoritma durur.

İteratif süreç
  1. Centroid
  2. Mesafe
  3. Atama
  4. Yeni
    Centroid
  5. Kontrol

4. K-Means Hangi Mesafeyi Kullanır?

K-Means uygulamalarında en yaygın yaklaşım, Öklid mesafesi üzerinden gözlemlerin merkezlere uzaklığını değerlendirmektir. İki boyutlu bir örnek için:

d=(x2−x1)2+(y2−y1)2d = \sqrt{(x_2 - x_1)^2 + (y_2 - y_1)^2}

Algoritmanın temel amacı, gözlemlerin kendi kümelerindeki centroid'e olan uzaklıklarını mümkün olduğunca küçük tutmaktır.

5. Elbow Method Nedir?

Elbow method, uygun K değerini seçmek için kullanılan yaygın yöntemlerden biri. Farklı K değerleri için model çalıştırılır ve her çözümün küme içi hata değeri, yani inertia incelenir.

K arttıkça kümeler daha küçük parçalara ayrıldığı için inertia genellikle düşer. Ancak bir noktadan sonra K'yı artırmak çok daha az ek fayda sağlar.

Elbow method'un fikri: K değerini artırdıkça inertia'daki düşüşün belirgin biçimde yavaşladığı noktayı aramak. Bu nokta tek başına “matematiksel olarak kesin doğru K” değildir; verinin yapısı ve iş problemi birlikte değerlendirilmelidir.

6. Python ile Elbow Method

Python'da scikit-learn ile K-Means uygulamak oldukça kolay.

K değerlerini karşılaştırmak
from sklearn.cluster import KMeans

inertias = []

for k in range(2, 9):
    model = KMeans(
        n_clusters=k,
        random_state=42,
        n_init="auto"
    )
    model.fit(X)
    inertias.append(model.inertia_)

# inertias listesini K değerlerine karşı çizdir

Sonuçları bir çizgi grafik üzerinde gösterdiğimizde dirsek şeklindeki kırılma noktasını arayabiliriz.

7. Gerçek Hayat Örneği: Müşteri Segmentasyonu

Şimdi küçük bir müşteri segmentasyonu problemi kuralım. Her müşteri için iki değişkenimiz olduğunu düşünelim:

DeğişkenAçıklama
Yıllık harcamaMüşterinin yıllık toplam alışveriş tutarı
Alışveriş sıklığıMüşterinin yıllık işlem sayısı

Veriyi hazırlamak

Değişkenleri seçmek
import pandas as pd

df = pd.read_csv("musteriler.csv")

X = df[["yillik_harcama", "alisveris_sikligi"]]

Ölçeklendirme neden önemli?

Örneğin yıllık harcama 5.000–100.000 aralığındayken alışveriş sıklığı 1–50 aralığında olabilir. Değişkenlerin ölçekleri çok farklıysa, büyük ölçekli değişken mesafe hesabını gereğinden fazla etkiler.

StandardScaler ile ölçekleme
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

K-Means modelini kurmak

Modeli çalıştırmak
from sklearn.cluster import KMeans

kmeans = KMeans(
    n_clusters=3,
    random_state=42,
    n_init="auto"
)

df["cluster"] = kmeans.fit_predict(X_scaled)

Artık her müşterinin hangi kümeye atandığını cluster sütununda görebiliriz.

8. Kümeleri Yorumlamak

Algoritmanın ürettiği “0, 1, 2” gibi etiketler tek başına anlam taşımaz. Analistin görevi, kümelerin ortalama özelliklerini inceleyerek bu gruplara anlam vermektir.

Küme ortalamaları
segment_ozet = df.groupby("cluster")[
    ["yillik_harcama", "alisveris_sikligi"]
].mean()

print(segment_ozet)
KümeHarcamaSıklıkÖrnek yorum
0YüksekYüksekYüksek değerli müşteriler
1OrtaYüksekAktif / potansiyel müşteriler
2DüşükDüşükDüşük etkileşimli müşteriler

Buradaki segment isimleri algoritmanın çıktısı değil; veri analistinin kümeleri iş bağlamında yorumlamasıyla ortaya çıkar.

9. K-Means Sonuçlarını Görselleştirmek

İki boyutlu bir müşteri segmentasyonunda kümeleri scatter plot ile görmek oldukça faydalı.

Kümeleri çizdirmek
import matplotlib.pyplot as plt

plt.scatter(
    X_scaled[:, 0],
    X_scaled[:, 1],
    c=df["cluster"]
)

plt.xlabel("Yıllık Harcama")
plt.ylabel("Alışveriş Sıklığı")
plt.title("Müşteri Segmentleri")
plt.show()

Görselleştirme sayesinde kümelerin birbirinden ne kadar ayrıştığını ve bazı gözlemlerin neden sınıra yakın kaldığını daha kolay inceleyebiliriz.

10. K-Means'in Güçlü ve Zayıf Yönleri

Avantaj

Basit ve hızlı

Temel mantığı kolay anlaşılır ve büyük veri setlerinde pratik olabilir.

Avantaj

Yorumlanabilir

Kümelerin ortalamaları üzerinden segmentleri anlamlandırmak kolaydır.

Sınırlılık

K seçilmelidir

Küme sayısı baştan belirlenir; uygun K'yı seçmek analizin önemli bir parçasıdır.

Sınırlılık

Geometrik yapı

Küme şekilleri ve değişken ölçekleri yöntemin sonucunu etkileyebilir.

11. Nelere Dikkat Etmeli?

Ölçekleme

Ölçeği kontrol et

Mesafe tabanlı yöntemlerde değişkenlerin ölçeği sonucu doğrudan etkiler.

K seçimi

Tek metriğe güvenme

Elbow method'u iş problemi ve diğer metriklerle birlikte değerlendir.

Başlangıç

Seed kullan

Farklı başlangıç merkezlerinin etkisini düşün ve sonucu tekrarlanabilir yap.

Aykırı değer

Uç gözlemler

Çok uç gözlemler centroid'leri belirgin biçimde çekebilir.

Doğrulama

Ayrışmayı ölç

Silhouette score gibi ölçülerle küme ayrışmasını ayrıca incele.

İş yorumu

Segmente dönüştür

Matematiksel kümeleri gerçek hayatta anlamlı segmentlere çevir.

12. Silhouette Score ile Ek Kontrol

Elbow method'un yanında kümeleme kalitesini değerlendirmek için silhouette score gibi metrikler de kullanılabilir. Bu skor, gözlemlerin kendi kümesine ne kadar yakın ve diğer kümelerden ne kadar uzak olduğunu değerlendirmeye yardımcı olur.

Silhouette score hesaplamak
from sklearn.metrics import silhouette_score

score = silhouette_score(
    X_scaled,
    df["cluster"]
)

print("Silhouette Score:", score)

Bu tür metrikler K seçiminde destekleyici bilgi sağlar; ancak tek başına iş probleminin gerektirdiği segment sayısını belirlemez.

13. K-Means Ne Zaman Kullanılmalı?

SenaryoUygunluk
Müşteri segmentasyonuBenzer davranışlara göre gruplama için kullanılabilir.
Ürün gruplamaBenzer özelliklere sahip ürünleri keşfetmek için değerlendirilebilir.
Davranış analiziKullanıcıları aktivite desenlerine göre ayırmak için kullanılabilir.
Çok farklı şekilli kümelerK-Means yerine farklı kümeleme yöntemleri incelenebilir.

Sonuç

K-Means, kümeleme problemlerine giriş yapmak için güçlü bir algoritma. Temel mantık oldukça sade olsa da iyi bir analiz için yalnızca modeli çalıştırmak yeterli değil.

Doğru değişkenleri seçmek, gerekirse ölçekleme yapmak, uygun K değerini belirlemek, elbow method ve silhouette score gibi yöntemlerle sonucu incelemek ve son olarak kümeleri iş bağlamında yorumlamak gerekiyor.

“K-Means sana kümeleri verir; bu kümelerin ne anlama geldiğini söylemek ise analistin işidir.”

Özellikle müşteri segmentasyonu gibi problemlerde bu yaklaşım, ham veriyi daha anlamlı davranış gruplarına dönüştürmek için iyi bir başlangıç noktası olabilir.

Not every data set comes with predefined classes. Sometimes the goal is not to answer “Which class does this record belong to?” but to explore “Which groups form naturally inside the data?”

Clustering algorithms are built exactly for this problem, and K-Means is one of the best known and easiest to start with. It can be used for customer segmentation, product grouping, behaviour analysis and many other problems where similar observations need to be brought together.

“The goal of K-Means is not to label the data; it is to reveal the natural groups using the similarities inside it.”

What Is K-Means?

K-Means is a clustering algorithm that belongs to unsupervised learning. The core idea is to split the observations into K clusters so that the members of each cluster are similar to one another, and to update the centre of each cluster — its centroid — iteratively.

The basic logic of K-Means
  1. Take the
    Data
  2. Choose
    K
  3. Initialise
    Centroids
  4. Assign
    Clusters
  5. Update
    Centroids
  6. Repeat

1. Understanding Clustering with a Simple Example

Imagine we want to group the customers of a shop by annual spending and purchase frequency, and that the data set carries no “VIP”, “standard” or “potential” label at all.

K-Means tries to form natural clusters using the similarity of customers across these two features.

Cluster 1

High value

Customers with high spending and high purchase frequency.

Cluster 2

Potential

Customers who buy often but with a lower average basket.

Cluster 3

Low engagement

Customers with both low spending and low purchase frequency.

Clustering

Labels come later

The algorithm does not name the clusters; the analyst interprets the groups that form.

2. What Is the K Value?

The most critical parameter of K-Means is K, which states how many clusters the data will be split into.

KMeaning
2Split the data into 2 clusters
3Split the data into 3 clusters
4Split the data into 4 clusters
5Split the data into 5 clusters

The real question is this: how do we know the value of K? This is where methods such as the elbow method help.

3. How Does the K-Means Algorithm Work?

Step 1 — Choose the initial centroids

The algorithm picks K starting centres. At this first stage they can be chosen randomly or with a suitable initialisation method.

Step 2 — Assign every point to the nearest centre

The distance of each observation to the centroids is calculated, and the observation joins the cluster of the nearest centre.

Step 3 — Update the centroids

The new centre is calculated as the mean of the observations in each cluster.

Step 4 — Assign again

Using the new centres, the observations are assigned to clusters once more.

Step 5 — Continue until convergence

When the centres and the cluster assignments no longer change significantly, the algorithm stops.

The iterative process
  1. Centroid
  2. Distance
  3. Assignment
  4. New
    Centroid
  5. Check

4. Which Distance Does K-Means Use?

The most common approach in K-Means is to evaluate how far observations are from the centres using Euclidean distance. For a two-dimensional example:

d=(x2−x1)2+(y2−y1)2d = \sqrt{(x_2 - x_1)^2 + (y_2 - y_1)^2}

The main goal of the algorithm is to keep the distances between observations and the centroid of their own cluster as small as possible.

5. What Is the Elbow Method?

The elbow method is one of the common ways to choose a suitable K. The model is run for different values of K and the within-cluster error of each solution — its inertia — is examined.

As K grows the clusters break into smaller pieces, so inertia usually falls. After a certain point, however, increasing K adds far less benefit.

The idea behind the elbow method: look for the point where the drop in inertia clearly slows down as K increases. That point is not on its own “the mathematically correct K”; the structure of the data and the business problem should be considered together with it.

6. The Elbow Method in Python

Applying K-Means with scikit-learn in Python is straightforward.

Comparing values of K
from sklearn.cluster import KMeans

inertias = []

for k in range(2, 9):
    model = KMeans(
        n_clusters=k,
        random_state=42,
        n_init="auto"
    )
    model.fit(X)
    inertias.append(model.inertia_)

# plot the inertias against the K values

Plotting the results on a line chart lets us look for the elbow-shaped break point.

7. A Real-Life Example: Customer Segmentation

Let us set up a small customer segmentation problem. Assume we have two variables for every customer:

VariableDescription
Annual spendingThe customer's total purchase amount for the year
Purchase frequencyThe customer's number of transactions per year

Preparing the data

Selecting the variables
import pandas as pd

df = pd.read_csv("customers.csv")

X = df[["annual_spending", "purchase_frequency"]]

Why does scaling matter?

Annual spending may range between 5,000 and 100,000 while purchase frequency ranges between 1 and 50. When the scales differ this much, the larger variable dominates the distance calculation.

Scaling with StandardScaler
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

Building the K-Means model

Running the model
from sklearn.cluster import KMeans

kmeans = KMeans(
    n_clusters=3,
    random_state=42,
    n_init="auto"
)

df["cluster"] = kmeans.fit_predict(X_scaled)

We can now see which cluster each customer belongs to in the cluster column.

8. Interpreting the Clusters

Labels such as “0, 1, 2” carry no meaning on their own. It is the analyst's job to give those groups meaning by examining the average properties of each cluster.

Cluster averages
segment_summary = df.groupby("cluster")[
    ["annual_spending", "purchase_frequency"]
].mean()

print(segment_summary)
ClusterSpendingFrequencyExample interpretation
0HighHighHigh-value customers
1MediumHighActive / potential customers
2LowLowLow-engagement customers

These segment names are not the output of the algorithm; they appear when the analyst interprets the clusters in a business context.

9. Visualising the K-Means Results

In a two-dimensional customer segmentation, seeing the clusters on a scatter plot is very useful.

Plotting the clusters
import matplotlib.pyplot as plt

plt.scatter(
    X_scaled[:, 0],
    X_scaled[:, 1],
    c=df["cluster"]
)

plt.xlabel("Annual Spending")
plt.ylabel("Purchase Frequency")
plt.title("Customer Segments")
plt.show()

The visualisation makes it easier to see how well the clusters separate and why some observations sit close to a boundary.

10. Strengths and Weaknesses of K-Means

Strength

Simple and fast

The core logic is easy to grasp and it stays practical on large data sets.

Strength

Interpretable

Making sense of the segments through cluster averages is straightforward.

Limitation

K must be chosen

The number of clusters is set upfront, and choosing a suitable K is a real part of the analysis.

Limitation

Geometry matters

Cluster shapes and variable scales can affect the outcome of the method.

11. What to Watch Out For

Scaling

Check the scale

In distance-based methods the scale of the variables affects the result directly.

Choosing K

Do not trust one metric

Weigh the elbow method together with the business problem and other metrics.

Initialisation

Use a seed

Consider the effect of different starting centres and make the result reproducible.

Outliers

Extreme observations

Very extreme observations can pull the centroids noticeably.

Validation

Measure separation

Examine cluster separation separately with measures such as the silhouette score.

Business view

Turn them into segments

Translate mathematical clusters into segments that mean something in real life.

12. An Extra Check with the Silhouette Score

Alongside the elbow method, metrics such as the silhouette score can be used to evaluate clustering quality. The score helps assess how close an observation is to its own cluster and how far it is from the others.

Calculating the silhouette score
from sklearn.metrics import silhouette_score

score = silhouette_score(
    X_scaled,
    df["cluster"]
)

print("Silhouette Score:", score)

Such metrics give supporting information when choosing K, but they do not on their own decide the number of segments the business problem needs.

13. When Should K-Means Be Used?

ScenarioSuitability
Customer segmentationCan be used to group customers by similar behaviour.
Product groupingWorth considering to discover products with similar properties.
Behaviour analysisCan be used to separate users by activity patterns.
Very irregular cluster shapesOther clustering methods may fit better than K-Means.

Conclusion

K-Means is a strong algorithm for getting into clustering problems. The core logic is quite simple, but running the model is not enough for a good analysis.

You need to choose the right variables, scale them where necessary, decide on a suitable K, examine the result with methods such as the elbow method and the silhouette score, and finally interpret the clusters in a business context.

“K-Means gives you the clusters; saying what those clusters mean is the analyst's job.”

Especially in problems such as customer segmentation, this approach is a good starting point for turning raw data into more meaningful behaviour groups.

SQL ile Veri Analizi: Temel Sorgulardan Window Functions'a

Tableau Dashboard Tasarlarken Dikkat Edilmesi Gerekenler

Python ile Veri Temizleme: Pratik İpuçları

Power BI Dashboard Tasarlarken Dikkat Edilmesi Gerekenler

Veri Analistinin Bir Günlük Çalışma Akışı

Python ile veri görselleştirme yöntemleri

Python ile veri görselleştirme yöntemleri

Big Data (Büyük Veri) Nedir?