Hangi Müşteriler Ertesi Yıl Geri Geliyor?Which Customers Come Back the Next Year?

Bir online perakendecinin iki yıllık satış verisinden RFM ve K-means ile dört müşteri segmenti çıkardım, sonra segmentleri ertesi yılın verisiyle sınadım.I built four customer segments with RFM and K-means from two years of an online retailer's sales, then tested the segments against the following year.

  • Python
  • RFM
  • K-means
  • scikit-learn

Bir müşteri listesine bakıp kimin ertesi yıl geri geleceğini söyleyebilir misiniz? Bu çalışmada bir online perakendecinin ilk yıl verisinden dört müşteri segmenti çıkardım, sonra bu segmentleri modelin hiç görmediği ikinci yılla sınadım.

İşlem satırı1.067.371Aralık 2009 – Aralık 2011
Segmentlenen müşteri4.229ilk yılda alışveriş yapanlar
Ertesi yıl geri gelen%64,2segmente göre %40 ile %96 arası

Veri ve temizlik

Veri, UCI'ın açık Online Retail II veri seti: İngiltere merkezli, müşterilerinin birçoğu toptancı olan bir online hediyelik eşya satıcısının iki yıllık tüm faturaları. Ham veri analize hazır değildi; dört adımda temizledim.

AdımKalan satırAtılan
Ham veri1.067.371
Birebir tekrar eden satırlar atıldı1.033.03634.335
Müşteri numarası olmayanlar atıldı797.885235.151
Ürün olmayan satırlar atıldı (kargo, düzeltme, masraf)794.2233.662
Fiyatı sıfır olanlar atıldı794.16360
  • İadeleri atmadım. 17.586 iade satırını silmek yerine müşterinin cirosundan düştüm (toplam £709.953). Silinseydi, büyük bir sipariş verip sonra iptal eden müşteri en değerli müşterilerden biri gibi görünürdü.
  • Müşterisiz satırlar dışarıda kaldı. Satırların %22,8 kadarında müşteri numarası yok; bunlar cironun %13,7 kadarı. Kime ait oldukları bilinmediği için segmentlenemiyorlar.
  • Çakışan günler. Dosyanın iki sayfası 1–9 Aralık 2010'u ikişer kez içeriyor. Birebir tekrar eden 34.335 satırın çoğu buradan geliyor.

Yöntem

Analiz adımları
  1. Veriyi
    temizle
  2. RFM
    hesapla
  3. K-means ile
    kümele
  4. Ertesi yılla
    sına

Her müşteri için üç değişken hesapladım: son alışverişten bu yana geçen gün (recency), fatura sayısı (frequency) ve iadeler düşülmüş net ciro (monetary). Üçü de sağa çarpık: müşterilerin %33,5 kadarı yalnızca bir sipariş vermiş, birkaç müşteri ise yüz binlerce sterlin harcamış. Bu yüzden logaritma alıp standartlaştırdım; aksi halde K-means birkaç dev müşteriye kilitleniyor.

Tasarım kararı: Segmentleri yalnızca ilk yılın verisiyle kurdum (1 Aralık 2009 – 30 Kasım 2010). İkinci yıl, segmentlerin gelecekteki davranışı ayırıp ayırmadığını görmek için kenarda bekledi.

Kaç küme?

Küme sayısını iki ölçüyle seçtim: siluet skoru ve kararlılık, yani veriyi yeniden örnekleyip modeli baştan kurduğumda müşterilerin aynı kümeye düşme derecesi.

Küme sayısıSiluetKararlılık (ARI)
20,410,97
30,320,90
4 (seçilen)0,340,93
50,330,85
60,300,84
70,280,60

En yüksek siluet iki kümede, ama o çözüm yalnızca aktif ve pasif müşteriyi ayırıyor; üzerine aksiyon kurmak zor. İkiden sonraki en iyi siluet ve yüksek kararlılık dört kümede, bu yüzden dördü seçtim. Siluetin 0,34 olması kümelerin keskin sınırlarla ayrılmadığını da gösteriyor: müşteriler bir süreklilik üzerinde duruyor, segmentler bu sürekliliği kullanışlı parçalara bölüyor. Yöntemin ayrıntıları için K-means yazısına bakabilirsiniz.

Dört segment

709 müşteri

Şampiyonlar

Sık ve yakın zamanda alışveriş yapan büyük müşteriler. Medyan 10 sipariş, £3.531 ciro, son alışveriş 8 gün önce.

1.184 müşteri

Düzenli

Yılda birkaç kez alışveriş yapan orta büyüklükte müşteriler. Medyan 4 sipariş, £1.286 ciro, son alışveriş 58 gün önce.

833 müşteri

Yeni ve küçük

Son aylarda gelmiş, henüz az harcamış müşteriler. Medyan 2 sipariş, £520 ciro, ilk alışverişten bu yana 63 gün.

1.503 müşteri

Uykudakiler

Çoğu tek sipariş verip aylardır dönmemiş. Medyan 1 sipariş, £257 ciro, son alışveriş 163 gün önce.

Müşterilerin %17'si cironun %64'ünü getiriyor

Segmentlerin ilk yıldaki müşteri ve ciro payı
  • Müşteri payı
  • Ciro payı
Şampiyonlar
%16,8
%63,9
Düzenli
%28,0
%24,2
Yeni ve küçük
%19,7
%6,5
Uykudakiler
%35,5
%5,5

Şampiyonlar müşterilerin altıda biri, cironun ise neredeyse üçte ikisi. Uykudakiler en kalabalık grup (%35,5), ama cirodaki payları yalnızca %5,5.

Segmentler ertesi yılı ayırıyor

Ertesi yıl yeniden alışveriş yapanların oranı
  • Şampiyonlar%96,2
  • Düzenli%74,0
  • Yeni ve küçük%66,7
  • Uykudakiler%40,1

Segmentler ilk yılın verisinden başka hiçbir şey görmedi; yine de ikinci yılda davranış net ayrışıyor. Şampiyonların %96'sı geri geldi ve müşteri başına ortalama £6.348 harcadı. Uykudakilerde geri gelme oranı %40, ortalama harcama £247.

SegmentGeri gelenlerMüşteri başına ortalama ciroİkinci yıl cirosundaki pay
Şampiyonlar%96,2£6.348%66,5
Düzenli%74,0£1.157%20,2
Yeni ve küçük%66,7£629%7,7
Uykudakiler%40,1£247%5,5
İki sürpriz: “Uykuda” dediğim müşterilerin %40'ı bir yıl içinde geri geldi; kayıp sayılmamalılar. “Yeni ve küçük” grubun ise üçte ikisi geri geldi.

Bu segmentlerle ne yapılır?

  • Şampiyonlar. Zaten geliyorlar. İndirim marjı eritir; öncelik onları kaybetmemek: öncelikli hizmet, erken erişim, stok güvencesi.
  • Düzenli. Dörtte biri ertesi yıl dönmedi. Sipariş aralığı uzadığında hatırlatma göndermeyi denemek için en uygun grup.
  • Yeni ve küçük. Üçte biri ikinci yıl hiç alışveriş yapmadı. Hedef, ikinci ve üçüncü siparişi hızlandırmak.
  • Uykudakiler. Düşük maliyetli yeniden kazanma kampanyaları. Bu grubun %40'ı zaten geri geldiği için kampanyanın etkisi mutlaka kontrol grubuyla ölçülmeli.

Sınırlar

Tek şirket, çok sayıda toptancıMüşterilerin birçoğu toptancı. Bireysel tüketici davranışına doğrudan genellenemez.
Müşterisiz satırlarCironun %13,7 kadarı müşteri numarası olmadığı için analiz dışında.
Kümeler keskin değilSiluet 0,34. Sınırdaki müşteriler iki segmente de yakın duruyor.
İkinci yıl biraz uzunDoğrulama dönemi 374 gün; verinin bittiği 9 Aralık 2011'e kadar gidiyor.
İsimler yorumModel yalnızca dört grup verir. “Şampiyon” ve “Uykuda” etiketlerini profillere bakarak ben koydum.
Öngörü, nedensellik değilSegmentler kimin geri geldiğini ayırıyor, neden geri geldiğini açıklamıyor.

Sonuç

Üç basit değişken ve dört küme, bir yıl sonraki davranışı ayırmaya yetti: en iyi segmentte her yüz müşteriden 96 tanesi geri gelirken en zayıf segmentte 40 tanesi geldi. Segmentasyonun değeri de burada: müşteri listesini, her biri için farklı bir karar gerektiren gruplara çeviriyor.

Benim için en öğretici kısım doğrulama oldu. Kümeleme her veride bir sonuç üretir; o sonucun işe yarayıp yaramadığını ancak modelin görmediği bir dönem gösterir.

Veri: Chen, D. (2019). Online Retail II. UCI Machine Learning Repository. Tutarlar sterlin cinsindendir.

Can you look at a customer list and tell who will come back next year? In this study I derived four customer segments from an online retailer's first year of data, then tested them against a second year the model never saw.

Transaction rows1,067,371December 2009 – December 2011
Customers segmented4,229those who bought in the first year
Came back the next year64.2%from 40% to 96% by segment

Data and cleaning

The data is UCI's open Online Retail II dataset: two years of invoices from a UK-based online gift retailer, many of whose customers are wholesalers. The raw data was not ready for analysis; I cleaned it in four steps.

StepRows leftRemoved
Raw data1,067,371
Exact duplicate rows removed1,033,03634,335
Rows without a customer ID removed797,885235,151
Non-product rows removed (postage, adjustments, fees)794,2233,662
Zero-price rows removed794,16360
  • Returns were kept. Instead of deleting 17,586 return rows, I deducted them from each customer's revenue (£709,953 in total). Had they been deleted, a customer who placed a huge order and then cancelled it would look like one of the most valuable.
  • Rows without a customer were left out. 22.8% of rows have no customer ID; they account for 13.7% of revenue. They cannot be segmented because nobody knows whose they are.
  • Overlapping days. The file's two sheets both contain December 1–9, 2010. Most of the 34,335 exact duplicate rows come from there.

Method

Analysis steps
  1. Clean
    the data
  2. Compute
    RFM
  3. Cluster with
    K-means
  4. Test against
    the next year

For each customer I computed three variables: days since the last purchase (recency), number of invoices (frequency) and net revenue after returns (monetary). All three are right-skewed: 33.5% of customers placed a single order, while a few spent hundreds of thousands of pounds. So I took logarithms and standardized them; otherwise K-means locks onto a handful of giant customers.

Design decision: The segments were built from the first year only (December 1, 2009 – November 30, 2010). The second year was held back to see whether the segments separate future behavior.

How many clusters?

I chose the number of clusters with two measures: the silhouette score and stability, meaning how consistently customers land in the same cluster when the data is resampled and the model rebuilt.

ClustersSilhouetteStability (ARI)
20.410.97
30.320.90
4 (chosen)0.340.93
50.330.85
60.300.84
70.280.60

The highest silhouette is at two clusters, but that solution only separates active from inactive customers, which is hard to act on. The best silhouette after two, together with high stability, is at four, so I chose four. A silhouette of 0.34 also shows the clusters are not sharply separated: customers sit on a continuum, and the segments cut that continuum into usable pieces. The K-means post covers the method in detail.

Four segments

709 customers

Champions

Large customers who buy often and bought recently. Median 10 orders, £3,531 revenue, last purchase 8 days ago.

1,184 customers

Regulars

Mid-sized customers who buy a few times a year. Median 4 orders, £1,286 revenue, last purchase 58 days ago.

833 customers

New and small

Customers who arrived in recent months and have spent little so far. Median 2 orders, £520 revenue, 63 days since first purchase.

1,503 customers

Dormant

Most placed a single order and have not returned for months. Median 1 order, £257 revenue, last purchase 163 days ago.

17% of customers bring 64% of revenue

Each segment's share of customers and revenue in the first year
  • Share of customers
  • Share of revenue
Champions
16.8%
63.9%
Regulars
28.0%
24.2%
New and small
19.7%
6.5%
Dormant
35.5%
5.5%

Champions are one in six customers and nearly two thirds of revenue. Dormant customers are the largest group (35.5%), yet they account for only 5.5% of revenue.

The segments separate the next year

Share of customers who bought again the next year
  • Champions96.2%
  • Regulars74.0%
  • New and small66.7%
  • Dormant40.1%

The segments saw nothing but first-year data, yet behavior in the second year separates clearly. 96% of Champions came back and spent £6,348 per customer on average. Among Dormant customers the return rate is 40% and the average spend £247.

SegmentCame backAverage revenue per customerShare of second-year revenue
Champions96.2%£6,34866.5%
Regulars74.0%£1,15720.2%
New and small66.7%£6297.7%
Dormant40.1%£2475.5%
Two surprises: 40% of the customers I labeled “Dormant” came back within a year; they should not be written off. And two thirds of the “New and small” group came back.

What to do with these segments

  • Champions. They come back anyway. Discounts erode margin; the priority is not losing them: priority service, early access, guaranteed stock.
  • Regulars. A quarter did not return the next year. The best group for testing reminders when the gap between orders grows.
  • New and small. A third bought nothing in the second year. The goal is to speed up the second and third order.
  • Dormant. Low-cost win-back campaigns. Because 40% come back anyway, any campaign's effect must be measured against a control group.

Limitations

One company, many wholesalersMany customers are wholesalers. The results do not transfer directly to individual consumers.
Rows without a customer13.7% of revenue is outside the analysis because it has no customer ID.
Clusters are not sharpSilhouette 0.34. Customers near a boundary sit close to two segments.
The second year runs longThe validation period is 374 days; it runs to December 9, 2011, where the data ends.
The names are interpretationThe model only returns four groups. I assigned “Champions” and “Dormant” by reading the profiles.
Prediction, not causationThe segments separate who comes back; they do not explain why.

Conclusion

Three simple variables and four clusters were enough to separate behavior a year later: 96 of every hundred customers in the best segment came back, against 40 in the weakest. That is where segmentation earns its keep: it turns a customer list into groups that each call for a different decision.

The most instructive part for me was the validation. Clustering produces a result on any data; only a period the model has not seen shows whether that result is useful.

Data: Chen, D. (2019). Online Retail II. UCI Machine Learning Repository. Amounts are in pounds sterling.