Hangi Müşteriler Ertesi Yıl Geri Geliyor?Which Customers Come Back the Next Year?
Bir online perakendecinin iki yıllık satış verisinden RFM ve K-means ile dört müşteri segmenti çıkardım, sonra segmentleri ertesi yılın verisiyle sınadım.I built four customer segments with RFM and K-means from two years of an online retailer's sales, then tested the segments against the following year.
- Python
- RFM
- K-means
- scikit-learn
Bir müşteri listesine bakıp kimin ertesi yıl geri geleceğini söyleyebilir misiniz? Bu çalışmada bir online perakendecinin ilk yıl verisinden dört müşteri segmenti çıkardım, sonra bu segmentleri modelin hiç görmediği ikinci yılla sınadım.
Veri ve temizlik
Veri, UCI'ın açık Online Retail II veri seti: İngiltere merkezli, müşterilerinin birçoğu toptancı olan bir online hediyelik eşya satıcısının iki yıllık tüm faturaları. Ham veri analize hazır değildi; dört adımda temizledim.
| Adım | Kalan satır | Atılan |
|---|---|---|
| Ham veri | 1.067.371 | |
| Birebir tekrar eden satırlar atıldı | 1.033.036 | 34.335 |
| Müşteri numarası olmayanlar atıldı | 797.885 | 235.151 |
| Ürün olmayan satırlar atıldı (kargo, düzeltme, masraf) | 794.223 | 3.662 |
| Fiyatı sıfır olanlar atıldı | 794.163 | 60 |
- İadeleri atmadım. 17.586 iade satırını silmek yerine müşterinin cirosundan düştüm (toplam £709.953). Silinseydi, büyük bir sipariş verip sonra iptal eden müşteri en değerli müşterilerden biri gibi görünürdü.
- Müşterisiz satırlar dışarıda kaldı. Satırların %22,8 kadarında müşteri numarası yok; bunlar cironun %13,7 kadarı. Kime ait oldukları bilinmediği için segmentlenemiyorlar.
- Çakışan günler. Dosyanın iki sayfası 1–9 Aralık 2010'u ikişer kez içeriyor. Birebir tekrar eden 34.335 satırın çoğu buradan geliyor.
Yöntem
- Veriyi
temizle - RFM
hesapla - K-means ile
kümele - Ertesi yılla
sına
Her müşteri için üç değişken hesapladım: son alışverişten bu yana geçen gün (recency), fatura sayısı (frequency) ve iadeler düşülmüş net ciro (monetary). Üçü de sağa çarpık: müşterilerin %33,5 kadarı yalnızca bir sipariş vermiş, birkaç müşteri ise yüz binlerce sterlin harcamış. Bu yüzden logaritma alıp standartlaştırdım; aksi halde K-means birkaç dev müşteriye kilitleniyor.
Kaç küme?
Küme sayısını iki ölçüyle seçtim: siluet skoru ve kararlılık, yani veriyi yeniden örnekleyip modeli baştan kurduğumda müşterilerin aynı kümeye düşme derecesi.
| Küme sayısı | Siluet | Kararlılık (ARI) |
|---|---|---|
| 2 | 0,41 | 0,97 |
| 3 | 0,32 | 0,90 |
| 4 (seçilen) | 0,34 | 0,93 |
| 5 | 0,33 | 0,85 |
| 6 | 0,30 | 0,84 |
| 7 | 0,28 | 0,60 |
En yüksek siluet iki kümede, ama o çözüm yalnızca aktif ve pasif müşteriyi ayırıyor; üzerine aksiyon kurmak zor. İkiden sonraki en iyi siluet ve yüksek kararlılık dört kümede, bu yüzden dördü seçtim. Siluetin 0,34 olması kümelerin keskin sınırlarla ayrılmadığını da gösteriyor: müşteriler bir süreklilik üzerinde duruyor, segmentler bu sürekliliği kullanışlı parçalara bölüyor. Yöntemin ayrıntıları için K-means yazısına bakabilirsiniz.
Dört segment
Şampiyonlar
Sık ve yakın zamanda alışveriş yapan büyük müşteriler. Medyan 10 sipariş, £3.531 ciro, son alışveriş 8 gün önce.
Düzenli
Yılda birkaç kez alışveriş yapan orta büyüklükte müşteriler. Medyan 4 sipariş, £1.286 ciro, son alışveriş 58 gün önce.
Yeni ve küçük
Son aylarda gelmiş, henüz az harcamış müşteriler. Medyan 2 sipariş, £520 ciro, ilk alışverişten bu yana 63 gün.
Uykudakiler
Çoğu tek sipariş verip aylardır dönmemiş. Medyan 1 sipariş, £257 ciro, son alışveriş 163 gün önce.
Müşterilerin %17'si cironun %64'ünü getiriyor
- Müşteri payı
- Ciro payı
Şampiyonlar müşterilerin altıda biri, cironun ise neredeyse üçte ikisi. Uykudakiler en kalabalık grup (%35,5), ama cirodaki payları yalnızca %5,5.
Segmentler ertesi yılı ayırıyor
Segmentler ilk yılın verisinden başka hiçbir şey görmedi; yine de ikinci yılda davranış net ayrışıyor. Şampiyonların %96'sı geri geldi ve müşteri başına ortalama £6.348 harcadı. Uykudakilerde geri gelme oranı %40, ortalama harcama £247.
| Segment | Geri gelenler | Müşteri başına ortalama ciro | İkinci yıl cirosundaki pay |
|---|---|---|---|
| Şampiyonlar | %96,2 | £6.348 | %66,5 |
| Düzenli | %74,0 | £1.157 | %20,2 |
| Yeni ve küçük | %66,7 | £629 | %7,7 |
| Uykudakiler | %40,1 | £247 | %5,5 |
Bu segmentlerle ne yapılır?
- Şampiyonlar. Zaten geliyorlar. İndirim marjı eritir; öncelik onları kaybetmemek: öncelikli hizmet, erken erişim, stok güvencesi.
- Düzenli. Dörtte biri ertesi yıl dönmedi. Sipariş aralığı uzadığında hatırlatma göndermeyi denemek için en uygun grup.
- Yeni ve küçük. Üçte biri ikinci yıl hiç alışveriş yapmadı. Hedef, ikinci ve üçüncü siparişi hızlandırmak.
- Uykudakiler. Düşük maliyetli yeniden kazanma kampanyaları. Bu grubun %40'ı zaten geri geldiği için kampanyanın etkisi mutlaka kontrol grubuyla ölçülmeli.
Sınırlar
Sonuç
Üç basit değişken ve dört küme, bir yıl sonraki davranışı ayırmaya yetti: en iyi segmentte her yüz müşteriden 96 tanesi geri gelirken en zayıf segmentte 40 tanesi geldi. Segmentasyonun değeri de burada: müşteri listesini, her biri için farklı bir karar gerektiren gruplara çeviriyor.
Benim için en öğretici kısım doğrulama oldu. Kümeleme her veride bir sonuç üretir; o sonucun işe yarayıp yaramadığını ancak modelin görmediği bir dönem gösterir.
Veri: Chen, D. (2019). Online Retail II. UCI Machine Learning Repository. Tutarlar sterlin cinsindendir.
Can you look at a customer list and tell who will come back next year? In this study I derived four customer segments from an online retailer's first year of data, then tested them against a second year the model never saw.
Data and cleaning
The data is UCI's open Online Retail II dataset: two years of invoices from a UK-based online gift retailer, many of whose customers are wholesalers. The raw data was not ready for analysis; I cleaned it in four steps.
| Step | Rows left | Removed |
|---|---|---|
| Raw data | 1,067,371 | |
| Exact duplicate rows removed | 1,033,036 | 34,335 |
| Rows without a customer ID removed | 797,885 | 235,151 |
| Non-product rows removed (postage, adjustments, fees) | 794,223 | 3,662 |
| Zero-price rows removed | 794,163 | 60 |
- Returns were kept. Instead of deleting 17,586 return rows, I deducted them from each customer's revenue (£709,953 in total). Had they been deleted, a customer who placed a huge order and then cancelled it would look like one of the most valuable.
- Rows without a customer were left out. 22.8% of rows have no customer ID; they account for 13.7% of revenue. They cannot be segmented because nobody knows whose they are.
- Overlapping days. The file's two sheets both contain December 1–9, 2010. Most of the 34,335 exact duplicate rows come from there.
Method
- Clean
the data - Compute
RFM - Cluster with
K-means - Test against
the next year
For each customer I computed three variables: days since the last purchase (recency), number of invoices (frequency) and net revenue after returns (monetary). All three are right-skewed: 33.5% of customers placed a single order, while a few spent hundreds of thousands of pounds. So I took logarithms and standardized them; otherwise K-means locks onto a handful of giant customers.
How many clusters?
I chose the number of clusters with two measures: the silhouette score and stability, meaning how consistently customers land in the same cluster when the data is resampled and the model rebuilt.
| Clusters | Silhouette | Stability (ARI) |
|---|---|---|
| 2 | 0.41 | 0.97 |
| 3 | 0.32 | 0.90 |
| 4 (chosen) | 0.34 | 0.93 |
| 5 | 0.33 | 0.85 |
| 6 | 0.30 | 0.84 |
| 7 | 0.28 | 0.60 |
The highest silhouette is at two clusters, but that solution only separates active from inactive customers, which is hard to act on. The best silhouette after two, together with high stability, is at four, so I chose four. A silhouette of 0.34 also shows the clusters are not sharply separated: customers sit on a continuum, and the segments cut that continuum into usable pieces. The K-means post covers the method in detail.
Four segments
Champions
Large customers who buy often and bought recently. Median 10 orders, £3,531 revenue, last purchase 8 days ago.
Regulars
Mid-sized customers who buy a few times a year. Median 4 orders, £1,286 revenue, last purchase 58 days ago.
New and small
Customers who arrived in recent months and have spent little so far. Median 2 orders, £520 revenue, 63 days since first purchase.
Dormant
Most placed a single order and have not returned for months. Median 1 order, £257 revenue, last purchase 163 days ago.
17% of customers bring 64% of revenue
- Share of customers
- Share of revenue
Champions are one in six customers and nearly two thirds of revenue. Dormant customers are the largest group (35.5%), yet they account for only 5.5% of revenue.
The segments separate the next year
The segments saw nothing but first-year data, yet behavior in the second year separates clearly. 96% of Champions came back and spent £6,348 per customer on average. Among Dormant customers the return rate is 40% and the average spend £247.
| Segment | Came back | Average revenue per customer | Share of second-year revenue |
|---|---|---|---|
| Champions | 96.2% | £6,348 | 66.5% |
| Regulars | 74.0% | £1,157 | 20.2% |
| New and small | 66.7% | £629 | 7.7% |
| Dormant | 40.1% | £247 | 5.5% |
What to do with these segments
- Champions. They come back anyway. Discounts erode margin; the priority is not losing them: priority service, early access, guaranteed stock.
- Regulars. A quarter did not return the next year. The best group for testing reminders when the gap between orders grows.
- New and small. A third bought nothing in the second year. The goal is to speed up the second and third order.
- Dormant. Low-cost win-back campaigns. Because 40% come back anyway, any campaign's effect must be measured against a control group.
Limitations
Conclusion
Three simple variables and four clusters were enough to separate behavior a year later: 96 of every hundred customers in the best segment came back, against 40 in the weakest. That is where segmentation earns its keep: it turns a customer list into groups that each call for a different decision.
The most instructive part for me was the validation. Clustering produces a result on any data; only a period the model has not seen shows whether that result is useful.
Data: Chen, D. (2019). Online Retail II. UCI Machine Learning Repository. Amounts are in pounds sterling.