E-posta Kampanyası Gerçekten Satış Getirdi mi?Did the Email Campaign Really Drive Sales?
64.000 müşterinin rastgele gruplara ayrıldığı bir kampanya deneyinde etkiyi ölçtüm, kime işe yaradığına baktım ve bir hedefleme modelini en basit kuralla yarıştırdım.In a campaign experiment where 64,000 customers were split into random groups, I measured the effect, looked at who it worked for and pitted a targeting model against the simplest rule.
- Python
- A/B Test
- Uplift
- Scikit-learn
Bir kampanya raporu genellikle şunu söyler: “E-posta gönderilen müşteriler şu kadar harcadı.” Ama o müşterilerin bir kısmı e-posta olmasa da alışveriş yapacaktı. Bu çalışmada, müşterilerin rastgele gruplara ayrıldığı bir e-posta kampanyası deneyiyle üç soruya baktım: kampanya gerçekten satış getirdi mi, kime işe yaradı ve hedefleme modeli bir şey kazandırır mı?
Veri ve deney
Veri, Kevin Hillstrom'un 2008'de açık bir yarışma için paylaştığı bir ABD perakendecisine ait (MineThatData). Son bir yılda alışveriş yapmış 64.000 müşteri rastgele üç gruba ayrılmış: birine erkek ürünlerini, birine kadın ürünlerini tanıtan bir e-posta gönderilmiş, üçüncüsüne hiçbir şey gönderilmemiş. Sonraki iki haftada her müşterinin siteye girip girmediği, alışveriş yapıp yapmadığı ve ne kadar harcadığı ölçülmüş. Ayrıca son alışverişten bu yana geçen süre, geçen yılki harcama, daha önce erkek mi kadın mı ürünü aldığı, yeni müşteri olup olmadığı, ikamet bölgesi ve alışveriş kanalı biliniyor.
- Deneyi
denetle - Ortalama
etki - Atıf
yanılgısı - Kime
işe yaradı? - Hedefleme
ve kâr
Önce deneyin kendisi
Sonuca bakmadan önce deneyin düzgün kurulduğunu kontrol ettim. Grup büyüklükleri birbirine çok yakın (21.306, 21.307 ve 21.387); bu dağılımın tesadüfen çıkma olasılığı yüksek (p = 0,90). Gruplar müşteri özellikleri bakımından da benzer: hiçbir özellikte standartlaştırılmış fark 0,014'ü geçmiyor; 0,1'in altı genelde dengeli kabul edilir. Yani gruplar arasındaki fark e-postadan geliyor.
Kampanya satış getirdi mi?
| Grup | Müşteri | Siteyi ziyaret | Alışveriş | Müşteri başına harcama |
|---|---|---|---|---|
| E-posta yok | 21.306 | %10,62 | %0,57 | $0,65 |
| Erkek ürünleri e-postası | 21.307 | %18,28 | %1,25 | $1,42 |
| Kadın ürünleri e-postası | 21.387 | %15,14 | %0,88 | $1,08 |
| Etki (e-posta yok grubuna göre) | Erkek e-postası | Kadın e-postası |
|---|---|---|
| Ziyaret oranı | +7,66 puan+%72 · %95 aralık 7,00 – 8,32 | +4,52 puan+%43 · %95 aralık 3,89 – 5,16 |
| Alışveriş oranı | +0,68 puan+%119 · %95 aralık 0,50 – 0,86 | +0,31 puan+%54 · %95 aralık 0,15 – 0,47 |
| Müşteri başına harcama | +$0,77+%118 · %95 aralık $0,49 – $1,05 | +$0,42+%65 · %95 aralık $0,17 – $0,68 |
Evet, iki e-posta da işe yaradı. Erkek e-postası siteyi ziyaret edenlerin oranını %72, alışveriş yapanların oranını iki katından fazla artırdı ve müşteri başına $0,77 ek harcama getirdi. Kadın e-postasının etkisi daha küçük: müşteri başına $0,42.
Harcamanın hepsi e-postanın değil
E-posta almayan müşteriler de iki haftada kişi başı $0,65 harcadı. Erkek e-postası alanların $1,42'lik harcamasının yalnızca %54'ü e-postanın etkisi. Kampanya raporu e-posta alanların bütün harcamasını kampanyaya yazsaydı, etkiyi 1,8 kat büyük gösterirdi; kadın e-postasında 2,5 kat.
Bu, deney olmadan yapılan kampanya ölçümünün en yaygın hatası. Kontrol grubu olmadan “bu satışlar zaten olacak mıydı” sorusu cevaplanamaz.
Etki nereden geliyor?
Daha fazla müşteri alışveriş yaptı; alışveriş yapanlar daha fazla harcamadı. Erkek e-postası alıp alışveriş yapanların ortalama sepeti $113,53, e-posta almayanlarınki $114,00. Bu karşılaştırma yalnızca 267 ve 122 alıcıya dayandığı için küçük bir farkı göremez, ama sepet büyüdüğüne dair bir işaret yok.
Hangi e-posta daha iyi?
Erkek e-postası müşteri başına $0,35 daha fazla harcama getirdi (aralık $0,03 – $0,66, p = 0,03). Fark var, ama deney bunu ancak görebiliyor. Harcama bu kadar oynak olduğunda, bu büyüklükte bir farkı %80 olasılıkla yakalamak için grup başına yaklaşık 28.974 müşteri gerekir; bu deneyde grup başına 21.307 müşteri var. Aynı karşılaştırma tekrarlansa farkın çıkmaması şaşırtıcı olmaz.
Kime işe yaradı?
Müşterileri geçmişte aldıkları ürüne göre üçe ayırdım. Burada ziyaret oranına bakıyorum, çünkü harcama alt gruplarda çok gürültülü.
Nokta etkiyi, çizgi %95 aralığı gösteriyor.
- Kadın e-postası, yalnızca erkek ürünü almış müşterilerde neredeyse hiç işe yaramadı: ziyaret oranı +1,1 puan. Yalnızca kadın ürünü almışlarda etki +7,4 puan. Bu fark tesadüf olamayacak kadar büyük (p < 0,001).
- Erkek e-postası iki grupta da aynı ölçüde işe yaradı: +6,9 puan ve +7,1 puan.
Harcamada da aynı yönde bir fark var, ama aralığı sıfırı içeriyor. Diğer müşteri özelliklerinde tablo daha belirsiz:
| Müşteri grubu | Müşteri | Erkek e-postası | Kadın e-postası |
|---|---|---|---|
| EvetSon 12 ayda yeni müşteri | 32.144 | +$0,96$0,59 – $1,33 | +$0,73$0,40 – $1,06 |
| HayırSon 12 ayda yeni müşteri | 31.856 | +$0,58$0,14 – $1,01 | +$0,11−$0,28 – $0,50 |
| WebGeçmiş alışveriş kanalı | 28.217 | +$0,85$0,40 – $1,30 | +$0,42$0,02 – $0,81 |
| TelefonGeçmiş alışveriş kanalı | 28.021 | +$0,56$0,16 – $0,97 | +$0,23−$0,12 – $0,58 |
| İkisi deGeçmiş alışveriş kanalı | 7.762 | +$1,21$0,39 – $2,03 | +$1,16$0,26 – $2,06 |
| 1–3 aySon alışverişten bu yana | 22.393 | +$1,05$0,46 – $1,63 | +$0,36−$0,12 – $0,83 |
| 4–6 aySon alışverişten bu yana | 14.192 | +$0,49$0,01 – $0,97 | +$0,67$0,11 – $1,22 |
| 7–9 aySon alışverişten bu yana | 14.014 | +$0,61$0,06 – $1,15 | +$0,30−$0,25 – $0,85 |
| 10–12 aySon alışverişten bu yana | 13.401 | +$0,77$0,21 – $1,33 | +$0,41−$0,02 – $0,84 |
| <$100Geçen yılki harcama | 22.970 | +$0,54$0,11 – $0,96 | +$0,63$0,19 – $1,06 |
| $100–200Geçen yılki harcama | 14.254 | +$0,70$0,20 – $1,21 | +$0,46$0,04 – $0,87 |
| $200–500Geçen yılki harcama | 18.698 | +$0,89$0,27 – $1,50 | −$0,25−$0,70 – $0,21 |
| $500+Geçen yılki harcama | 8.078 | +$1,27$0,37 – $2,17 | +$1,35$0,39 – $2,31 |
Her hücrede etki ve %95 aralık. 26 karşılaştırma bir arada olduğu için birinin tesadüfen ‘anlamlı’ görünmesi beklenir.
Aralıkların neredeyse hepsi geniş ve birbirine biniyor. Bu verinin söylediği şu: erkek e-postası hemen herkeste aşağı yukarı aynı ölçüde işe yarıyor.
Hedefleme modeli bir şey kazandırır mı?
Kampanyalarda sık sorulan soru: “Herkese göndermek yerine, kime hangi e-postayı göndereceğimizi bir modelle seçsek?” Bunu denemek için her grup için ayrı bir gradyan artırma modeli kurdum (T-learner). Model her müşteri için üç seçeneğin sonucunu tahmin ediyor ve en iyisini seçiyor. Her müşteri, kendisini görmemiş modellerle puanlandı (5 katlı çapraz uydurma).
Gruplar rastgele olduğu için her politikanın sonucu doğrudan ölçülebiliyor: politikanın önerdiği e-postayı tesadüfen almış müşterilerin ortalamasına bakmak yeterli.
| Politika | Ziyaret oranı | Müşteri başına harcama |
|---|---|---|
| Kimseye gönderme | %10,62 | $0,65$0,50 – $0,82 |
| Herkese kadın e-postası | %15,14 | $1,08$0,87 – $1,28 |
| Herkese erkek e-postası | %18,28 | $1,42$1,17 – $1,64 |
| Kural: yalnızca kadın ürünü alanlara kadın, diğerlerine erkek | %18,42 | $1,43$1,19 – $1,66 |
| Model: ziyareti en çok artıracak e-posta | %18,12 | $1,39$1,15 – $1,61 |
| Model: harcamayı en çok artıracak e-posta | %17,33 | $1,32$1,08 – $1,52 |
- En iyi seçeneklerden biri en basit olanı: herkese erkek e-postası. Ürün kategorisine göre kural ondan ayırt edilemiyor (fark +$0,006, aralık −$0,18 – $0,20).
- Modeller daha iyi değil. Harcamayı tahmin eden model, ziyaret oranında herkese erkek e-postasından anlamlı biçimde kötü (−0,9 puan); gürültüyü örüntü sandı.
Müşteriler modelin tahmin ettiği etkiye göre sıralandı. Kesikli çizgi rastgele sıra.
Model ziyaret için rastgele sıradan biraz iyi, harcama için değil: müşterilerin %30'una gönderilseydi harcama etkisinin yalnızca %19'u yakalanırdı. Etki müşteriler arasında bu kadar düzgün dağıldığında, hedefleme modelinin ayıklayacak bir şeyi kalmıyor.
Kampanya kâr ettirir mi?
Veride e-posta maliyeti ve kâr marjı yok. Bu yüzden şu soruyu sordum: e-posta başına maliyet en fazla ne kadar olabilir ki kampanya zarar ettirmesin? Cevap, ek harcama çarpı brüt kâr marjı.
| Brüt kâr marjı varsayımı | Erkek e-postası | Kadın e-postası |
|---|---|---|
| %20 | $0,15$0,10 – $0,21 | $0,09$0,03 – $0,14 |
| %30 | $0,23$0,15 – $0,32 | $0,13$0,05 – $0,20 |
| %40 | $0,31$0,19 – $0,42 | $0,17$0,07 – $0,27 |
| %50 | $0,39$0,24 – $0,53 | $0,21$0,09 – $0,34 |
Örneğin marj %30 ise, erkek e-postası müşteri başına $0,23'ün altında bir maliyetle kâr ettirir. E-posta göndermenin doğrudan maliyeti genelde bunun çok altında; asıl maliyet, abonelikten çıkan ya da e-postaları önemsemez olan müşteriler olabilir. Bu deney iki haftayı kapsadığı için o etkiyi göremiyor.
Sınırlar
Sonuç
Kampanya işe yaradı, ama raporda görünecek rakamın yarısı kadar. Kontrol grubu olmasaydı, e-posta alan müşterilerin zaten yapacağı alışverişler de kampanyaya yazılacaktı.
İkinci ders hedeflemeyle ilgili. Karmaşık bir model, “herkese aynı e-postayı gönder” kuralını geçemedi; çünkü etki müşteriden müşteriye pek değişmiyordu. Modelden önce sorulacak soru, ayıklanacak bir fark olup olmadığı.
Veri: Kevin Hillstrom, MineThatData E-Mail Analytics and Data Mining Challenge (2008). Tutarlar ABD doları.
A campaign report usually says: “Customers who received the email spent this much.” But some of those customers would have bought even without the email. In this study I used an email campaign experiment, in which customers were split into random groups, to look at three questions: did the campaign really drive sales, who did it work for, and does a targeting model add anything?
Data and experiment
The data belongs to a US retailer and was shared by Kevin Hillstrom in 2008 for an open challenge (MineThatData). 64,000 customers who had bought in the last year were split at random into three groups: one received an email featuring men's products, one an email featuring women's products, and the third received nothing. Over the next two weeks, whether each customer visited the site, whether they bought and how much they spent were recorded. Also known: months since last purchase, last year's spend, whether they had bought men's or women's products before, whether they were a new customer, area type and purchase channel.
- Check the
experiment - Average
effect - Attribution
trap - Who did it
work for? - Targeting
and profit
First, the experiment itself
Before looking at results I checked that the experiment was set up properly. The group sizes are very close (21,306, 21,307 and 21,387); a split like this is very likely to arise by chance (p = 0.90). The groups are also alike in customer characteristics: no standardised difference exceeds 0.014, where below 0.1 is generally considered balanced. So differences between the groups come from the email.
Did the campaign drive sales?
| Group | Customers | Visited the site | Purchased | Spend per customer |
|---|---|---|---|---|
| No email | 21,306 | 10.62% | 0.57% | $0.65 |
| Men's products email | 21,307 | 18.28% | 1.25% | $1.42 |
| Women's products email | 21,387 | 15.14% | 0.88% | $1.08 |
| Effect (vs no email) | Men's email | Women's email |
|---|---|---|
| Visit rate | +7.66 pts+72% · 95% CI 7.00 – 8.32 | +4.52 pts+43% · 95% CI 3.89 – 5.16 |
| Purchase rate | +0.68 pts+119% · 95% CI 0.50 – 0.86 | +0.31 pts+54% · 95% CI 0.15 – 0.47 |
| Spend per customer | +$0.77+118% · 95% CI $0.49 – $1.05 | +$0.42+65% · 95% CI $0.17 – $0.68 |
Yes, both emails worked. The men's email raised the share of customers visiting the site by 72%, more than doubled the share who bought, and added $0.77 of spend per customer. The women's email had a smaller effect: $0.42 per customer.
Not all of the spend belongs to the email
Customers who received no email also spent $0.65 each over the two weeks. Of the $1.42 spent by men's email recipients, only 54% is the email's effect. A campaign report crediting all of the recipients' spend to the campaign would overstate the effect 1.8 times; 2.5 times for the women's email.
This is the most common error in campaign measurement without an experiment. Without a control group, “would these sales have happened anyway?” cannot be answered.
Where does the effect come from?
More customers bought; those who bought did not spend more. The average basket of buyers who received the men's email was $113.53, against $114.00 for those who received nothing. This comparison rests on only 267 and 122 buyers, so it cannot see a small difference, but there is no sign of bigger baskets.
Which email is better?
The men's email brought $0.35 more spend per customer (interval $0.03 – $0.66, p = 0.03). There is a difference, but the experiment can only just see it. With spend this volatile, detecting a difference of this size with 80% probability needs about 28,974 customers per group; this experiment has 21,307. If the comparison were repeated, it would not be surprising for the difference to disappear.
Who did it work for?
I split customers into three by the products they had bought before. Here I look at the visit rate, because spend is too noisy within subgroups.
The dot is the effect, the line the 95% interval.
- The women's email barely worked on customers who had bought only men's products: visit rate +1.1 pts. For customers who had bought only women's products the effect was +7.4 pts. The gap is too large to be chance (p < 0.001).
- The men's email worked equally well in both groups: +6.9 pts and +7.1 pts.
Spend shows a difference in the same direction, but its interval includes zero. For other customer characteristics the picture is less clear:
| Customer group | Customers | Men's email | Women's email |
|---|---|---|---|
| YesNew customer in the last 12 months | 32,144 | +$0.96$0.59 – $1.33 | +$0.73$0.40 – $1.06 |
| NoNew customer in the last 12 months | 31,856 | +$0.58$0.14 – $1.01 | +$0.11−$0.28 – $0.50 |
| WebPast purchase channel | 28,217 | +$0.85$0.40 – $1.30 | +$0.42$0.02 – $0.81 |
| PhonePast purchase channel | 28,021 | +$0.56$0.16 – $0.97 | +$0.23−$0.12 – $0.58 |
| BothPast purchase channel | 7,762 | +$1.21$0.39 – $2.03 | +$1.16$0.26 – $2.06 |
| 1–3 monthsTime since last purchase | 22,393 | +$1.05$0.46 – $1.63 | +$0.36−$0.12 – $0.83 |
| 4–6 monthsTime since last purchase | 14,192 | +$0.49$0.01 – $0.97 | +$0.67$0.11 – $1.22 |
| 7–9 monthsTime since last purchase | 14,014 | +$0.61$0.06 – $1.15 | +$0.30−$0.25 – $0.85 |
| 10–12 monthsTime since last purchase | 13,401 | +$0.77$0.21 – $1.33 | +$0.41−$0.02 – $0.84 |
| <$100Last year's spend | 22,970 | +$0.54$0.11 – $0.96 | +$0.63$0.19 – $1.06 |
| $100–200Last year's spend | 14,254 | +$0.70$0.20 – $1.21 | +$0.46$0.04 – $0.87 |
| $200–500Last year's spend | 18,698 | +$0.89$0.27 – $1.50 | −$0.25−$0.70 – $0.21 |
| $500+Last year's spend | 8,078 | +$1.27$0.37 – $2.17 | +$1.35$0.39 – $2.31 |
Each cell shows the effect and its 95% interval. With 26 comparisons side by side, one is expected to look ‘significant’ by chance.
Almost all intervals are wide and overlap. What this data says is that the men's email works to roughly the same degree for almost everyone.
Does a targeting model add anything?
A common question in campaigns: “Instead of sending to everyone, what if a model chose which email each customer gets?” To test this I built a separate gradient boosting model for each group (a T-learner). For each customer the model predicts the outcome of all three options and picks the best. Every customer was scored by models that had not seen them (5-fold cross-fitting).
Because the groups are random, each policy's result can be measured directly: it is the average of the customers who happened to receive the email the policy recommends.
| Policy | Visit rate | Spend per customer |
|---|---|---|
| Email nobody | 10.62% | $0.65$0.50 – $0.82 |
| Women's email to everyone | 15.14% | $1.08$0.87 – $1.28 |
| Men's email to everyone | 18.28% | $1.42$1.17 – $1.64 |
| Rule: women's email to women-only buyers, men's to the rest | 18.42% | $1.43$1.19 – $1.66 |
| Model: email that raises visits most | 18.12% | $1.39$1.15 – $1.61 |
| Model: email that raises spend most | 17.33% | $1.32$1.08 – $1.52 |
- One of the best options is the simplest: the men's email to everyone. The product-category rule cannot be told apart from it (difference +$0.006, interval −$0.18 – $0.20).
- The models are not better. The model that predicts spend is significantly worse on visits than the men's email for everyone (−0.9 pts); it took noise for pattern.
Customers ranked by the model's predicted effect. The dashed line is a random order.
For visits the model is slightly better than a random order, for spend it is not: mailing 30% of customers would capture only 19% of the spend effect. When the effect is spread this evenly across customers, a targeting model has nothing to sort out.
Is the campaign profitable?
The data has no email cost or profit margin. So I asked: what is the highest cost per email at which the campaign does not lose money? The answer is the extra spend times the gross margin.
| Assumed gross margin | Men's email | Women's email |
|---|---|---|
| 20% | $0.15$0.10 – $0.21 | $0.09$0.03 – $0.14 |
| 30% | $0.23$0.15 – $0.32 | $0.13$0.05 – $0.20 |
| 40% | $0.31$0.19 – $0.42 | $0.17$0.07 – $0.27 |
| 50% | $0.39$0.24 – $0.53 | $0.21$0.09 – $0.34 |
For example, at a 30% margin the men's email is profitable at any cost below $0.23 per customer. The direct cost of sending an email is usually far below that; the real cost may be customers who unsubscribe or start ignoring the emails. Because this experiment covers two weeks, it cannot see that.
Limitations
Conclusion
The campaign worked, but about half as well as a report would show. Without a control group, purchases the recipients would have made anyway would have been credited to the campaign.
The second lesson is about targeting. A complex model could not beat the rule “send everyone the same email”, because the effect hardly varied between customers. The question to ask before the model is whether there is any difference to sort out.
Data: Kevin Hillstrom, MineThatData E-Mail Analytics and Data Mining Challenge (2008). Amounts in US dollars.