E-posta Kampanyası Gerçekten Satış Getirdi mi?Did the Email Campaign Really Drive Sales?

64.000 müşterinin rastgele gruplara ayrıldığı bir kampanya deneyinde etkiyi ölçtüm, kime işe yaradığına baktım ve bir hedefleme modelini en basit kuralla yarıştırdım.In a campaign experiment where 64,000 customers were split into random groups, I measured the effect, looked at who it worked for and pitted a targeting model against the simplest rule.

  • Python
  • A/B Test
  • Uplift
  • Scikit-learn

Bir kampanya raporu genellikle şunu söyler: “E-posta gönderilen müşteriler şu kadar harcadı.” Ama o müşterilerin bir kısmı e-posta olmasa da alışveriş yapacaktı. Bu çalışmada, müşterilerin rastgele gruplara ayrıldığı bir e-posta kampanyası deneyiyle üç soruya baktım: kampanya gerçekten satış getirdi mi, kime işe yaradı ve hedefleme modeli bir şey kazandırır mı?

Erkek e-postasının etkisi+$0,77müşteri başına harcama, +%118
Harcamanın e-postadan gelen kısmı%54gerisi e-posta olmadan da olacaktı
Hedefleme modelinin katkısıYokherkese aynı e-posta kadar iyi

Veri ve deney

Veri, Kevin Hillstrom'un 2008'de açık bir yarışma için paylaştığı bir ABD perakendecisine ait (MineThatData). Son bir yılda alışveriş yapmış 64.000 müşteri rastgele üç gruba ayrılmış: birine erkek ürünlerini, birine kadın ürünlerini tanıtan bir e-posta gönderilmiş, üçüncüsüne hiçbir şey gönderilmemiş. Sonraki iki haftada her müşterinin siteye girip girmediği, alışveriş yapıp yapmadığı ve ne kadar harcadığı ölçülmüş. Ayrıca son alışverişten bu yana geçen süre, geçen yılki harcama, daha önce erkek mi kadın mı ürünü aldığı, yeni müşteri olup olmadığı, ikamet bölgesi ve alışveriş kanalı biliniyor.

Çalışmanın adımları
  1. Deneyi
    denetle
  2. Ortalama
    etki
  3. Atıf
    yanılgısı
  4. Kime
    işe yaradı?
  5. Hedefleme
    ve kâr

Önce deneyin kendisi

Sonuca bakmadan önce deneyin düzgün kurulduğunu kontrol ettim. Grup büyüklükleri birbirine çok yakın (21.306, 21.307 ve 21.387); bu dağılımın tesadüfen çıkma olasılığı yüksek (p = 0,90). Gruplar müşteri özellikleri bakımından da benzer: hiçbir özellikte standartlaştırılmış fark 0,014'ü geçmiyor; 0,1'in altı genelde dengeli kabul edilir. Yani gruplar arasındaki fark e-postadan geliyor.

Kampanya satış getirdi mi?

GrupMüşteriSiteyi ziyaretAlışverişMüşteri başına harcama
E-posta yok21.306%10,62%0,57$0,65
Erkek ürünleri e-postası21.307%18,28%1,25$1,42
Kadın ürünleri e-postası21.387%15,14%0,88$1,08
Etki (e-posta yok grubuna göre)Erkek e-postasıKadın e-postası
Ziyaret oranı+7,66 puan+%72 · %95 aralık 7,00 – 8,32+4,52 puan+%43 · %95 aralık 3,89 – 5,16
Alışveriş oranı+0,68 puan+%119 · %95 aralık 0,50 – 0,86+0,31 puan+%54 · %95 aralık 0,15 – 0,47
Müşteri başına harcama+$0,77+%118 · %95 aralık $0,49 – $1,05+$0,42+%65 · %95 aralık $0,17 – $0,68

Evet, iki e-posta da işe yaradı. Erkek e-postası siteyi ziyaret edenlerin oranını %72, alışveriş yapanların oranını iki katından fazla artırdı ve müşteri başına $0,77 ek harcama getirdi. Kadın e-postasının etkisi daha küçük: müşteri başına $0,42.

Neden aralıkları da yazıyorum? 64.000 müşteriden yalnızca 578'i alışveriş yaptı. Harcama birkaç yüz kişiye dayandığı için belirsiz: erkek e-postasının etkisi müşteri başına $0,49 da olabilir, $1,05 da.

Harcamanın hepsi e-postanın değil

Müşteri başına harcama: e-posta olmadan da olacak kısım ve e-postanın getirdiği
  • E-posta yok$0,65$0,65
  • Erkek ürünleri e-postasıharcamanın %54'ü e-postadan$0,65+$0,77$1,42
  • Kadın ürünleri e-postasıharcamanın %39'u e-postadan$0,65+$0,42$1,08

E-posta almayan müşteriler de iki haftada kişi başı $0,65 harcadı. Erkek e-postası alanların $1,42'lik harcamasının yalnızca %54'ü e-postanın etkisi. Kampanya raporu e-posta alanların bütün harcamasını kampanyaya yazsaydı, etkiyi 1,8 kat büyük gösterirdi; kadın e-postasında 2,5 kat.

Bu, deney olmadan yapılan kampanya ölçümünün en yaygın hatası. Kontrol grubu olmadan “bu satışlar zaten olacak mıydı” sorusu cevaplanamaz.

Etki nereden geliyor?

Daha fazla müşteri alışveriş yaptı; alışveriş yapanlar daha fazla harcamadı. Erkek e-postası alıp alışveriş yapanların ortalama sepeti $113,53, e-posta almayanlarınki $114,00. Bu karşılaştırma yalnızca 267 ve 122 alıcıya dayandığı için küçük bir farkı göremez, ama sepet büyüdüğüne dair bir işaret yok.

Hangi e-posta daha iyi?

Erkek e-postası müşteri başına $0,35 daha fazla harcama getirdi (aralık $0,03 – $0,66, p = 0,03). Fark var, ama deney bunu ancak görebiliyor. Harcama bu kadar oynak olduğunda, bu büyüklükte bir farkı %80 olasılıkla yakalamak için grup başına yaklaşık 28.974 müşteri gerekir; bu deneyde grup başına 21.307 müşteri var. Aynı karşılaştırma tekrarlansa farkın çıkmaması şaşırtıcı olmaz.

Kime işe yaradı?

Müşterileri geçmişte aldıkları ürüne göre üçe ayırdım. Burada ziyaret oranına bakıyorum, çünkü harcama alt gruplarda çok gürültülü.

Ziyaret oranına etki (puan), geçmiş alışverişe göre
  • Yalnızca erkek ürünü almışerkek e-postası+6,9+5,9 – +7,9
  •  kadın e-postası+1,1+0,2 – +2,0
  • Yalnızca kadın ürünü almışerkek e-postası+7,1+6,1 – +8,0
  •  kadın e-postası+7,4+6,4 – +8,3
  • İkisini de almışerkek e-postası+13,4+10,9 – +16,0
  •  kadın e-postası+7,1+4,7 – +9,6

Nokta etkiyi, çizgi %95 aralığı gösteriyor.

  • Kadın e-postası, yalnızca erkek ürünü almış müşterilerde neredeyse hiç işe yaramadı: ziyaret oranı +1,1 puan. Yalnızca kadın ürünü almışlarda etki +7,4 puan. Bu fark tesadüf olamayacak kadar büyük (p < 0,001).
  • Erkek e-postası iki grupta da aynı ölçüde işe yaradı: +6,9 puan ve +7,1 puan.

Harcamada da aynı yönde bir fark var, ama aralığı sıfırı içeriyor. Diğer müşteri özelliklerinde tablo daha belirsiz:

Müşteri başına harcamaya etki, müşteri gruplarına göre
Müşteri grubuMüşteriErkek e-postasıKadın e-postası
EvetSon 12 ayda yeni müşteri32.144+$0,96$0,59 – $1,33+$0,73$0,40 – $1,06
HayırSon 12 ayda yeni müşteri31.856+$0,58$0,14 – $1,01+$0,11−$0,28 – $0,50
WebGeçmiş alışveriş kanalı28.217+$0,85$0,40 – $1,30+$0,42$0,02 – $0,81
TelefonGeçmiş alışveriş kanalı28.021+$0,56$0,16 – $0,97+$0,23−$0,12 – $0,58
İkisi deGeçmiş alışveriş kanalı7.762+$1,21$0,39 – $2,03+$1,16$0,26 – $2,06
1–3 aySon alışverişten bu yana22.393+$1,05$0,46 – $1,63+$0,36−$0,12 – $0,83
4–6 aySon alışverişten bu yana14.192+$0,49$0,01 – $0,97+$0,67$0,11 – $1,22
7–9 aySon alışverişten bu yana14.014+$0,61$0,06 – $1,15+$0,30−$0,25 – $0,85
10–12 aySon alışverişten bu yana13.401+$0,77$0,21 – $1,33+$0,41−$0,02 – $0,84
<$100Geçen yılki harcama22.970+$0,54$0,11 – $0,96+$0,63$0,19 – $1,06
$100–200Geçen yılki harcama14.254+$0,70$0,20 – $1,21+$0,46$0,04 – $0,87
$200–500Geçen yılki harcama18.698+$0,89$0,27 – $1,50−$0,25−$0,70 – $0,21
$500+Geçen yılki harcama8.078+$1,27$0,37 – $2,17+$1,35$0,39 – $2,31

Her hücrede etki ve %95 aralık. 26 karşılaştırma bir arada olduğu için birinin tesadüfen ‘anlamlı’ görünmesi beklenir.

Aralıkların neredeyse hepsi geniş ve birbirine biniyor. Bu verinin söylediği şu: erkek e-postası hemen herkeste aşağı yukarı aynı ölçüde işe yarıyor.

Hedefleme modeli bir şey kazandırır mı?

Kampanyalarda sık sorulan soru: “Herkese göndermek yerine, kime hangi e-postayı göndereceğimizi bir modelle seçsek?” Bunu denemek için her grup için ayrı bir gradyan artırma modeli kurdum (T-learner). Model her müşteri için üç seçeneğin sonucunu tahmin ediyor ve en iyisini seçiyor. Her müşteri, kendisini görmemiş modellerle puanlandı (5 katlı çapraz uydurma).

Gruplar rastgele olduğu için her politikanın sonucu doğrudan ölçülebiliyor: politikanın önerdiği e-postayı tesadüfen almış müşterilerin ortalamasına bakmak yeterli.

PolitikaZiyaret oranıMüşteri başına harcama
Kimseye gönderme%10,62$0,65$0,50 – $0,82
Herkese kadın e-postası%15,14$1,08$0,87 – $1,28
Herkese erkek e-postası%18,28$1,42$1,17 – $1,64
Kural: yalnızca kadın ürünü alanlara kadın, diğerlerine erkek%18,42$1,43$1,19 – $1,66
Model: ziyareti en çok artıracak e-posta%18,12$1,39$1,15 – $1,61
Model: harcamayı en çok artıracak e-posta%17,33$1,32$1,08 – $1,52
  • En iyi seçeneklerden biri en basit olanı: herkese erkek e-postası. Ürün kategorisine göre kural ondan ayırt edilemiyor (fark +$0,006, aralık −$0,18 – $0,20).
  • Modeller daha iyi değil. Harcamayı tahmin eden model, ziyaret oranında herkese erkek e-postasından anlamlı biçimde kötü (−0,9 puan); gürültüyü örüntü sandı.
Erkek e-postası, modelin önerdiği sırayla gönderilseydi: yakalanan etki
%0%25%50%75%100
%0%20%40%60%80%100

Müşteriler modelin tahmin ettiği etkiye göre sıralandı. Kesikli çizgi rastgele sıra.

Model ziyaret için rastgele sıradan biraz iyi, harcama için değil: müşterilerin %30'una gönderilseydi harcama etkisinin yalnızca %19'u yakalanırdı. Etki müşteriler arasında bu kadar düzgün dağıldığında, hedefleme modelinin ayıklayacak bir şeyi kalmıyor.

Ders: Hedefleme modeli kurmadan önce, etkinin müşteriden müşteriye gerçekten değişip değişmediğine bakmak gerekiyor. Değişmiyorsa en iyi hedefleme, herkese en iyi e-postayı göndermek.

Kampanya kâr ettirir mi?

Veride e-posta maliyeti ve kâr marjı yok. Bu yüzden şu soruyu sordum: e-posta başına maliyet en fazla ne kadar olabilir ki kampanya zarar ettirmesin? Cevap, ek harcama çarpı brüt kâr marjı.

Brüt kâr marjı varsayımıErkek e-postasıKadın e-postası
%20$0,15$0,10 – $0,21$0,09$0,03 – $0,14
%30$0,23$0,15 – $0,32$0,13$0,05 – $0,20
%40$0,31$0,19 – $0,42$0,17$0,07 – $0,27
%50$0,39$0,24 – $0,53$0,21$0,09 – $0,34

Örneğin marj %30 ise, erkek e-postası müşteri başına $0,23'ün altında bir maliyetle kâr ettirir. E-posta göndermenin doğrudan maliyeti genelde bunun çok altında; asıl maliyet, abonelikten çıkan ya da e-postaları önemsemez olan müşteriler olabilir. Bu deney iki haftayı kapsadığı için o etkiyi göremiyor.

Sınırlar

Eski ve tek bir şirket2008, ABD'de bir perakendeci. Oranlar bugünkü kampanyalara aynen taşınamaz; yöntem taşınır.
İki haftalık pencereUzun vadeli etki, abonelikten çıkma ve sonraki alışverişler görülmüyor.
Az alıcı578 alışveriş. Harcamaya dair her sonuç geniş aralıklı.
Maliyet yokKâr hesabı varsayılan marjlarla yapıldı.
Çoklu karşılaştırmaAlt grup sonuçları ipucu; önceden tek bir hipotez kurulmadan doğrulanmış sayılmaz.
E-postanın içeriği bilinmiyorErkek e-postasının neden herkeste işe yaradığını veri söylemiyor.

Sonuç

Kampanya işe yaradı, ama raporda görünecek rakamın yarısı kadar. Kontrol grubu olmasaydı, e-posta alan müşterilerin zaten yapacağı alışverişler de kampanyaya yazılacaktı.

İkinci ders hedeflemeyle ilgili. Karmaşık bir model, “herkese aynı e-postayı gönder” kuralını geçemedi; çünkü etki müşteriden müşteriye pek değişmiyordu. Modelden önce sorulacak soru, ayıklanacak bir fark olup olmadığı.

Veri: Kevin Hillstrom, MineThatData E-Mail Analytics and Data Mining Challenge (2008). Tutarlar ABD doları.

A campaign report usually says: “Customers who received the email spent this much.” But some of those customers would have bought even without the email. In this study I used an email campaign experiment, in which customers were split into random groups, to look at three questions: did the campaign really drive sales, who did it work for, and does a targeting model add anything?

Effect of the men's email+$0.77spend per customer, +118%
Share of spend caused by the email54%the rest would have happened anyway
Value of a targeting modelNoneno better than one email for everyone

Data and experiment

The data belongs to a US retailer and was shared by Kevin Hillstrom in 2008 for an open challenge (MineThatData). 64,000 customers who had bought in the last year were split at random into three groups: one received an email featuring men's products, one an email featuring women's products, and the third received nothing. Over the next two weeks, whether each customer visited the site, whether they bought and how much they spent were recorded. Also known: months since last purchase, last year's spend, whether they had bought men's or women's products before, whether they were a new customer, area type and purchase channel.

Steps of the study
  1. Check the
    experiment
  2. Average
    effect
  3. Attribution
    trap
  4. Who did it
    work for?
  5. Targeting
    and profit

First, the experiment itself

Before looking at results I checked that the experiment was set up properly. The group sizes are very close (21,306, 21,307 and 21,387); a split like this is very likely to arise by chance (p = 0.90). The groups are also alike in customer characteristics: no standardised difference exceeds 0.014, where below 0.1 is generally considered balanced. So differences between the groups come from the email.

Did the campaign drive sales?

GroupCustomersVisited the sitePurchasedSpend per customer
No email21,30610.62%0.57%$0.65
Men's products email21,30718.28%1.25%$1.42
Women's products email21,38715.14%0.88%$1.08
Effect (vs no email)Men's emailWomen's email
Visit rate+7.66 pts+72% · 95% CI 7.00 – 8.32+4.52 pts+43% · 95% CI 3.89 – 5.16
Purchase rate+0.68 pts+119% · 95% CI 0.50 – 0.86+0.31 pts+54% · 95% CI 0.15 – 0.47
Spend per customer+$0.77+118% · 95% CI $0.49 – $1.05+$0.42+65% · 95% CI $0.17 – $0.68

Yes, both emails worked. The men's email raised the share of customers visiting the site by 72%, more than doubled the share who bought, and added $0.77 of spend per customer. The women's email had a smaller effect: $0.42 per customer.

Why do I also give intervals? Of 64,000 customers only 578 bought anything. Because spend rests on a few hundred people it is uncertain: the men's email effect could be as low as $0.49 per customer or as high as $1.05.

Not all of the spend belongs to the email

Spend per customer: what would have happened anyway and what the email added
  • No email$0.65$0.65
  • Men's products email54% of spend from the email$0.65+$0.77$1.42
  • Women's products email39% of spend from the email$0.65+$0.42$1.08

Customers who received no email also spent $0.65 each over the two weeks. Of the $1.42 spent by men's email recipients, only 54% is the email's effect. A campaign report crediting all of the recipients' spend to the campaign would overstate the effect 1.8 times; 2.5 times for the women's email.

This is the most common error in campaign measurement without an experiment. Without a control group, “would these sales have happened anyway?” cannot be answered.

Where does the effect come from?

More customers bought; those who bought did not spend more. The average basket of buyers who received the men's email was $113.53, against $114.00 for those who received nothing. This comparison rests on only 267 and 122 buyers, so it cannot see a small difference, but there is no sign of bigger baskets.

Which email is better?

The men's email brought $0.35 more spend per customer (interval $0.03 – $0.66, p = 0.03). There is a difference, but the experiment can only just see it. With spend this volatile, detecting a difference of this size with 80% probability needs about 28,974 customers per group; this experiment has 21,307. If the comparison were repeated, it would not be surprising for the difference to disappear.

Who did it work for?

I split customers into three by the products they had bought before. Here I look at the visit rate, because spend is too noisy within subgroups.

Effect on visit rate (points) by purchase history
  • Bought men's products onlymen's email+6.9+5.9 – +7.9
  •  women's email+1.1+0.2 – +2.0
  • Bought women's products onlymen's email+7.1+6.1 – +8.0
  •  women's email+7.4+6.4 – +8.3
  • Bought bothmen's email+13.4+10.9 – +16.0
  •  women's email+7.1+4.7 – +9.6

The dot is the effect, the line the 95% interval.

  • The women's email barely worked on customers who had bought only men's products: visit rate +1.1 pts. For customers who had bought only women's products the effect was +7.4 pts. The gap is too large to be chance (p < 0.001).
  • The men's email worked equally well in both groups: +6.9 pts and +7.1 pts.

Spend shows a difference in the same direction, but its interval includes zero. For other customer characteristics the picture is less clear:

Effect on spend per customer by customer group
Customer groupCustomersMen's emailWomen's email
YesNew customer in the last 12 months32,144+$0.96$0.59 – $1.33+$0.73$0.40 – $1.06
NoNew customer in the last 12 months31,856+$0.58$0.14 – $1.01+$0.11−$0.28 – $0.50
WebPast purchase channel28,217+$0.85$0.40 – $1.30+$0.42$0.02 – $0.81
PhonePast purchase channel28,021+$0.56$0.16 – $0.97+$0.23−$0.12 – $0.58
BothPast purchase channel7,762+$1.21$0.39 – $2.03+$1.16$0.26 – $2.06
1–3 monthsTime since last purchase22,393+$1.05$0.46 – $1.63+$0.36−$0.12 – $0.83
4–6 monthsTime since last purchase14,192+$0.49$0.01 – $0.97+$0.67$0.11 – $1.22
7–9 monthsTime since last purchase14,014+$0.61$0.06 – $1.15+$0.30−$0.25 – $0.85
10–12 monthsTime since last purchase13,401+$0.77$0.21 – $1.33+$0.41−$0.02 – $0.84
<$100Last year's spend22,970+$0.54$0.11 – $0.96+$0.63$0.19 – $1.06
$100–200Last year's spend14,254+$0.70$0.20 – $1.21+$0.46$0.04 – $0.87
$200–500Last year's spend18,698+$0.89$0.27 – $1.50−$0.25−$0.70 – $0.21
$500+Last year's spend8,078+$1.27$0.37 – $2.17+$1.35$0.39 – $2.31

Each cell shows the effect and its 95% interval. With 26 comparisons side by side, one is expected to look ‘significant’ by chance.

Almost all intervals are wide and overlap. What this data says is that the men's email works to roughly the same degree for almost everyone.

Does a targeting model add anything?

A common question in campaigns: “Instead of sending to everyone, what if a model chose which email each customer gets?” To test this I built a separate gradient boosting model for each group (a T-learner). For each customer the model predicts the outcome of all three options and picks the best. Every customer was scored by models that had not seen them (5-fold cross-fitting).

Because the groups are random, each policy's result can be measured directly: it is the average of the customers who happened to receive the email the policy recommends.

PolicyVisit rateSpend per customer
Email nobody10.62%$0.65$0.50 – $0.82
Women's email to everyone15.14%$1.08$0.87 – $1.28
Men's email to everyone18.28%$1.42$1.17 – $1.64
Rule: women's email to women-only buyers, men's to the rest18.42%$1.43$1.19 – $1.66
Model: email that raises visits most18.12%$1.39$1.15 – $1.61
Model: email that raises spend most17.33%$1.32$1.08 – $1.52
  • One of the best options is the simplest: the men's email to everyone. The product-category rule cannot be told apart from it (difference +$0.006, interval −$0.18 – $0.20).
  • The models are not better. The model that predicts spend is significantly worse on visits than the men's email for everyone (−0.9 pts); it took noise for pattern.
If the men's email were sent in the order the model suggests: share of the effect captured
0%25%50%75%100%
0%20%40%60%80%100%

Customers ranked by the model's predicted effect. The dashed line is a random order.

For visits the model is slightly better than a random order, for spend it is not: mailing 30% of customers would capture only 19% of the spend effect. When the effect is spread this evenly across customers, a targeting model has nothing to sort out.

Lesson: Before building a targeting model, check whether the effect actually varies from customer to customer. If it does not, the best targeting is sending everyone the best email.

Is the campaign profitable?

The data has no email cost or profit margin. So I asked: what is the highest cost per email at which the campaign does not lose money? The answer is the extra spend times the gross margin.

Assumed gross marginMen's emailWomen's email
20%$0.15$0.10 – $0.21$0.09$0.03 – $0.14
30%$0.23$0.15 – $0.32$0.13$0.05 – $0.20
40%$0.31$0.19 – $0.42$0.17$0.07 – $0.27
50%$0.39$0.24 – $0.53$0.21$0.09 – $0.34

For example, at a 30% margin the men's email is profitable at any cost below $0.23 per customer. The direct cost of sending an email is usually far below that; the real cost may be customers who unsubscribe or start ignoring the emails. Because this experiment covers two weeks, it cannot see that.

Limitations

Old and a single company2008, a US retailer. The rates do not carry over to today's campaigns; the method does.
A two-week windowLong-term effects, unsubscribes and later purchases are not visible.
Few buyers578 purchases. Every result about spend has a wide interval.
No costsProfit was computed with assumed margins.
Multiple comparisonsSubgroup results are hints; without a single hypothesis set in advance they are not confirmed.
Email content unknownThe data does not say why the men's email worked for everyone.

Conclusion

The campaign worked, but about half as well as a report would show. Without a control group, purchases the recipients would have made anyway would have been credited to the campaign.

The second lesson is about targeting. A complex model could not beat the rule “send everyone the same email”, because the effect hardly varied between customers. The question to ask before the model is whether there is any difference to sort out.

Data: Kevin Hillstrom, MineThatData E-Mail Analytics and Data Mining Challenge (2008). Amounts in US dollars.