A/B Testi Nedir? Veri Odaklı Karar Verme

Hipotez kurmaktan kontrol ve test gruplarına, dönüşüm oranından istatistiksel anlamlılığa kadar bir A/B testinin tüm adımlarını örnekle inceliyorum.

  • İstatistik
  • Araçlar

“Hangisi daha iyi?” sorusunu tahminle değil, kontrollü bir deneyle yanıtlamak. A/B testi; web sitelerinden reklam kampanyalarına, e-posta başlıklarından ürün akışlarına kadar birçok alanda kullanıcı davranışını ölçerek daha bilinçli kararlar vermeyi sağlıyor.

Bu yazıda hipotez kurmaktan kontrol/test gruplarına, dönüşüm oranından istatistiksel anlamlılığa kadar bir A/B testinin tüm adımlarını gerçek hayata yakın bir e-ticaret örneğiyle ele alıyorum.

“İyi bir deney yalnızca ‘hangi grup kazandı?' sorusuna cevap vermez; farkın büyüklüğünü, belirsizliğini ve iş açısından değerini de ortaya koyar.”

1. A/B Testi Nedir?

A/B testi, bir ürünün, sayfanın, reklamın veya mesajın iki farklı versiyonunu kullanıcı gruplarına göstererek hangi versiyonun hedeflenen davranışı daha iyi gerçekleştirdiğini ölçen kontrollü deney yöntemidir.

En basit senaryoda kullanıcıların bir bölümü A versiyonunu (kontrol), diğer bölümü B versiyonunu (test) görür. Sonra satın alma, kayıt olma, form gönderme veya reklama tıklama gibi önceden belirlenmiş bir hedef karşılaştırılır.

A

Kontrol

Mevcut durumda kullanılan versiyon; karşılaştırmanın referans noktası.

B

Test

Değişiklik içeren versiyon; etkisi kontrol grubuna göre ölçülür.

KPI

Hedef metrik

Sonucu belirleyen ana ölçü: dönüşüm, tıklama, gelir veya başka bir gösterge.

A/B testi süreci
  1. Hipotez
    Oluştur
  2. Testi
    Kur
  3. Veriyi
    Topla
  4. Analiz
    Et
  5. Karar
    Ver
Temel mantık: Aynı hedefe hizmet eden iki seçeneği mümkün olduğunca benzer koşullarda karşılaştırmak ve gözlenen farkın gerçek bir etkiden mi kaynaklandığını istatistikle değerlendirmek.

2. Veri Odaklı Karar Vermede Neden Önemli?

Pazarlama ve ürün kararlarında “bu tasarım daha güzel”, “bu başlık kesin daha çok tık alır” gibi sezgisel yorumlar faydalı başlangıç noktaları olabilir; ama tek başına kanıt değildir. A/B testi bu varsayımları ölçülebilir hipotezlere dönüştürüyor.

Böylece ekipler “B daha iyi görünüyor” demek yerine “B versiyonu, belirli koşullarda dönüşüm oranını ne ölçüde değiştiriyor ve bu fark istatistiksel olarak güvenilir mi?” sorusunu sorabiliyor.

3. Hipotez Nasıl Kurulur?

İyi bir deney ölçülebilir bir hipotezle başlar. Hipotez yalnızca yapılacak değişikliği değil, beklenen sonucu da açıklar.

Hipotezİfade
H₀ (sıfır hipotezi)A ve B versiyonlarının hedef metrik üzerinde anlamlı bir farkı yoktur.
H₁ (alternatif hipotez)B versiyonunun hedef metrik üzerinde A'dan farklı bir etkisi vardır.
Örnek hipotez: “Sepete ekle” butonunun daha görünür bir konuma alınması, ürün sayfasındaki satın alma dönüşüm oranını artıracaktır.

Hipotezin ölçülebilir olması şart. “Sayfa daha iyi olacak” yerine “satın alma dönüşüm oranı artacak” denildiğinde test edilecek net bir çıktı ortaya çıkıyor.

4. Kontrol ve Test Grupları

Kullanıcıların gruplara mümkün olduğunca rastgele atanması önemli. Amaç, iki grubun yaş, cihaz, trafik kaynağı ve önceki davranış gibi özellikler açısından gereksiz biçimde ayrışmasını önlemek.

A • Kontrol

Mevcut deneyim

Bugün kullanılan sayfa veya kampanya; yeni versiyon bununla karşılaştırılır.

B • Test

Değiştirilmiş deneyim

Hipotezi sınamak için belirli bir değişkenin farklılaştırıldığı versiyon.

Tasarım

Tek değişken

Yalnızca başlık değiştiyse, sonucun hangi değişiklikten geldiğini yorumlamak kolaydır.

Örneklem

Yeterli hacim

Çok az kullanıcıyla yapılan testlerde küçük farklar güvenilir görünmez.

5. Dönüşüm Oranı Nasıl Hesaplanır?

A/B testlerinde en sık kullanılan göstergelerden biri dönüşüm oranı. Hedef davranışı gerçekleştiren kullanıcıların toplam kullanıcıya oranını ifade eder.

CR=CN×100CR = \frac{C}{N} \times 100

CR: dönüşüm oranı, C: dönüşüm sayısı, N: toplam kullanıcı sayısı.

GrupKullanıcıDönüşümDönüşüm oranı
A • Kontrol100.0005.800%5,8
B • Test100.0007.100%7,1

Göreceli artış nasıl bulunur?

Lift=CRB−CRACRA×100=7.1−5.85.8×100≈22.4Lift = \frac{CR_B - CR_A}{CR_A} \times 100 = \frac{7.1 - 5.8}{5.8} \times 100 \approx 22.4

Bu örnekte B versiyonunun dönüşüm oranı yaklaşık %22,4 daha yüksek görünüyor. Fakat burada önemli bir ayrım var: gözlenen fark ile istatistiksel olarak güvenilir fark aynı şey değil.

6. İstatistiksel Anlamlılık Nedir?

Bir deneyde B'nin A'dan yüksek sonuç vermesi, tek başına B'nin gerçekten daha iyi olduğunu kanıtlamaz. Aradaki fark örneklem kaynaklı rastlantısal dalgalanmanın sonucu olabilir. İstatistiksel anlamlılık, gözlenen farkın yalnızca şans eseri ortaya çıkma ihtimalini değerlendirmeye yardımcı olur.

p

p-değeri

Test edilen varsayım altında gözlenen veya daha uç bir farkın ne kadar sıra dışı olduğunu değerlendirir. Tek başına “başarı yüzdesi” değildir.

CI

Güven aralığı

Etkinin ne büyüklükte olabileceğine dair aralık sunar; “anlamlı/anlamsız” demekten çok daha açıklayıcıdır.

Önemli: Yaygın eşik olan 0,05 birçok analizde karar için kullanılır; ancak testin tasarımı, örneklem büyüklüğü, çoklu karşılaştırmalar ve iş etkisi birlikte değerlendirilmeli.

İstatistiksel anlamlılık ≠ iş açısından anlamlılık

B versiyonu istatistiksel olarak anlamlı şekilde %1 daha iyi olabilir. Fakat bu iyileştirmenin maliyeti, kullanıcı deneyimine etkisi veya gelir katkısı çok düşükse iş kararı farklı olabilir. İyi bir A/B testi hem istatistiksel hem de iş açısından değerlendirilir.

7. Gerçek Hayata Yakın Bir Örnek

Bir e-ticaret sitesinde “Sepete Ekle” butonunun yeterince görünür olmadığı düşünülüyor. Pazarlama ekibi, butonu daha belirgin hale getiren yeni bir tasarımın satın alma oranını artıracağını öngörüyor.

Deney tasarımı

AlanKarar
HipotezDaha görünür buton, satın alma dönüşümünü artırır.
KontrolMevcut ürün sayfası
TestYeni buton konumu ve tasarımı
Ana metrikSatın alma dönüşüm oranı
İkincil metrikSepete ekleme oranı ve kullanıcı başına gelir
Dağılım%50 A / %50 B

Örnek sonuç

Mutlak

+1,3 puan

Dönüşüm oranındaki mutlak artış (%5,8 → %7,1).

Göreceli

+%22,4

Kontrol grubuna göre göreceli iyileşme.

Destek

2 ek metrik

Sepete ekleme ve gelir etkisi kararı destekler.

Son aşamada uygun istatistiksel test uygulanır ve farkın güvenilirliği incelenir. Sonuç, önceden belirlenen analiz planına göre anlamlı bulunursa B versiyonunun yayına alınması düşünülebilir. Ancak karar yalnızca p-değerine değil, gelir ve kullanıcı deneyimi göstergelerine de bakılarak verilmeli.

Python ile temel dönüşüm hesabı

Dönüşüm oranı ve lift
control_users = 100000
control_conversions = 5800

test_users = 100000
test_conversions = 7100

control_rate = control_conversions / control_users
test_rate = test_conversions / test_users

print("Kontrol:", control_rate)
print("Test:", test_rate)

lift = (test_rate - control_rate) / control_rate
print("Goreceli artis:", lift)

Farkın istatistiksel olarak anlamlı olup olmadığını değerlendirmek için iki oran arasındaki farkı test eden yöntemler kullanılır:

İki oran için z-testi
from statsmodels.stats.proportion import proportions_ztest

# dönüşümler ve kullanıcı sayıları
donusumler = [5800, 7100]
kullanicilar = [100000, 100000]

z, p = proportions_ztest(donusumler, kullanicilar)
print("z:", z, "p:", p)

8. Sonuçları Nasıl Yorumlamalı?

A/B testi sonuçlarını okurken tek bir yüzdeye odaklanmak yerine deneyin tamamını değerlendirmek gerekiyor. Şu dört soru iyi bir başlangıç:

1. Soru

Fark ne kadar?

Mutlak ve göreceli etkiyi ayrı ayrı hesapla.

2. Soru

Güvenilir mi?

İstatistiksel belirsizliği ve güven aralığını incele.

3. Soru

İşe etkisi ne?

Gelir, maliyet, elde tutma veya diğer göstergeleri değerlendir.

4. Soru

Deney doğru mu?

Gruplar, süre, trafik ve ölçüm kalitesi uygun mu kontrol et.

Örnek karar cümlesi: “B versiyonu %22,4 lift sağladı” demek yerine; “B versiyonu kontrol grubuna göre %22,4 daha yüksek dönüşüm gösterdi; sonuç planlanan analiz kapsamında istatistiksel olarak anlamlıysa ve gelir göstergeleri de olumluysa B tercih edilebilir” demek daha sağlıklı.

9. Sık Yapılan Hatalar

HataNeden problem?Daha iyi yaklaşım
Sonuca çok erken bakmakRastlantısal dalgalanmalar yanıltır.Önceden belirlenen deney planına ve örneklem ihtiyacına uy.
Aynı anda çok şey değiştirmekHangi değişkenin etki yarattığı belirsizleşir.Hipoteze bağlı net bir değişiklik yap.
Sadece p-değerine odaklanmakEtkinin büyüklüğü ve belirsizlik gözden kaçar.Etki büyüklüğü, güven aralığı ve iş metriklerini birlikte değerlendir.
Yanlış metrik seçmekKolay ölçülen bir metrik gerçek hedefi temsil etmeyebilir.Ana başarı metriğini iş hedefiyle ilişkilendir.
Grupların dengesiz olmasıKarşılaştırmanın adilliği bozulur.Rastgele atama ve dağılımı kontrol et.

10. İyi Bir A/B Testi İçin Kontrol Listesi

01

Net hipotez

Neyi değiştirdiğin ve hangi davranışı etkilemesini beklediğin açık mı?

02

Doğru metrik

Başarı ölçüsü iş hedefini gerçekten temsil ediyor mu?

03

Adil gruplar

Kullanıcılar rastgele ve dengeli dağıtılıyor mu?

04

Yeterli veri

Sonucun güvenilirliği için yeterli örneklem ve uygun süre var mı?

05

İstatistik

Belirsizlik, güven aralığı ve uygun hipotez testi dikkate alınıyor mu?

06

İş yorumu

Sonuç gerçek gelir, maliyet veya kullanıcı davranışı açısından anlamlı mı?

A/B testi web sitesi, mobil uygulama, reklam kreatifi, e-posta başlığı, açılış sayfası, buton metni, fiyatlandırma sunumu ve kullanıcı karşılama akışı gibi birçok alanda uygulanabiliyor.

Sonuç: Daha Az Tahmin, Daha Çok Kanıt

A/B testi, veri odaklı karar verme kültürünün en pratik yöntemlerinden biri. Kontrol ve test gruplarını karşılaştırarak bir değişikliğin gerçekten fayda sağlayıp sağlamadığını sistematik biçimde değerlendirmeye yardımcı oluyor.

Testleri başarılı kılan asıl unsur ise doğru hipotez, doğru deney tasarımı, doğru metrik ve doğru istatistiksel yorum kombinasyonu. Veri analistleri için bu süreç; ham kullanıcı davranışını ölçülebilir kanıta ve sonunda daha sağlam iş kararlarına dönüştürmenin en güçlü yollarından biri.

“A/B testinin değeri kazananı ilan etmek değil; kararın arkasındaki kanıtı görünür kılmaktır.”

Answering “which one is better?” with a controlled experiment instead of a guess. A/B testing measures user behaviour across websites, ad campaigns, email subject lines and product flows, and lets teams decide with evidence.

In this post I go through every step of an A/B test — from forming a hypothesis to control and test groups, from conversion rate to statistical significance — over a realistic e-commerce example.

“A good experiment does not only answer ‘which group won?'; it also shows the size of the difference, its uncertainty and its business value.”

1. What Is an A/B Test?

An A/B test is a controlled experiment that shows two different versions of a product, page, ad or message to user groups and measures which version performs the target behaviour better.

In the simplest setup, part of the users see version A (control) and the rest see version B (test). A predefined goal — a purchase, a sign-up, a form submission or an ad click — is then compared.

A

Control

The version currently in use; the reference point of the comparison.

B

Test

The version carrying the change; its effect is measured against the control.

KPI

Target metric

The main measure that decides the outcome: conversion, clicks, revenue or another indicator.

The A/B testing process
  1. Form a
    Hypothesis
  2. Set Up
    the Test
  3. Collect
    Data
  4. Analyse
  5. Decide
The core logic: compare two options serving the same goal under conditions as similar as possible, then use statistics to judge whether the observed difference comes from a real effect.

2. Why Does It Matter for Data-Driven Decisions?

In marketing and product decisions, intuitions such as “this design looks better” or “this headline will definitely get more clicks” can be useful starting points, but they are not evidence on their own. A/B testing turns those assumptions into measurable hypotheses.

Teams can then ask “by how much does version B change the conversion rate under these conditions, and is that difference statistically reliable?” instead of saying “B looks better”.

3. How Is the Hypothesis Formed?

A good experiment starts with a measurable hypothesis that explains not only the change but the expected outcome.

HypothesisStatement
H₀ (null)There is no meaningful difference between versions A and B on the target metric.
H₁ (alternative)Version B has a different effect on the target metric than A.
An example hypothesis: moving the “Add to cart” button to a more visible position will increase the purchase conversion rate on the product page.

The hypothesis has to be measurable. “The page will be better” gives nothing to test; “the purchase conversion rate will increase” gives a clear outcome.

4. Control and Test Groups

Assigning users to groups as randomly as possible matters. The aim is to prevent the two groups from differing needlessly in age, device, traffic source or previous behaviour.

A • Control

Current experience

The page or campaign in use today; the new version is compared against it.

B • Test

Changed experience

The version where one specific variable is changed to test the hypothesis.

Design

One variable

If only the headline changed, it is easier to attribute the result to that change.

Sample

Enough volume

In tests with very few users, small differences do not look reliable.

5. How Is the Conversion Rate Calculated?

One of the most used indicators in A/B tests is the conversion rate: the share of users who performed the target behaviour.

CR=CN×100CR = \frac{C}{N} \times 100

CR: conversion rate, C: number of conversions, N: total users.

GroupUsersConversionsConversion rate
A • Control100,0005,8005.8%
B • Test100,0007,1007.1%

How is the relative lift found?

Lift=CRB−CRACRA×100=7.1−5.85.8×100≈22.4Lift = \frac{CR_B - CR_A}{CR_A} \times 100 = \frac{7.1 - 5.8}{5.8} \times 100 \approx 22.4

In this example version B's conversion rate looks about 22.4% higher. But there is an important distinction here: an observed difference and a statistically reliable difference are not the same thing.

6. What Is Statistical Significance?

B scoring higher than A in one experiment does not prove on its own that B is genuinely better. The difference may be random fluctuation from sampling. Statistical significance helps evaluate how likely the observed difference is to have appeared by chance alone.

p

p-value

Evaluates how unusual the observed (or a more extreme) difference is under the tested assumption. It is not a “success percentage”.

CI

Confidence interval

Gives a range for how large the effect might be — far more informative than a plain “significant or not”.

Important: the common 0.05 threshold is used for decisions in many analyses, but the design of the test, the sample size, multiple comparisons and the business impact all have to be weighed together.

Statistical significance ≠ business significance

Version B may be statistically significantly 1% better. If the cost of that improvement, its effect on user experience or its revenue contribution is very small, the business decision can still go the other way. A good A/B test is judged both statistically and commercially.

7. A Realistic Example

On an e-commerce site the “Add to cart” button is thought not to be visible enough. The marketing team expects that a new design making the button more prominent will increase the purchase rate.

The experiment design

FieldDecision
HypothesisA more visible button increases purchase conversion.
ControlThe current product page
TestThe new button position and design
Primary metricPurchase conversion rate
Secondary metricAdd-to-cart rate and revenue per user
Split50% A / 50% B

The example result

Absolute

+1.3 points

The absolute increase in conversion rate (5.8% → 7.1%).

Relative

+22.4%

The relative improvement over the control group.

Support

2 extra metrics

Add-to-cart rate and revenue impact back the decision up.

In the final step the appropriate statistical test is run and the reliability of the difference is examined. If the result is significant under the planned analysis, rolling out version B can be considered — but the decision should also weigh revenue and user experience indicators, not the p-value alone.

A basic conversion calculation in Python

Conversion rate and lift
control_users = 100000
control_conversions = 5800

test_users = 100000
test_conversions = 7100

control_rate = control_conversions / control_users
test_rate = test_conversions / test_users

print("Control:", control_rate)
print("Test:", test_rate)

lift = (test_rate - control_rate) / control_rate
print("Relative lift:", lift)

To judge whether the difference is statistically significant, methods that test the difference between two proportions are used:

A z-test for two proportions
from statsmodels.stats.proportion import proportions_ztest

# conversions and user counts
conversions = [5800, 7100]
users = [100000, 100000]

z, p = proportions_ztest(conversions, users)
print("z:", z, "p:", p)

8. How Should the Results Be Read?

Rather than focusing on a single percentage, the whole experiment has to be judged. These four questions are a good start:

Q1

How big is the difference?

Calculate the absolute and relative effect separately.

Q2

Is it reliable?

Examine the statistical uncertainty and the confidence interval.

Q3

What is the business impact?

Evaluate revenue, cost, retention and other indicators.

Q4

Is the experiment sound?

Check the groups, duration, traffic and measurement quality.

A better way to state the result: instead of “version B gave a 22.4% lift”, say “version B converted 22.4% higher than the control; if the result is statistically significant within the planned analysis and the revenue indicators are also positive, B can be preferred”.

9. Common Mistakes

MistakeWhy is it a problem?Better approach
Looking at the result too earlyRandom fluctuations mislead.Follow the predefined plan and sample size requirement.
Changing too many things at onceIt becomes unclear which variable caused the effect.Make one clear change tied to the hypothesis.
Focusing only on the p-valueThe size of the effect and the uncertainty get missed.Weigh effect size, confidence interval and business metrics together.
Choosing the wrong metricAn easily measured metric may not represent the real goal.Tie the primary success metric to the business goal.
Imbalanced groupsThe fairness of the comparison breaks down.Check the randomisation and the split.

10. A Checklist for a Good A/B Test

01

A clear hypothesis

Is it obvious what you changed and which behaviour you expect it to affect?

02

The right metric

Does the success measure really represent the business goal?

03

Fair groups

Are users assigned randomly and in a balanced way?

04

Enough data

Is there enough sample and a suitable duration for a reliable result?

05

Statistics

Are uncertainty, the confidence interval and a suitable test taken into account?

06

Business reading

Is the outcome meaningful in terms of real revenue, cost or user behaviour?

A/B tests can be applied to websites, mobile apps, ad creatives, email subject lines, landing pages, button copy, pricing presentation and onboarding flows.

Conclusion: Less Guessing, More Evidence

A/B testing is one of the most practical methods in a data-driven decision culture. By comparing control and test groups it helps evaluate systematically whether a change really brings a benefit.

What makes a test successful is the combination of the right hypothesis, the right experiment design, the right metric and the right statistical reading. For data analysts it is one of the strongest ways of turning raw user behaviour into measurable evidence and, in the end, into sounder business decisions.

“The value of an A/B test is not declaring a winner; it is making the evidence behind the decision visible.”