Confusion Matrix Nedir? Accuracy, Precision, Recall ve F1-Score

Sınıflandırma modellerini doğru değerlendirmek için confusion matrix'i ve accuracy, precision, recall, F1-score metriklerini adım adım inceliyorum.

  • Makine Öğrenmesi
  • İstatistik

Sınıflandırma modellerinin gerçekten ne kadar iyi çalıştığını anlamak için yalnızca “kaç tanesini doğru bildi?” sorusu yeterli değil. Confusion matrix, modelin doğru ve yanlış tahminlerini ayrıntılı biçimde görmemizi sağlıyor.

Bu yazıda confusion matrix'in dört temel hücresinden başlayıp accuracy, precision, recall ve F1-score metriklerini; hangi problemde hangisine bakmak gerektiğini ve dengesiz veride accuracy'nin neden yanıltıcı olduğunu örneklerle ele alıyorum.

“Model %90 doğru” demek tek başına yeterli değildir; o %10'luk hatanın hangi sınıfta ve hangi yönde gerçekleştiğini de bilmek gerekir.”

Confusion Matrix Nedir?

Confusion matrix (karışıklık matrisi), bir sınıflandırma modelinin gerçek sınıflar ile tahmin ettiği sınıfları karşılaştıran tablodur. Özellikle ikili sınıflandırmada modelin yalnızca doğru tahminlerini değil, hangi tür hataları yaptığını görmemizi sağlar.

Bu ayrım önemli; çünkü bazı problemlerde yanlış pozitif üretmek, bazı problemlerde ise gerçek pozitifi kaçırmak çok daha maliyetli.

Model değerlendirme akışı
  1. Confusion
    Matrix
  2. Accuracy
  3. Precision
  4. Recall
  5. F1-Score

1. TP, TN, FP ve FN

Confusion matrix'in temelini dört sonuç oluşturur. Buradaki “pozitif” ve “negatif” problemin sınıflarıdır; otomatik olarak “iyi” ve “kötü” anlamına gelmez.

Tahmin: pozitifTahmin: negatif
Gerçek: pozitifTrue Positive (TP)
Pozitifi doğru buldu.
False Negative (FN)
Pozitifi negatif sandı.
Gerçek: negatifFalse Positive (FP)
Negatifi pozitif sandı.
True Negative (TN)
Negatifi doğru buldu.
TP

Doğru pozitif

Model pozitif dedi ve gerçekten pozitifti.

TN

Doğru negatif

Model negatif dedi ve gerçekten negatifti.

FP

Yanlış pozitif

Model pozitif dedi ama değildi — yanlış alarm.

FN

Yanlış negatif

Model negatif dedi ama pozitifti — kaçırılan vaka.

2. Gerçek Hayat Örneği: Spam E-posta Tespiti

Bir e-posta modelinin mesajları spam veya normal olarak sınıflandırdığını düşünelim. 1.000 e-posta üzerinde şu sonuçlar alınsın:

Tahmin: spamTahmin: normal
Gerçek: spamTP = 180FN = 20
Gerçek: normalFP = 30TN = 770

Model 180 spam mesajını doğru yakalıyor. Ancak 20 spam mesajını normal sanarak kaçırıyor, 30 normal e-postayı da yanlışlıkla spam klasörüne gönderiyor.

Neden bu ayrım önemli? Kullanıcı için önemli bir e-postanın spam diye işaretlenmesi FP hatasıdır; spam mesajın gelen kutusuna düşmesi ise FN. Hangi hatanın daha kritik olduğu iş kuralına göre değişir.

3. Accuracy, Precision, Recall ve F1-Score

Confusion matrix'teki dört değeri kullanarak model performansını farklı açılardan ölçebiliriz.

Accuracy — genel doğruluk

Modelin tüm örnekler içindeki doğru tahmin oranını gösterir.

Accuracy=TP+TNTP+TN+FP+FNAccuracy = \frac{TP + TN}{TP + TN + FP + FN}

Sınıflar dengeliyse ve FP ile FN hatalarının maliyeti birbirine yakınsa iyi bir başlangıç metriğidir.

Precision — pozitif tahminlerin kalitesi

Modelin pozitif dediği örneklerin ne kadarının gerçekten pozitif olduğunu ölçer.

Precision=TPTP+FPPrecision = \frac{TP}{TP + FP}

Yanlış pozitifleri azaltmak kritikse öne çıkar: gereksiz alarm, yanlış kampanya hedefleme veya yanlış işlem başlatma gibi durumlar.

Recall — pozitifleri yakalama gücü

Gerçekte pozitif olanların ne kadarını modelin yakalayabildiğini gösterir.

Recall=TPTP+FNRecall = \frac{TP}{TP + FN}

Pozitif vakaları kaçırmak kritikse öne çıkar: riskli işlem, arıza veya önemli bir vaka tespiti gibi senaryolar.

F1-score — precision ve recall dengesi

Precision ve recall'un harmonik ortalamasıdır; iki metriğin birlikte güçlü olmasını ister.

F1=2⋅Precision⋅RecallPrecision+RecallF_1 = 2 \cdot \frac{Precision \cdot Recall}{Precision + Recall}

Precision ve recall'u tek bir özet sayıyla birlikte değerlendirmek istediğinde, özellikle sınıflar dengesizken faydalıdır.

4. Örnek Üzerinden Hesaplama

Yukarıdaki spam örneğinde TP = 180, TN = 770, FP = 30 ve FN = 20 olduğunda metrikler şöyle çıkar:

MetrikHesapSonuç
Accuracy(180 + 770) / 10000,95 → %95
Precision180 / (180 + 30)≈ 0,857 → %85,7
Recall180 / (180 + 20)0,90 → %90
F1-score2 × (0,857 × 0,90) / (0,857 + 0,90)≈ 0,878 → %87,8

Dört metrik de yüksek görünüyor; ancak recall'un precision'dan yüksek olması modelin spam yakalamada cömert davrandığını, bunun karşılığında 30 normal e-postayı yanlışlıkla spam'e attığını söylüyor.

5. Hangi Metrik Ne Zaman Tercih Edilmeli?

Precision

Yanlış alarm pahalıysa

Modelin pozitif dediği kayıtların mümkün olduğunca gerçekten pozitif olmasını istersin.

Recall

Vaka kaçırmak pahalıysa

Gerçek pozitifleri kaçırmamak önceliklidir; FN hatasının maliyeti yüksektir.

F1-score

Denge arıyorsan

Precision ile recall arasında dengeli performans ve tek bir karşılaştırma metriği istersin.

Accuracy

Sınıflar dengeliyse

Yanlış pozitif ve negatifin maliyeti birbirine yakınsa genel doğruluk anlamlıdır.

SenaryoÖncelikli metrikNeden?
Dolandırıcılık alarmıPrecisionÇok fazla yanlış alarm operasyon ekibini yorar.
Hastalık taramasıRecallGerçek vakaları kaçırmamak daha kritiktir.
Metin sınıflandırmaF1-scoreYakalama ve yanlış alarm dengesine birlikte bakılır.
Dengeli iki sınıfAccuracy + diğerleriGenel doğruluk anlamlı bir özet sunar.

6. Dengesiz Veride Accuracy Neden Yanıltır?

Diyelim ki 10.000 müşterinin yalnızca 100'ü dolandırıcılık vakası. Model tüm müşterilere “dolandırıcılık değil” derse 9.900 doğru tahmin yapmış olur ve accuracy %99 çıkar.

İlk bakışta mükemmel görünür. Fakat model gerçek dolandırıcılık vakalarının hiçbirini yakalayamamıştır; recall %0'dır.

Ders: Sınıflar ciddi biçimde dengesizse accuracy'yi tek başına raporlamak risklidir. Confusion matrix, precision, recall ve F1-score ile birlikte değerlendirmek gerekir.

7. Python ile Confusion Matrix ve Metrikler

scikit-learn ile temel metrikleri birkaç satırda hesaplayabilirsin.

scikit-learn ile hesaplama
from sklearn.metrics import (
    confusion_matrix, accuracy_score,
    precision_score, recall_score, f1_score
)

y_true = [1, 1, 1, 1, 0, 0, 0, 0]
y_pred = [1, 1, 0, 1, 1, 0, 0, 0]

cm = confusion_matrix(y_true, y_pred)
accuracy = accuracy_score(y_true, y_pred)
precision = precision_score(y_true, y_pred)
recall = recall_score(y_true, y_pred)
f1 = f1_score(y_true, y_pred)

print(cm)
print(accuracy, precision, recall, f1)

Tek bir rapor satırında hepsini görmek istersen classification_report() sınıf bazında precision, recall ve F1 değerlerini birlikte verir.

Sınıf bazında rapor
from sklearn.metrics import classification_report

print(classification_report(y_true, y_pred))
Not: Bu kod yalnızca hesaplama örneğidir. Gerçek bir projede karar eşiği, sınıf dengesizliği, cross-validation ve iş maliyetleri de değerlendirilmelidir.

8. Model Seçiminde Confusion Matrix'i Kullanmak

İki modelin accuracy değeri aynı olsa bile hata profilleri tamamen farklı olabilir. Model karşılaştırırken şu sırayla düşünmek faydalı:

Karşılaştırma sırası
  1. Confusion
    Matrix
  2. Precision
  3. Recall
  4. F1-Score
  5. İş
    Maliyeti

Son aşamada metrikleri iş hedefleriyle birleştirmek gerekiyor. Pazarlama modelinde FP gereksiz hedefleme maliyeti yaratırken, risk tespitinde FN çok daha pahalıya mal olabilir.

Sonuç

Confusion matrix, bir sınıflandırma modelini yalnızca “kaç doğru yaptı?” seviyesinde değil, hangi tür hataları neden ürettiği açısından incelemeni sağlar.

Accuracy genel performansı, precision pozitif tahminlerin kalitesini, recall gerçek pozitifleri yakalama gücünü, F1-score ise bu ikisi arasındaki dengeyi özetler.

“En yüksek metriğe sahip modeli otomatik seçmek yerine, iş probleminin hangi hataya daha fazla tolerans gösterdiğini belirle.”

To understand how well a classification model really works, “how many did it get right?” is not enough on its own. A confusion matrix lets us see the model's correct and incorrect predictions in detail.

In this post I start from the four cells of the confusion matrix and go through accuracy, precision, recall and the F1-score: which metric to look at in which problem, and why accuracy misleads on imbalanced data.

“Saying ‘the model is 90% accurate' is not enough; you also need to know in which class and in which direction that 10% of error happened.”

What Is a Confusion Matrix?

A confusion matrix is the table that compares the real classes with the classes a classification model predicted. In binary classification especially, it shows not only the correct predictions but which kind of mistakes the model makes.

That distinction matters, because in some problems producing a false positive is costly, while in others missing a true positive is far worse.

The model evaluation flow
  1. Confusion
    Matrix
  2. Accuracy
  3. Precision
  4. Recall
  5. F1-Score

1. TP, TN, FP and FN

Four outcomes form the basis of the confusion matrix. “Positive” and “negative” here are the classes of the problem; they do not automatically mean “good” and “bad”.

Predicted: positivePredicted: negative
Actual: positiveTrue Positive (TP)
Found the positive correctly.
False Negative (FN)
Thought the positive was negative.
Actual: negativeFalse Positive (FP)
Thought the negative was positive.
True Negative (TN)
Found the negative correctly.
TP

True positive

The model said positive and it really was positive.

TN

True negative

The model said negative and it really was negative.

FP

False positive

The model said positive but it was not — a false alarm.

FN

False negative

The model said negative but it was positive — a missed case.

2. A Real-Life Example: Spam Detection

Imagine an email model that classifies messages as spam or normal. On 1,000 emails it produces these results:

Predicted: spamPredicted: normal
Actual: spamTP = 180FN = 20
Actual: normalFP = 30TN = 770

The model catches 180 spam messages correctly. But it misses 20 spam messages by treating them as normal, and it sends 30 normal emails to the spam folder by mistake.

Why does the distinction matter? An important email marked as spam is an FP error; a spam message landing in the inbox is an FN. Which error is more critical depends on the business rule.

3. Accuracy, Precision, Recall and F1-Score

Using the four values in the confusion matrix, we can measure model performance from different angles.

Accuracy — overall correctness

Shows the proportion of correct predictions across all examples.

Accuracy=TP+TNTP+TN+FP+FNAccuracy = \frac{TP + TN}{TP + TN + FP + FN}

It is a good starting metric when the classes are balanced and the cost of FP and FN errors is similar.

Precision — the quality of positive predictions

Measures how many of the examples the model called positive really are positive.

Precision=TPTP+FPPrecision = \frac{TP}{TP + FP}

It comes first when reducing false positives is critical: unnecessary alarms, wrong campaign targeting or triggering the wrong process.

Recall — the power to catch positives

Shows how many of the truly positive cases the model managed to catch.

Recall=TPTP+FNRecall = \frac{TP}{TP + FN}

It comes first when missing positive cases is critical: risky transactions, failures or detecting an important case.

F1-score — the balance of precision and recall

The harmonic mean of precision and recall; it wants both metrics to be strong together.

F1=2⋅Precision⋅RecallPrecision+RecallF_1 = 2 \cdot \frac{Precision \cdot Recall}{Precision + Recall}

It helps when you want to weigh precision and recall together in a single summary number, especially with imbalanced classes.

4. Calculating the Metrics on the Example

With TP = 180, TN = 770, FP = 30 and FN = 20 from the spam example, the metrics come out as:

MetricCalculationResult
Accuracy(180 + 770) / 10000.95 → 95%
Precision180 / (180 + 30)≈ 0.857 → 85.7%
Recall180 / (180 + 20)0.90 → 90%
F1-score2 × (0.857 × 0.90) / (0.857 + 0.90)≈ 0.878 → 87.8%

All four look high, but recall being above precision tells us the model is generous in flagging spam, and pays for it by sending 30 normal emails to the spam folder.

5. Which Metric to Prefer, and When?

Precision

When false alarms are costly

You want the records the model calls positive to really be positive.

Recall

When missing cases is costly

Not missing true positives comes first; the cost of FN is high.

F1-score

When you need balance

You want balanced performance between precision and recall in one comparison metric.

Accuracy

When classes are balanced

If FP and FN cost about the same, overall correctness is meaningful.

ScenarioPriority metricWhy?
Fraud alertingPrecisionToo many false alarms exhaust the operations team.
Disease screeningRecallNot missing real cases is more critical.
Text classificationF1-scoreCatching positives and avoiding false alarms are weighed together.
Two balanced classesAccuracy + the othersOverall correctness gives a meaningful summary.

6. Why Does Accuracy Mislead on Imbalanced Data?

Say only 100 of 10,000 customers are fraud cases. If the model calls every customer “not fraud”, it makes 9,900 correct predictions and accuracy lands at 99%.

At first glance that looks perfect. But the model caught none of the real fraud cases; recall is 0%.

The lesson: when classes are seriously imbalanced, reporting accuracy alone is risky. It has to be read together with the confusion matrix, precision, recall and the F1-score.

7. The Confusion Matrix and Metrics in Python

With scikit-learn you can calculate the core metrics in a few lines.

Calculating with scikit-learn
from sklearn.metrics import (
    confusion_matrix, accuracy_score,
    precision_score, recall_score, f1_score
)

y_true = [1, 1, 1, 1, 0, 0, 0, 0]
y_pred = [1, 1, 0, 1, 1, 0, 0, 0]

cm = confusion_matrix(y_true, y_pred)
accuracy = accuracy_score(y_true, y_pred)
precision = precision_score(y_true, y_pred)
recall = recall_score(y_true, y_pred)
f1 = f1_score(y_true, y_pred)

print(cm)
print(accuracy, precision, recall, f1)

If you want to see everything in one report, classification_report() gives precision, recall and F1 per class together.

A per-class report
from sklearn.metrics import classification_report

print(classification_report(y_true, y_pred))
Note: this code is only a calculation example. In a real project the decision threshold, class imbalance, cross-validation and business costs all need to be considered.

8. Using the Confusion Matrix in Model Selection

Two models can share the same accuracy and still have completely different error profiles. When comparing models, this order helps:

The comparison order
  1. Confusion
    Matrix
  2. Precision
  3. Recall
  4. F1-Score
  5. Business
    Cost

In the final step the metrics have to be combined with business goals. In a marketing model an FP creates wasted targeting cost, while in risk detection an FN can be far more expensive.

Conclusion

A confusion matrix lets you examine a classification model not only at the level of “how many did it get right?” but in terms of which kinds of errors it produces and why.

Accuracy summarises overall performance, precision the quality of positive predictions, recall the power to catch true positives, and the F1-score the balance between the last two.

“Instead of automatically picking the model with the highest metric, decide which error your business problem tolerates more.”