Keşifsel Veri Analizi (EDA) Nedir?

Veri setini tanımadan modele geçmek yerine; veriyi keşfetmenin, sorunları bulmanın ve analitik soruları doğru kurmanın temellerini inceliyorum.

  • Araçlar
  • İstatistik

Veri analizi sürecinde model kurmak, grafik üretmek veya istatistiksel test uygulamak çoğu zaman işin görünen kısmıdır. Fakat bütün bunlardan önce yapılması gereken çok daha temel bir aşama var: veriyi tanımak.

Keşifsel Veri Analizi, yani Exploratory Data Analysis (EDA), bir veri setinin yapısını anlamak, eksik ve sıra dışı değerleri tespit etmek, değişkenler arasındaki ilişkileri keşfetmek ve analiz öncesinde veriden mümkün olduğunca fazla bilgi çıkarmak için kullanılan yaklaşımdır.

“İyi bir analiz, iyi bir soruyla başlar; iyi bir soru ise çoğu zaman veriyi keşfederken ortaya çıkar.”

EDA Neden Önemlidir?

Bir veri setini doğrudan modele vermek cazip görünebilir. Ancak verinin yapısını anlamadan yapılan analizler yanıltıcı sonuçlara yol açabilir. EDA, bu riski analiz başlamadan görünür hale getirir.

01 — Veri

Eksik değerleri fark etmek

Hangi değişkende ne kadar eksik veri olduğunu ve bu eksikliğin analize etkisini görmek.

02 — Kalite

Hatalı kayıtları bulmak

Duplicate kayıtları, yanlış veri tiplerini veya mantıksız değerleri analiz öncesinde tespit etmek.

03 — Dağılım

Verinin davranışını görmek

Ortalamaları, dağılımları, aykırı değerleri ve olası çarpıklıkları incelemek.

04 — İlişki

Örüntüleri keşfetmek

Değişkenlerin birbirleriyle nasıl hareket ettiğini görerek yeni sorular oluşturmak.

EDA çalışma akışı
  1. Veriyi
    Tanı
  2. Temizle
  3. Keşfet
  4. İlişkileri
    İncele
  5. Yorumla

1. Veri Setini Tanımak

EDA'ya başlarken ilk olarak elimizde nasıl bir veri olduğunu anlamamız gerekir. Python tarafında bunun için pandas oldukça güçlü ve pratik bir araç.

İlk pandas kontrolleri
import pandas as pd

df = pd.read_csv("satislar.csv")

print(df.shape)
print(df.head())
df.info()
print(df.dtypes)
KomutNe gösterir?
df.shapeSatır ve sütun sayısını
df.head()İlk kayıtları
df.tail()Son kayıtları
df.info()Sütunları, veri tiplerini ve doluluk durumunu
df.dtypesDeğişkenlerin veri tiplerini

2. Temel İstatistikleri İncelemek

Veri setinin genel davranışını hızlıca görmek için describe() oldukça kullanışlı. Tek satırda merkezi eğilim, yayılım ve uç değerler hakkında fikir verir.

pandas ile özet istatistik
df.describe()

# Kategorik alanlar da dahil
df.describe(include="all")
mean

Ortalama

Değerlerin merkezi eğilimi hakkında fikir verir.

median

Medyan

Özellikle aykırı değerlerin etkisini anlamada yararlıdır.

std

Standart sapma

Değerlerin ortalama etrafında ne kadar yayıldığını gösterir.

min / max

Uç değerler

Dağılımın sınırlarını ve sıra dışı gözlemleri incelemeye yardımcı olur.

3. Eksik Değerleri Keşfetmek

EDA sırasındaki en önemli kontrollerden biri eksik veri analizi. Hangi sütunda ne kadar boşluk olduğunu bilmeden yapılan her hesaplama yanıltıcı olabilir.

Eksik değer kontrolü
# Sütun başına eksik değer sayısı
df.isna().sum()

# Oran olarak görmek daha okunaklı
df.isna().mean().sort_values(ascending=False)

Örneğin gelir sütununda çok sayıda eksik değer bulunuyorsa, bu değişkenin analizde nasıl kullanılacağına ayrıca karar vermek gerekir.

Önemli: Eksik değerleri otomatik olarak silmek her zaman doğru değildir. Eksikliğin nedeni, değişkenin önemi ve analiz amacı birlikte değerlendirilmelidir.

4. Duplicate Kayıtları Kontrol Etmek

Aynı kaydın birden fazla kez bulunması, özellikle farklı veri kaynakları birleştirildiğinde karşılaşılan sorunlardan biri.

Duplicate kontrolü
# Duplicate sayısı
df.duplicated().sum()

# Duplicate kayıtları görüntüle
df[df.duplicated()]

# Gerekliyse kaldır
df = df.drop_duplicates()

Birbirine benzeyen iki kaydın gerçekten duplicate olup olmadığını kontrol etmek önemli; aynı görünen iki işlem farklı olaylar olabilir.

5. Aykırı Değerleri İncelemek

Aykırı değerler, veri setindeki genel dağılımdan belirgin biçimde ayrılan gözlemlerdir. Ancak istatistiksel olarak sıra dışı olmak, değerin hatalı olduğu anlamına gelmez.

IQR yaklaşımı
Q1 = df["satis"].quantile(0.25)
Q3 = df["satis"].quantile(0.75)

IQR = Q3 - Q1
alt = Q1 - 1.5 * IQR
ust = Q3 + 1.5 * IQR

aykiri = df[(df["satis"] < alt) | (df["satis"] > ust)]
Dikkat: Her aykırı değer silinmemeli. Gerçek bir büyük sipariş de istatistiksel olarak aykırı görünebilir.

6. Değişkenlerin Dağılımını Keşfetmek

Bir değişkenin nasıl dağıldığını görmek için histogram, boxplot ve yoğunluk grafikleri kullanılabilir. Ortalama tek başına dağılımın şeklini anlatmaz.

histogram

Frekans

Değerlerin hangi aralıklarda yoğunlaştığını gösterir.

boxplot

Dağılım

Medyanı, çeyrekleri ve olası aykırı değerleri birlikte görmeyi sağlar.

Python ile hızlı histogram
import matplotlib.pyplot as plt

df["satis"].hist(bins=30)
plt.show()

7. Kategorik Değişkenleri İncelemek

EDA sadece sayısal değişkenlerden oluşmaz. Kategorik alanların frekanslarını görmek de veri setini anlamak için önemli.

Kategori frekansları
df["kategori"].value_counts()

# Oransal dağılım
df["kategori"].value_counts(normalize=True)

Bu analiz sayesinde veri setinin hangi kategorilerde yoğunlaştığını ve örneklemin dengeli olup olmadığını görebiliriz.

8. Değişkenler Arasındaki İlişkileri Keşfetmek

EDA'nın en değerli taraflarından biri, değişkenleri tek tek değil birlikte incelemek. Örneğin gelir ile satış arasında ilişki olup olmadığını sorgulayabiliriz.

Scatter plot ve korelasyon
df.plot(kind="scatter", x="gelir", y="satis")
plt.show()

# Sayısal değişkenler arası korelasyon
df.corr(numeric_only=True)
Unutma: Korelasyon tek başına nedensellik kanıtı değildir. İki değişkenin birlikte hareket etmesi, birinin diğerine sebep olduğunu göstermez.

9. Grup Bazında Analiz

Toplam değerler bazen veri setinin içindeki farklılıkları gizler. Bu yüzden kategoriler veya segmentler bazında analiz yapmak önemli.

groupby ile segment analizi
df.groupby("kategori")["satis"].agg(["count", "mean", "median", "std"])

EDA Sırasında Kendime Sorduğum Sorular

Veri

Boyut

Kaç satır ve kaç sütun var?

Kalite

Sorunlar

Eksik, duplicate veya mantıksız değer var mı?

Tipler

Veri tipleri

Değişken tipleri doğru mu?

Dağılım

Davranış

Değişkenler nasıl dağılıyor?

İlişki

Örüntüler

Değişkenler arasında hangi örüntüler var?

Soru

Sonraki adım

Bu keşifler hangi yeni soruları doğuruyor?

Python ile Basit Bir EDA Akışı

Tüm temel kontrolleri küçük bir akışta birleştirebiliriz. Yeni bir veri seti geldiğinde ilk çalıştırdığım şablon bu:

Mini EDA şablonu
import pandas as pd

df = pd.read_csv("satislar.csv")

print(df.shape)
print(df.head())
df.info()
print(df.isna().sum())
print(df.duplicated().sum())
print(df.describe())
print(df["kategori"].value_counts())
print(df.corr(numeric_only=True))

EDA Bir Sonuç Değil, Bir Başlangıçtır

EDA'nın amacı yalnızca “veri temiz” demek değil. Asıl amaç, verinin bize hangi soruları sormamız gerektiğini göstermesi.

Örneğin gelir ile satış arasında dikkat çekici bir ilişki bulabiliriz. Bu gözlem “Gelir arttıkça satış da artıyor mu?” veya “Bu ilişki her segmentte aynı mı?” gibi yeni sorular doğurur.

Bu nedenle EDA, basit bir veri kontrolünden çıkar ve analitik düşünme sürecinin başlangıç noktasına dönüşür.

Sonuç

Keşifsel Veri Analizi, veri analizi sürecinin en önemli aşamalarından biri. Çünkü model kurmadan, istatistiksel test uygulamadan veya dashboard hazırlamadan önce elimizdeki verinin gerçekten ne anlattığını anlamamız gerekiyor.

Python ve pandas bu süreçte güçlü araçlar sunuyor. head(), info(), describe(), isna(), duplicated(), groupby() ve korelasyon analizleri gibi temel fonksiyonlar bile veri seti hakkında çok şey anlatabiliyor.

EDA'nın özeti: Veriyi tanı → Sorunları bul → Dağılımları incele → İlişkileri keşfet → Soruları oluştur → Analize geç.

In a data analysis process, building a model, producing charts or running statistical tests is usually the visible part of the work. But there is a much more fundamental step that comes before all of it: getting to know the data.

Exploratory Data Analysis (EDA) is the approach used to understand the structure of a data set, spot missing and unusual values, discover relationships between variables and extract as much information as possible from the data before the analysis begins.

“A good analysis starts with a good question, and a good question usually appears while you are exploring the data.”

Why Does EDA Matter?

Feeding a data set straight into a model can look tempting. But analyses made without understanding the structure of the data can lead to misleading results. EDA makes that risk visible before the analysis starts.

01 — Data

Noticing missing values

Seeing how much data is missing in each variable and how that affects the analysis.

02 — Quality

Finding faulty records

Detecting duplicates, wrong data types or implausible values before the analysis.

03 — Distribution

Seeing how data behaves

Examining means, distributions, outliers and possible skewness.

04 — Relationship

Discovering patterns

Seeing how variables move together and forming new questions from it.

The EDA workflow
  1. Know
    the Data
  2. Clean
  3. Explore
  4. Examine
    Relationships
  5. Interpret

1. Getting to Know the Data Set

When starting EDA, the first thing to understand is what kind of data we actually have. On the Python side, pandas is a powerful and practical tool for this.

First pandas checks
import pandas as pd

df = pd.read_csv("sales.csv")

print(df.shape)
print(df.head())
df.info()
print(df.dtypes)
CommandWhat does it show?
df.shapeThe number of rows and columns
df.head()The first records
df.tail()The last records
df.info()Columns, data types and how many values are filled in
df.dtypesThe data types of the variables

2. Looking at Basic Statistics

To see the general behaviour of a data set quickly, describe() is very useful. In a single line it gives an idea about central tendency, spread and extreme values.

Summary statistics with pandas
df.describe()

# Including categorical columns
df.describe(include="all")
mean

Average

Gives an idea about the central tendency of the values.

median

Median

Especially useful for understanding the effect of outliers.

std

Standard deviation

Shows how widely the values spread around the mean.

min / max

Extreme values

Helps examine the limits of the distribution and unusual observations.

3. Exploring Missing Values

One of the most important checks during EDA is missing data analysis. Any calculation made without knowing how many gaps each column has can be misleading.

Checking missing values
# Number of missing values per column
df.isna().sum()

# Seeing it as a ratio is easier to read
df.isna().mean().sort_values(ascending=False)

If a column such as income has a large number of missing values, for example, you need a separate decision about how that variable will be used in the analysis.

Important: Deleting missing values automatically is not always right. The reason for the gap, the importance of the variable and the purpose of the analysis should be considered together.

4. Checking Duplicate Records

The same record appearing more than once is one of the problems you meet especially when different data sources are combined.

Duplicate check
# Number of duplicates
df.duplicated().sum()

# Display the duplicate records
df[df.duplicated()]

# Remove them if needed
df = df.drop_duplicates()

It is important to check whether two similar-looking records really are duplicates; two transactions that look the same can be different events.

5. Examining Outliers

Outliers are observations that clearly separate from the general distribution in a data set. But being statistically unusual does not mean the value is wrong.

The IQR approach
Q1 = df["sales"].quantile(0.25)
Q3 = df["sales"].quantile(0.75)

IQR = Q3 - Q1
lower = Q1 - 1.5 * IQR
upper = Q3 + 1.5 * IQR

outliers = df[(df["sales"] < lower) | (df["sales"] > upper)]
Careful: Not every outlier should be deleted. A genuinely large order can also look statistically extreme.

6. Exploring the Distribution of Variables

Histograms, boxplots and density plots can be used to see how a variable is distributed. The mean alone does not describe the shape of a distribution.

histogram

Frequency

Shows the ranges in which the values concentrate.

boxplot

Spread

Lets you see the median, the quartiles and possible outliers together.

A quick histogram in Python
import matplotlib.pyplot as plt

df["sales"].hist(bins=30)
plt.show()

7. Examining Categorical Variables

EDA is not only about numerical variables. Seeing the frequencies of categorical fields is also important for understanding the data set.

Category frequencies
df["category"].value_counts()

# Proportional distribution
df["category"].value_counts(normalize=True)

This analysis shows which categories the data set concentrates in and whether the sample is balanced.

8. Discovering Relationships Between Variables

One of the most valuable sides of EDA is examining variables together rather than one by one. For example, we can ask whether there is a relationship between income and sales.

Scatter plot and correlation
df.plot(kind="scatter", x="income", y="sales")
plt.show()

# Correlation between numerical variables
df.corr(numeric_only=True)
Remember: Correlation alone is not proof of causation. Two variables moving together does not show that one causes the other.

9. Analysis by Group

Totals sometimes hide the differences inside a data set. That is why analysing by categories or segments matters.

Segment analysis with groupby
df.groupby("category")["sales"].agg(["count", "mean", "median", "std"])

The Questions I Ask Myself During EDA

Data

Size

How many rows and columns are there?

Quality

Problems

Are there missing, duplicate or implausible values?

Types

Data types

Are the variable types correct?

Distribution

Behaviour

How are the variables distributed?

Relationship

Patterns

Which patterns exist between the variables?

Question

Next step

Which new questions do these findings raise?

A Simple EDA Flow in Python

All the basic checks can be combined into a small flow. This is the template I run first whenever a new data set arrives:

Mini EDA template
import pandas as pd

df = pd.read_csv("sales.csv")

print(df.shape)
print(df.head())
df.info()
print(df.isna().sum())
print(df.duplicated().sum())
print(df.describe())
print(df["category"].value_counts())
print(df.corr(numeric_only=True))

EDA Is Not a Result, It Is a Beginning

The purpose of EDA is not only to say “the data is clean”. The real purpose is for the data to show us which questions we should be asking.

We may find a striking relationship between income and sales, for example. That observation raises new questions such as “Do sales grow as income grows?” or “Is this relationship the same in every segment?”.

This is why EDA moves beyond a simple data check and becomes the starting point of analytical thinking.

Conclusion

Exploratory Data Analysis is one of the most important stages of the data analysis process. Before building a model, running a statistical test or preparing a dashboard, we need to understand what the data in front of us really says.

Python and pandas offer strong tools for this. Even basic functions such as head(), info(), describe(), isna(), duplicated(), groupby() and correlation analysis can tell a lot about a data set.

EDA in short: Know the data → Find the problems → Examine the distributions → Discover the relationships → Form the questions → Move to analysis.

Python ile veri görselleştirme yöntemleri

Big Data (Büyük Veri) Nedir?