İstanbul'da Bir Airbnb İlanının Fiyatını Ne Belirliyor?What Sets the Price of an Airbnb Listing in Istanbul?
26.631 ilanlık ham veriyi temizledim, fiyat sütununun iki ayrı pazarı ölçtüğünü buldum ve gecelik fiyatı tahmin eden modeli hiç görmediği ev sahipleriyle sınadım.I cleaned the raw data of 26,631 listings, found that the price column measures two different markets, and tested a nightly-price model on hosts it had never seen.
- Python
- Pandas
- Jupyter
- Scikit-learn
Bir veri setinin ilk hali neredeyse hiçbir zaman analize hazır değildir. Bu çalışmada İstanbul'daki 26.631 Airbnb ilanının ham verisini temizledim, fiyatın nasıl dağıldığına baktım ve gecelik fiyatı tahmin eden bir model kurdum. En önemli bulgu modelden önce, temizlik sırasında çıktı.
Veri
Veri, Inside Airbnb'nin İstanbul için yayımladığı ayrıntılı ilan dosyası (CC BY 4.0). Her satır bir ilan: konum, oda tipi, kapasite, olanaklar, ev sahibi bilgisi, yorum puanları ve fiyat.
- Ham veriyi
tanı - Temizle
- Keşfet
- Modelle
ve sına
Tüm adımlar tek bir Jupyter defterinde, çıktılarıyla birlikte: defteri indir (.ipynb).
Neler bozuktu, ne yaptım?
İlk bakışta sekiz ayrı sorun çıktı. Hiçbiri tek başına zor değil; zor olan, her birinde neyi neden yaptığını yazabilmek.
| Sorun | Örnek | Ne yaptım |
|---|---|---|
| Fiyat metin, işareti yanıltıcı | $4,296.25 | İşareti ve binlik ayracını temizledim. Teklifin ham halinde tutarlar ₺ ile yazılı olduğu için birimi TL aldım. |
| Banyo sayısı metin | 1.5 shared baths, Half-bath | Sayı ve “ortak banyo” işareti olarak ikiye ayırdım. Boş oranı %45,8'den %0,5'e indi. |
| Olanaklar tek hücrede liste | ["Wifi", "Kitchen", …] | 4.372 farklı olanak var. Sayısını ve 11 seçili olanağı ayrı sütunlara çıkardım. |
| Mülk tipi dağınık | 90 farklı değer | Beş gruba indirdim: daire, otel/pansiyon, müstakil ev/villa, rezidans, diğer. |
| Aynı değer farklı yazılmış | İstanbul, Turkey, Istanbul, Türkiye | Tek değere topladım. |
| İzin numarası serbest metin | 2022-34-1997, 12-3456 | Biçimine göre beş sınıfa ayırdım. 149 ilanda numara yerine 12-3456 yazılmış. |
| Tamamen boş sütunlar | 13 sütun | Attım. |
| Uç değerler | 50 yatak odası, gecelik ₺1.706.000 | Oda, yatak ve banyo sayısını 10'da kırptım; fiyatın en uç %1'ini çıkardım. |
Eksik değerler rastgele değil
Fiyatın %10'u, yatak ve banyo sayısının neredeyse yarısı boş. Bunları doldurmadan önce boşluğun nereden geldiğine baktım. Inside Airbnb her ilan için bilginin bu taramadan mı, önceki taramadan mı geldiğini yazıyor; eksikler neredeyse tamamen ikinci grupta.
| Bilginin kaynağı | İlan | Fiyatı boş | Yatak sayısı boş | Banyo sayısı boş |
|---|---|---|---|---|
| Bu tarama | 15.395 | %0,5 | %1,4 | %6,2 |
| Önceki tarama | 11.236 | %23,2 | %100,0 | %100,0 |
Değerlendirme puanlarında da durum aynı: puanı olmayan her ilan, hiç yorumu olmayan bir ilan.
Aynı sütunda iki ayrı fiyat
Fiyat, ilanın kabul ettiği en kısa konaklama için alınmış bir tekliften hesaplanıyor. Teklifin giriş ve çıkış tarihleri veride var; aradaki gece sayısına bakınca beklemediğim bir şey çıktı.
İlanların %40'ı en az 100 gece koşulu koymuş. Bu ilanlarda fiyat 100 gecelik bir konaklama için alınıp geceye bölünmüş; yani gecelik fiyat yerine aylık kiranın günlüğü. İki grup aynı sütunda duruyor ama kıyaslanamaz.
Sütunu olduğu gibi kullansaydım “İstanbul'da medyan gecelik fiyat ₺2.853” derdim; bu rakam iki pazarın karışımı. Analizin geri kalanını teklifi 28 geceden kısa olan ilanlarla yaptım.
100 gece nereden geliyor?
Türkiye'de 1 Ocak 2024'te yürürlüğe giren 7464 sayılı kanun, 100 gün ve daha kısa süreli konut kiralamaları için izin belgesi istiyor (Airbnb'nin açıklaması). İlanlardaki izin numarası alanına bakınca örüntü çok net.
| İzin alanında yazan | İlan | En az 100 gece isteyenler |
|---|---|---|
| Boş | 10.293 | %98,9 |
| 34 ile başlayan numara | 9.642 | %2,1 |
| Başka biçimde numara | 4.330 | %3,7 |
| “Konut dışı ilan” (otel vb.) | 2.082 | %1,3 |
| “Muaf” | 284 | %25,7 |
İzin alanı boş olan ilanların %99'u en az 100 gece istiyor; 34 ile başlayan numarası olanlarda bu oran %2. Veri nedeni söylemiyor, ama tablo kanunla uyumlu: izin belgesi olmayan ilanlar, kapsam dışında kalacak kadar uzun kiralamaya geçmiş görünüyor.
Analiz kümesi
| Adım | Kalan ilan | Çıkan |
|---|---|---|
| Ham veri | 26.631 | |
| Fiyatı olmayanlar çıkarıldı | 23.943 | 2.688 |
| Uzun konaklama teklifleri (28+ gece) ayrıldı | 15.226 | 8.717 |
| Aynı ev sahibinin birebir aynı ilanları teke indirildi | 14.923 | 303 |
| Fiyatın en uç %1'i çıkarıldı | 14.774 | 149 |
Geriye 14.774 ilan ve 3.387 ev sahibi kaldı. Fiyat sınırları ₺1.137 ile ₺40.000.
Fiyat nasıl dağılıyor?
Kısa konaklamada medyan gecelik fiyat ₺4.300; ilanların yarısı ₺2.852 ile ₺6.499 arasında.
| Oda tipi | İlan | Medyan gecelik fiyat |
|---|---|---|
| Evin tamamı | 10.759 | ₺4.688 |
| Özel oda | 3.644 | ₺3.007 |
| Otel odası | 269 | ₺5.699 |
| Paylaşımlı oda | 102 | ₺2.729 |
İlanların %66'sı üç ilçede: Beyoğlu, Fatih ve Şişli. Fiyatta ise iki sayfiye ilçesi, Adalar ve Şile önde.
Geri kalanında ilçe farkının bir kısmı büyüklükten geliyor. Şişli'nin medyanı Kadıköy'den ₺1.524 yüksek, ama Şişli'de tipik ilan 4 kişilik, Kadıköy'de 2 kişilik. Kişi başına bakınca ikisi neredeyse aynı: ₺1.390 ve ₺1.405.
Fiyat modeli
Hedef, gecelik fiyatın logaritması. Fiyattan türeyen sütunları (tahmini gelir, teklif toplamı) modele almadım; alırsam model cevabı görmüş olur. Üç modeli karşılaştırdım ve hepsini aynı şekilde sınadım: ilanları ev sahibine göre gruplayıp beşe böldüm, böylece model her seferinde hiç görmediği ev sahiplerinin ilanlarını tahmin etti.
| Model | R² | Medyan hata | ±%25 içindeki tahminler |
|---|---|---|---|
| Taban: ilçe × oda tipinin medyanı | 0,10 | %35,3 | %37,2 |
| Doğrusal model (Ridge) | 0,49 | %26,1 | %48,0 |
| Gradyan artırmalı ağaçlar | 0,64 | %21,6 | %56,2 |
| Aynı model, rastgele bölmeyle | 0,76 | %16,1 | %67,5 |
- Taban neredeyse hiçbir şey açıklamıyor. “Bu ilçede bu oda tipinin medyanı” demek, fiyat farklarının yalnızca onda birini yakalıyor.
- En iyi model tahminlerin yarısında ±%22 içinde. Gradyan artırmalı ağaçlar doğrusal modelden belirgin iyi; ilişki doğrusal değil.
- Rastgele bölme modeli olduğundan iyi gösteriyor. Aynı model rastgele bölmeyle sınanınca medyan hata %16'ya iniyor. Aradaki fark başarı değil; model, eğitimde gördüğü ev sahibinin öbür ilanlarını tanıyor.
Model ortalamaya çekiyor: en ucuz beşte birlik dilimde fiyatı medyan %29 fazla, en pahalı dilimde %23 eksik tahmin ediyor. Ortadaki üç dilimde sapma küçük.
Model neye yaslanıyor?
Değişkenleri anlam gruplarına ayırıp her grubu test verisinde birlikte karıştırdım. Bir grup karıştırılınca R² ne kadar düşüyorsa, model o bilgiye o kadar yaslanıyor.
Büyüklük açık ara önde; konum ve olanaklar onu izliyor. Konaklama tipi ve ev sahibi bilgisi çok az şey ekliyor. Değişkenler birbiriyle ilişkili olduğu için bu sıralama modelin neye yaslandığını gösterir, nedensel etkiyi değil.
Konumun etkisi “merkeze yakınlık” kadar basit de değil. 2–4 kişilik evlerde Taksim'e uzaklık fiyatı sıralamıyor:
| Taksim'e uzaklık | İlan | Medyan gecelik fiyat |
|---|---|---|
| 0–2 km | 3.391 | ₺3.957 |
| 2–5 km | 1.646 | ₺3.710 |
| 5–10 km | 894 | ₺4.424 |
| 10–20 km | 677 | ₺3.996 |
| 20+ km | 552 | ₺3.774 |
Sınırlar
Sonuç
Bu çalışmada en çok işe yarayan şey bir model değil, bir sütunun ne ölçtüğünü sorgulamak oldu. Fiyat sütunu iki ayrı pazarı aynı adla tutuyordu; ayırmadan hesaplanan her ortalama ve kurulan her model yanlış çıkardı.
İkinci ders doğrulamadan geldi. Aynı model, sınama biçimine göre %16 ya da %22 hata veriyor. Doğru olan ikincisi, çünkü gerçek hayatta model hiç görmediği bir ev sahibinin ilanını fiyatlayacak.
Veri: Inside Airbnb, İstanbul, 30 Haziran 2026 (CC BY 4.0). Tutarlar Türk lirası.
A dataset is almost never ready for analysis in its first form. In this study I cleaned the raw data of 26,631 Airbnb listings in Istanbul, looked at how prices are distributed and built a model that predicts the nightly price. The most important finding came before the model, during cleaning.
Data
The data is the detailed listings file that Inside Airbnb publishes for Istanbul (CC BY 4.0). Each row is a listing: location, room type, capacity, amenities, host details, review scores and price.
- Get to know
the raw data - Clean
- Explore
- Model
and test
Every step is in a single Jupyter notebook, with its outputs: download the notebook (.ipynb). Its comments are in Turkish.
What was broken, and what I did
A first look turned up eight separate problems. None is hard on its own; the hard part is being able to write down what you did in each case and why.
| Problem | Example | What I did |
|---|---|---|
| Price is text with a misleading sign | $4,296.25 | Stripped the sign and the thousands separator. The raw quote writes the same amounts with ₺, so I treated the unit as Turkish lira. |
| Bathroom count is text | 1.5 shared baths, Half-bath | Split into a number and a “shared bath” flag. Missing values fell from 45.8% to 0.5%. |
| Amenities are a list in one cell | ["Wifi", "Kitchen", …] | There are 4,372 distinct amenities. I extracted their count and 11 selected ones into columns. |
| Property type is scattered | 90 distinct values | Reduced to five groups: apartment, hotel/guesthouse, house/villa, serviced residence, other. |
| Same value spelled differently | İstanbul, Turkey, Istanbul, Türkiye | Collapsed into one value. |
| Licence number is free text | 2022-34-1997, 12-3456 | Classified into five types by format. 149 listings have 12-3456 in place of a number. |
| Completely empty columns | 13 columns | Dropped. |
| Extreme values | 50 bedrooms, ₺1,706,000 per night | Capped rooms, beds and baths at 10; removed the most extreme 1% of prices. |
Missing values are not random
10% of prices and nearly half of the bed and bathroom counts are missing. Before filling anything I looked at where the gaps come from. Inside Airbnb records whether each listing's information comes from this scrape or a previous one; the gaps sit almost entirely in the second group.
| Where the record comes from | Listings | Price missing | Beds missing | Bathrooms missing |
|---|---|---|---|---|
| This scrape | 15,395 | 0.5% | 1.4% | 6.2% |
| Previous scrape | 11,236 | 23.2% | 100.0% | 100.0% |
Review scores follow the same logic: every listing without a score is a listing with no reviews at all.
Two different prices in one column
The price is calculated from a quote for the shortest stay the listing accepts. The quote's check-in and check-out dates are in the data, and the number of nights between them showed something I did not expect.
40% of listings require at least 100 nights. For those, the price was quoted for a 100-night stay and divided by the nights; it is a monthly rent expressed per day, not a nightly rate. Both groups sit in the same column but cannot be compared.
Had I used the column as it came, I would have said “the median nightly price in Istanbul is ₺2,853”; that figure is a blend of two markets. The rest of the analysis uses listings whose quote covers fewer than 28 nights.
Where does 100 nights come from?
Law no. 7464, in force in Türkiye since January 1, 2024, requires a permit for residential rentals of 100 days or less (Airbnb's explanation, in Turkish). Looking at the licence field of the listings, the pattern is very clear.
| What the licence field says | Listings | Requiring 100+ nights |
|---|---|---|
| Empty | 10,293 | 98.9% |
| Number starting with 34 | 9,642 | 2.1% |
| Number in another format | 4,330 | 3.7% |
| “Non-real estate listing” (hotels etc.) | 2,082 | 1.3% |
| “Exempt” | 284 | 25.7% |
99% of listings with an empty licence field require at least 100 nights; among those with a number starting with 34 the share is 2%. The data does not state the reason, but the table is consistent with the law: listings without a permit appear to have moved to rentals long enough to fall outside its scope.
The analysis set
| Step | Listings left | Removed |
|---|---|---|
| Raw data | 26,631 | |
| Listings without a price removed | 23,943 | 2,688 |
| Long-stay quotes (28+ nights) set aside | 15,226 | 8,717 |
| Identical listings from the same host collapsed | 14,923 | 303 |
| Most extreme 1% of prices removed | 14,774 | 149 |
That leaves 14,774 listings and 3,387 hosts. Prices range from ₺1,137 to ₺40,000.
How prices are distributed
For short stays the median nightly price is ₺4,300; half of the listings fall between ₺2,852 and ₺6,499.
| Room type | Listings | Median nightly price |
|---|---|---|
| Entire home | 10,759 | ₺4,688 |
| Private room | 3,644 | ₺3,007 |
| Hotel room | 269 | ₺5,699 |
| Shared room | 102 | ₺2,729 |
66% of listings are in three districts: Beyoğlu, Fatih and Şişli. On price, two resort districts lead: Adalar (the Princes' Islands) and Şile.
Elsewhere, part of the gap between districts is really a difference in size. Şişli's median is ₺1,524 above Kadıköy's, but the typical listing sleeps 4 in Şişli and 2 in Kadıköy. Per guest the two are nearly identical: ₺1,390 and ₺1,405.
The price model
The target is the logarithm of the nightly price. I kept columns derived from the price (estimated revenue, the quote total) out of the model; with them the model would have seen the answer. I compared three models and tested all of them the same way: listings were grouped by host and split into five folds, so each time the model predicted listings from hosts it had never seen.
| Model | R² | Median error | Predictions within ±25% |
|---|---|---|---|
| Baseline: median of district × room type | 0.10 | 35.3% | 37.2% |
| Linear model (Ridge) | 0.49 | 26.1% | 48.0% |
| Gradient-boosted trees | 0.64 | 21.6% | 56.2% |
| Same model, random split | 0.76 | 16.1% | 67.5% |
- The baseline explains almost nothing. Saying “the median for this room type in this district” captures only a tenth of the variation in price.
- The best model is within ±22% for half of its predictions. Gradient-boosted trees clearly beat the linear model; the relationship is not linear.
- A random split flatters the model. Tested with a random split, the same model's median error drops to 16%. The gap is not skill; the model recognizes other listings of a host it saw in training.
The model pulls toward the middle: in the cheapest fifth it predicts a median 29% too high, in the most expensive fifth 23% too low. In the three middle fifths the bias is small.
What the model relies on
I sorted the variables into meaningful groups and shuffled each group together on the test data. The more R² falls when a group is shuffled, the more the model relies on that information.
Size leads by far; location and amenities follow. Type of stay and host details add very little. Because the variables are correlated, this ranking shows what the model relies on, not a causal effect.
Nor is the effect of location as simple as “close to the centre”. For homes sleeping 2–4, distance to Taksim does not order prices:
| Distance to Taksim | Listings | Median nightly price |
|---|---|---|
| 0–2 km | 3,391 | ₺3,957 |
| 2–5 km | 1,646 | ₺3,710 |
| 5–10 km | 894 | ₺4,424 |
| 10–20 km | 677 | ₺3,996 |
| 20+ km | 552 | ₺3,774 |
Limitations
Conclusion
What helped most in this study was questioning what a column measures, more than any model. The price column held two different markets under one name; every average computed and every model built without separating them would have been wrong.
The second lesson came from validation. The same model shows an error of 16% or 22% depending on how it is tested. The second is the right one, because in real use the model will price a listing from a host it has never seen.
Data: Inside Airbnb, Istanbul, June 30, 2026 (CC BY 4.0). Amounts are in Turkish lira.