E-ticaret Uygulamalarında Kullanıcılar Neden Şikayet Ediyor?Why Do Users Complain About E-commerce Apps?

Trendyol, Hepsiburada ve n11'in Google Play'deki 10.457 yorumunu duygu ve konu açısından inceledim. Yıldız puanının tek başına göstermediği şikayetler ortaya çıktı.I analyzed 10,457 Google Play reviews of Trendyol, Hepsiburada and n11 for sentiment and topic. They surface complaints the star rating alone does not show.

  • Python
  • NLP
  • BERT
  • scikit-learn

Bir uygulamanın mağaza puanı 4,2 ise kullanıcılar memnun mu? Kısmen. Yıldız ortalaması, insanların neden bir yıldız verdiğini ya da beş yıldız verip altına ne yazdığını söylemez. Bu çalışmada üç büyük e-ticaret uygulamasının Google Play yorumlarını metin üzerinden inceledim.

Yorum10.4572 Ağustos – 4 Ekim 2026
Uygulama3Trendyol, Hepsiburada, n11
Yıldızla uyum%92,3macro-F1 0,89

Veri

Google Play'den Trendyol, Hepsiburada ve n11 uygulamalarının Türkçe yorumlarını çektim. Karşılaştırma adil olsun diye üç uygulamayı da aynı döneme kırptım. Kullanıcı adı ve profil fotoğrafı toplamadım.

UygulamaYorumOrtalama yıldız
Hepsiburada1.3113,48
Trendyol7.8194,19
n111.3274,49

Trendyol'un yorum hacmi diğer ikisinin yaklaşık altı katı. Bu yüzden aşağıdaki bütün rakamlar adet değil, her uygulamanın kendi yorumları içindeki oran.

Yöntem

Analiz adımları
  1. Yorumları
    çek
  2. Duyguyu
    skorla
  3. Modeli veriye
    uyarla
  4. Konuları
    etiketle

Hazır model yetmedi

Duygu için Türkçe BERT tabanlı hazır bir model kullandım (BERTurk, pozitif / negatif). Sonucu yıldız puanıyla karşılaştırınca sorun ortaya çıktı: model 4-5 yıldızlı yorumların %17'sini olumsuz sayıyordu, “memnunum” yazanlar dahil. Model film ve ürün yorumlarıyla eğitilmiş; uygulama yorumlarının dili farklı.

Bunun üzerine modeli bu veri setine uyarladım. BERT skorunu ve yorum metnini (karakter n-gram TF-IDF) bir lojistik regresyona verdim, etiket olarak yıldız puanını kullandım: 1-2 yıldız olumsuz, 4-5 yıldız olumlu, 3 yıldız hariç. Skorlar 5 katlı çapraz doğrulamadan geliyor; yani hiçbir yorum, kendi yıldızını görmüş bir modelden skor almıyor.

Yıldız puanıyla uyum (10.111 yorum)
  • Hazır model%85,3
  • Uyarlanmış model%92,3

Macro-F1 de 0,81'den 0,89'a çıktı. Asıl kazanç olumlu yorumlarda: yanlışlıkla olumsuz sayılan 4-5 yıldızlı yorum sayısı 1.379'dan 629'a indi.

Konu etiketleme

Olumsuz yorumları on konuya ayırmak için kelime köklerinden oluşan bir sözlük kurdum. Sözlük, Türkçe karakter kullanılmadan yazılan yorumları da yakalıyor (“urun”, “musteri”). Bir yorum birden fazla konuya girebiliyor; “uygulama” ve “ürün” gibi her yorumda geçen genel kelimeler sözlükte yok.

Hepsiburada yorumlarının neredeyse yarısı olumsuz

Olumsuz yorum oranı
  • Hepsiburada%44,5
  • Trendyol%23,9
  • n11%15,4

Sıralama yıldız ortalamasıyla tutarlı: en düşük puanlı uygulama en çok şikayet alan. Ama aradaki fark yıldızda göründüğünden büyük. Hepsiburada ile n11 arasında ortalama puan farkı bir yıldız kadarken olumsuz yorum oranı neredeyse üç katı.

Beş yıldız verip şikayet yazanlar

4-5 yıldızlı 8.163 yorumun 629'u (%7,7) metin olarak olumsuz çıktı.

Hepsiburada%12,74-5 yıldızlı yorumlar içinde
Trendyol%7,64-5 yıldızlı yorumlar içinde
n11%4,94-5 yıldızlı yorumlar içinde

Bir kısmı nedenini açıkça yazıyor: yorumu üstte görünsün diye beş yıldız veriyor, altına uygulamanın günlerdir açılmadığını ekliyor. Bir kısmı “güzel ama…” diye başlayan istekler: kargo hızlansın, bildirimler azalsın, arama düzelsin. Bu grubun bir bölümü model hatası, yine de tablo net.

Çıkarım: Yalnızca yıldız ortalamasını izleyen bir ekip bu şikayetleri hiç görmez. Yıldız, memnuniyetsizliği olduğundan düşük gösteriyor.

En büyük şikayet uygulamanın kendisi

Olumsuz yorumların yüzde kaçı bu konudan bahsediyor
KonuHepsiburadaTrendyoln11
Uygulama hataları%28,3%24,9%30,2
Ürün ve satıcı%17,8%23,0%21,0
Kargo ve teslimat%14,8%18,4%17,1
Fiyat ve kampanya%11,1%11,4%20,0
İade ve iptal%16,1%12,2%13,7
Müşteri hizmetleri%10,3%12,6%14,1
Ödeme ve üyelik%10,6%4,3%9,3
Reklam ve bildirim%4,6%8,7%8,3
Güven ve dolandırıcılık%6,2%6,9%7,3
Arama ve ürün bulma%5,1%5,9%6,3

Uygulama hataları (açılmama, donma, giriş yapamama) üç markada da ilk sırada ve olumsuz yorumların dörtte birinden fazlasında geçiyor. Mağaza yorumu olduğu için beklenen bir sonuç, yine de kargo ve iadenin önünde olması dikkat çekici.

Markalar arasındaki farklar daha ilginç:

  • n11'de fiyat. Olumsuz yorumların %20'si fiyattan bahsediyor; diğer iki uygulamada bu oran %11.
  • Hepsiburada'da iade ve ödeme. İade ve iptal %16, ödeme ve üyelik %11. Trendyol'da ödeme ve üyelik yalnızca %4.
  • Trendyol'da ürün, satıcı ve kargo. Ürün ve satıcı %23, kargo %18. Reklam ve bildirim şikayeti de Hepsiburada'nın yaklaşık iki katı.

Haftadan haftaya

Haftalık olumsuz yorum oranı
  • Hepsiburada
  • Trendyol
  • n11
%0%20%40%603 Ağu17 Ağu31 Ağu14 Eyl28 EylHepsiburada %40Trendyol %25n11 %12

Hafta başlangıcına göre; her hafta en az 110 yorum

Haftalık değerleri tablo olarak göster
HaftaHepsiburadaTrendyoln11
3 Ağu%42,4%17,0%21,6
10 Ağu%51,3%19,9%19,1
17 Ağu%46,0%22,2%19,6
24 Ağu%40,3%28,0%9,9
31 Ağu%42,5%22,1%12,8
7 Eyl%42,7%26,0%18,8
14 Eyl%44,6%25,5%8,3
21 Eyl%48,2%32,2%17,1
28 Eyl%39,7%25,1%12,1

Hepsiburada dönem boyunca %40 ile %51 arasında kaldı. Trendyol'da olumsuz oran ağustos başındaki %17'den eylül sonunda %32'ye kadar çıktı. n11 düşük ama dalgalı; haftalık yorum sayısı 150 civarında olduğu için tek haftalık iniş çıkışlara fazla anlam yüklememek gerekiyor.

Sınırlar

Yıldız kusursuz bir etiket değilYüksek yıldızlı şikayetler bunun kanıtı. Doğruluk rakamı “yıldızla uyum” olarak okunmalı.
Model ikili çalışıyorNötr sınıfı yok. “Güzel ama kargo yavaş” gibi karışık yorumlar çoğunlukla olumsuz sayılıyor.
Konu sözlüğü kural tabanlıOlumsuz yorumların %23'üne konu atanamadı; bir kısmı “berbat” gibi tek kelimelik yorumlar.
n11 örneklemi küçük205 olumsuz yorum var; konu payları yaklaşık ±6 puan oynayabilir.
Yalnızca Google PlayUygulamayı kullanan ve yorum yazanların sesi. Tüm müşterileri temsil etmez.
İki aylık dönemKampanya ve mevsim etkisini ayırmak için daha uzun bir seri gerekir.

Sonuç

Yorum metni, yıldız puanının anlatmadığı iki şeyi gösterdi: şikayetin konusu ve yüksek puanın arkasına saklanan memnuniyetsizlik. Üç uygulamada da en çok konuşulan sorun teslimat ya da fiyat değil, uygulamanın çalışmaması. Marka bazında ise öncelikler ayrışıyor: n11 için fiyat algısı, Hepsiburada için iade ve ödeme süreci, Trendyol için satıcı kalitesi ve kargo.

Yöntem tarafında çıkardığım ders, hazır bir duygu modelini doğrulamadan kullanmamak. Aynı model, veriye uyarlanmadan önce beş yıldızlı “memnunum” yorumlarını olumsuz sayıyordu.

Bu çalışma bağımsızdır; adı geçen markalarla bir bağlantısı yoktur ve yalnızca herkese açık yorumlara dayanır.

If an app is rated 4.2 in the store, are its users happy? Partly. The average rating does not tell you why people give one star, or what they write underneath a five-star rating. In this study I analyzed the text of Google Play reviews for three major Turkish e-commerce apps.

Reviews10,457August 2 – October 4, 2026
Apps3Trendyol, Hepsiburada, n11
Agreement with stars92.3%macro-F1 0.89

Data

I collected Turkish-language reviews of the Trendyol, Hepsiburada and n11 apps from Google Play. To keep the comparison fair, all three apps are trimmed to the same period. User names and profile photos were not collected.

AppReviewsAverage rating
Hepsiburada1,3113.48
Trendyol7,8194.19
n111,3274.49

Trendyol has roughly six times the review volume of the other two. For that reason every figure below is a share within each app's own reviews, not a count.

Method

Analysis steps
  1. Collect
    reviews
  2. Score
    sentiment
  3. Adapt model
    to the data
  4. Tag
    topics

The off-the-shelf model was not enough

For sentiment I used a ready-made Turkish BERT model (BERTurk, positive / negative). Comparing its output with star ratings exposed a problem: it labeled 17% of 4-5 star reviews as negative, including ones that simply say “memnunum” (“I'm satisfied”). The model was trained on movie and product reviews; app reviews are written differently.

So I adapted the model to this dataset. The BERT score and the review text (character n-gram TF-IDF) feed a logistic regression, with star ratings as labels: 1-2 stars negative, 4-5 stars positive, 3 stars excluded. Scores come from 5-fold cross-validation, so no review is scored by a model that has seen its own rating.

Agreement with star ratings (10,111 reviews)
  • Off-the-shelf model85.3%
  • Adapted model92.3%

Macro-F1 rose from 0.81 to 0.89. Most of the gain is on positive reviews: the number of 4-5 star reviews wrongly labeled negative fell from 1,379 to 629.

Topic tagging

To sort negative reviews into ten topics I built a dictionary of word stems. It also catches reviews typed without Turkish characters (“urun”, “musteri”). A review can fall under more than one topic, and generic words that appear in almost every review, such as “app” and “product”, are left out.

Nearly half of Hepsiburada's reviews are negative

Share of negative reviews
  • Hepsiburada44.5%
  • Trendyol23.9%
  • n1115.4%

The ranking matches the average ratings: the lowest-rated app gets the most complaints. But the gap is wider than the stars suggest. Hepsiburada and n11 are about one star apart on average, yet Hepsiburada's share of negative reviews is almost three times higher.

Five stars, followed by a complaint

Of 8,163 reviews rated 4-5 stars, 629 (7.7%) are negative in their text.

Hepsiburada12.7%of 4-5 star reviews
Trendyol7.6%of 4-5 star reviews
n114.9%of 4-5 star reviews

Some say why outright: they give five stars so the review shows up at the top, then add that the app has not opened for days. Others are requests that start with “good, but…”: faster delivery, fewer notifications, better search. Part of this group is model error, but the picture is still clear.

Takeaway: A team that only tracks the average rating never sees these complaints. Stars understate dissatisfaction.

The biggest complaint is the app itself

Share of negative reviews that mention each topic
TopicHepsiburadaTrendyoln11
App errors28.3%24.9%30.2%
Product & seller17.8%23.0%21.0%
Shipping & delivery14.8%18.4%17.1%
Price & promotions11.1%11.4%20.0%
Returns & cancellations16.1%12.2%13.7%
Customer service10.3%12.6%14.1%
Payment & membership10.6%4.3%9.3%
Ads & notifications4.6%8.7%8.3%
Trust & fraud6.2%6.9%7.3%
Search & availability5.1%5.9%6.3%

App errors (not opening, freezing, failing to log in) rank first for all three brands and appear in more than a quarter of negative reviews. That is partly expected from store reviews, but it is notable that they outrank shipping and returns.

The differences between brands are more interesting:

  • Price at n11. 20% of negative reviews mention price, against 11% for the other two apps.
  • Returns and payment at Hepsiburada. Returns and cancellations 16%, payment and membership 11%. At Trendyol, payment and membership is only 4%.
  • Products, sellers and shipping at Trendyol. Product and seller 23%, shipping 18%. Complaints about ads and notifications are also about twice Hepsiburada's.

Week by week

Weekly share of negative reviews
  • Hepsiburada
  • Trendyol
  • n11
0%20%40%60%Aug 3Aug 17Aug 31Sep 14Sep 28Hepsiburada 40%Trendyol 25%n11 12%

By week start; at least 110 reviews per brand each week

Show weekly values as a table
WeekHepsiburadaTrendyoln11
Aug 342.4%17.0%21.6%
Aug 1051.3%19.9%19.1%
Aug 1746.0%22.2%19.6%
Aug 2440.3%28.0%9.9%
Aug 3142.5%22.1%12.8%
Sep 742.7%26.0%18.8%
Sep 1444.6%25.5%8.3%
Sep 2148.2%32.2%17.1%
Sep 2839.7%25.1%12.1%

Hepsiburada stayed between 40% and 51% throughout. Trendyol's negative share climbed from 17% in early August to as high as 32% in late September. n11 is low but volatile; with around 150 reviews a week, single-week swings should not be over-read.

Limitations

Stars are not a perfect labelHigh-star complaints prove it. The accuracy figure should be read as “agreement with stars”.
The model is binaryThere is no neutral class. Mixed reviews such as “good, but delivery is slow” are mostly counted as negative.
Topic tagging is rule-based23% of negative reviews got no topic; some are one-word reviews such as “terrible”.
n11's sample is small205 negative reviews; topic shares can move by about ±6 points.
Google Play onlyThe voice of people who use the app and write reviews, not of all customers.
A two-month windowSeparating campaign and seasonal effects needs a longer series.

Conclusion

Review text showed two things the star rating does not: what the complaint is about, and the dissatisfaction hiding behind high ratings. Across all three apps the most discussed problem is not delivery or price but the app failing to work. By brand, priorities diverge: price perception for n11, returns and payment for Hepsiburada, seller quality and shipping for Trendyol.

On the method side, the lesson is not to use a ready-made sentiment model without validating it. Before adaptation, the same model was labeling five-star “I'm satisfied” reviews as negative.

This is an independent study. It has no affiliation with the brands named and relies only on publicly available reviews.