E-ticaret Uygulamalarında Kullanıcılar Neden Şikayet Ediyor?Why Do Users Complain About E-commerce Apps?
Trendyol, Hepsiburada ve n11'in Google Play'deki 10.457 yorumunu duygu ve konu açısından inceledim. Yıldız puanının tek başına göstermediği şikayetler ortaya çıktı.I analyzed 10,457 Google Play reviews of Trendyol, Hepsiburada and n11 for sentiment and topic. They surface complaints the star rating alone does not show.
- Python
- NLP
- BERT
- scikit-learn
Bir uygulamanın mağaza puanı 4,2 ise kullanıcılar memnun mu? Kısmen. Yıldız ortalaması, insanların neden bir yıldız verdiğini ya da beş yıldız verip altına ne yazdığını söylemez. Bu çalışmada üç büyük e-ticaret uygulamasının Google Play yorumlarını metin üzerinden inceledim.
Veri
Google Play'den Trendyol, Hepsiburada ve n11 uygulamalarının Türkçe yorumlarını çektim. Karşılaştırma adil olsun diye üç uygulamayı da aynı döneme kırptım. Kullanıcı adı ve profil fotoğrafı toplamadım.
| Uygulama | Yorum | Ortalama yıldız |
|---|---|---|
| Hepsiburada | 1.311 | 3,48 |
| Trendyol | 7.819 | 4,19 |
| n11 | 1.327 | 4,49 |
Trendyol'un yorum hacmi diğer ikisinin yaklaşık altı katı. Bu yüzden aşağıdaki bütün rakamlar adet değil, her uygulamanın kendi yorumları içindeki oran.
Yöntem
- Yorumları
çek - Duyguyu
skorla - Modeli veriye
uyarla - Konuları
etiketle
Hazır model yetmedi
Duygu için Türkçe BERT tabanlı hazır bir model kullandım (BERTurk, pozitif / negatif). Sonucu yıldız puanıyla karşılaştırınca sorun ortaya çıktı: model 4-5 yıldızlı yorumların %17'sini olumsuz sayıyordu, “memnunum” yazanlar dahil. Model film ve ürün yorumlarıyla eğitilmiş; uygulama yorumlarının dili farklı.
Bunun üzerine modeli bu veri setine uyarladım. BERT skorunu ve yorum metnini (karakter n-gram TF-IDF) bir lojistik regresyona verdim, etiket olarak yıldız puanını kullandım: 1-2 yıldız olumsuz, 4-5 yıldız olumlu, 3 yıldız hariç. Skorlar 5 katlı çapraz doğrulamadan geliyor; yani hiçbir yorum, kendi yıldızını görmüş bir modelden skor almıyor.
Macro-F1 de 0,81'den 0,89'a çıktı. Asıl kazanç olumlu yorumlarda: yanlışlıkla olumsuz sayılan 4-5 yıldızlı yorum sayısı 1.379'dan 629'a indi.
Konu etiketleme
Olumsuz yorumları on konuya ayırmak için kelime köklerinden oluşan bir sözlük kurdum. Sözlük, Türkçe karakter kullanılmadan yazılan yorumları da yakalıyor (“urun”, “musteri”). Bir yorum birden fazla konuya girebiliyor; “uygulama” ve “ürün” gibi her yorumda geçen genel kelimeler sözlükte yok.
Hepsiburada yorumlarının neredeyse yarısı olumsuz
Sıralama yıldız ortalamasıyla tutarlı: en düşük puanlı uygulama en çok şikayet alan. Ama aradaki fark yıldızda göründüğünden büyük. Hepsiburada ile n11 arasında ortalama puan farkı bir yıldız kadarken olumsuz yorum oranı neredeyse üç katı.
Beş yıldız verip şikayet yazanlar
4-5 yıldızlı 8.163 yorumun 629'u (%7,7) metin olarak olumsuz çıktı.
Bir kısmı nedenini açıkça yazıyor: yorumu üstte görünsün diye beş yıldız veriyor, altına uygulamanın günlerdir açılmadığını ekliyor. Bir kısmı “güzel ama…” diye başlayan istekler: kargo hızlansın, bildirimler azalsın, arama düzelsin. Bu grubun bir bölümü model hatası, yine de tablo net.
En büyük şikayet uygulamanın kendisi
| Konu | Hepsiburada | Trendyol | n11 |
|---|---|---|---|
| Uygulama hataları | %28,3 | %24,9 | %30,2 |
| Ürün ve satıcı | %17,8 | %23,0 | %21,0 |
| Kargo ve teslimat | %14,8 | %18,4 | %17,1 |
| Fiyat ve kampanya | %11,1 | %11,4 | %20,0 |
| İade ve iptal | %16,1 | %12,2 | %13,7 |
| Müşteri hizmetleri | %10,3 | %12,6 | %14,1 |
| Ödeme ve üyelik | %10,6 | %4,3 | %9,3 |
| Reklam ve bildirim | %4,6 | %8,7 | %8,3 |
| Güven ve dolandırıcılık | %6,2 | %6,9 | %7,3 |
| Arama ve ürün bulma | %5,1 | %5,9 | %6,3 |
Uygulama hataları (açılmama, donma, giriş yapamama) üç markada da ilk sırada ve olumsuz yorumların dörtte birinden fazlasında geçiyor. Mağaza yorumu olduğu için beklenen bir sonuç, yine de kargo ve iadenin önünde olması dikkat çekici.
Markalar arasındaki farklar daha ilginç:
- n11'de fiyat. Olumsuz yorumların %20'si fiyattan bahsediyor; diğer iki uygulamada bu oran %11.
- Hepsiburada'da iade ve ödeme. İade ve iptal %16, ödeme ve üyelik %11. Trendyol'da ödeme ve üyelik yalnızca %4.
- Trendyol'da ürün, satıcı ve kargo. Ürün ve satıcı %23, kargo %18. Reklam ve bildirim şikayeti de Hepsiburada'nın yaklaşık iki katı.
Haftadan haftaya
- Hepsiburada
- Trendyol
- n11
Hafta başlangıcına göre; her hafta en az 110 yorum
Haftalık değerleri tablo olarak göster
| Hafta | Hepsiburada | Trendyol | n11 |
|---|---|---|---|
| 3 Ağu | %42,4 | %17,0 | %21,6 |
| 10 Ağu | %51,3 | %19,9 | %19,1 |
| 17 Ağu | %46,0 | %22,2 | %19,6 |
| 24 Ağu | %40,3 | %28,0 | %9,9 |
| 31 Ağu | %42,5 | %22,1 | %12,8 |
| 7 Eyl | %42,7 | %26,0 | %18,8 |
| 14 Eyl | %44,6 | %25,5 | %8,3 |
| 21 Eyl | %48,2 | %32,2 | %17,1 |
| 28 Eyl | %39,7 | %25,1 | %12,1 |
Hepsiburada dönem boyunca %40 ile %51 arasında kaldı. Trendyol'da olumsuz oran ağustos başındaki %17'den eylül sonunda %32'ye kadar çıktı. n11 düşük ama dalgalı; haftalık yorum sayısı 150 civarında olduğu için tek haftalık iniş çıkışlara fazla anlam yüklememek gerekiyor.
Sınırlar
Sonuç
Yorum metni, yıldız puanının anlatmadığı iki şeyi gösterdi: şikayetin konusu ve yüksek puanın arkasına saklanan memnuniyetsizlik. Üç uygulamada da en çok konuşulan sorun teslimat ya da fiyat değil, uygulamanın çalışmaması. Marka bazında ise öncelikler ayrışıyor: n11 için fiyat algısı, Hepsiburada için iade ve ödeme süreci, Trendyol için satıcı kalitesi ve kargo.
Yöntem tarafında çıkardığım ders, hazır bir duygu modelini doğrulamadan kullanmamak. Aynı model, veriye uyarlanmadan önce beş yıldızlı “memnunum” yorumlarını olumsuz sayıyordu.
Bu çalışma bağımsızdır; adı geçen markalarla bir bağlantısı yoktur ve yalnızca herkese açık yorumlara dayanır.
If an app is rated 4.2 in the store, are its users happy? Partly. The average rating does not tell you why people give one star, or what they write underneath a five-star rating. In this study I analyzed the text of Google Play reviews for three major Turkish e-commerce apps.
Data
I collected Turkish-language reviews of the Trendyol, Hepsiburada and n11 apps from Google Play. To keep the comparison fair, all three apps are trimmed to the same period. User names and profile photos were not collected.
| App | Reviews | Average rating |
|---|---|---|
| Hepsiburada | 1,311 | 3.48 |
| Trendyol | 7,819 | 4.19 |
| n11 | 1,327 | 4.49 |
Trendyol has roughly six times the review volume of the other two. For that reason every figure below is a share within each app's own reviews, not a count.
Method
- Collect
reviews - Score
sentiment - Adapt model
to the data - Tag
topics
The off-the-shelf model was not enough
For sentiment I used a ready-made Turkish BERT model (BERTurk, positive / negative). Comparing its output with star ratings exposed a problem: it labeled 17% of 4-5 star reviews as negative, including ones that simply say “memnunum” (“I'm satisfied”). The model was trained on movie and product reviews; app reviews are written differently.
So I adapted the model to this dataset. The BERT score and the review text (character n-gram TF-IDF) feed a logistic regression, with star ratings as labels: 1-2 stars negative, 4-5 stars positive, 3 stars excluded. Scores come from 5-fold cross-validation, so no review is scored by a model that has seen its own rating.
Macro-F1 rose from 0.81 to 0.89. Most of the gain is on positive reviews: the number of 4-5 star reviews wrongly labeled negative fell from 1,379 to 629.
Topic tagging
To sort negative reviews into ten topics I built a dictionary of word stems. It also catches reviews typed without Turkish characters (“urun”, “musteri”). A review can fall under more than one topic, and generic words that appear in almost every review, such as “app” and “product”, are left out.
Nearly half of Hepsiburada's reviews are negative
The ranking matches the average ratings: the lowest-rated app gets the most complaints. But the gap is wider than the stars suggest. Hepsiburada and n11 are about one star apart on average, yet Hepsiburada's share of negative reviews is almost three times higher.
Five stars, followed by a complaint
Of 8,163 reviews rated 4-5 stars, 629 (7.7%) are negative in their text.
Some say why outright: they give five stars so the review shows up at the top, then add that the app has not opened for days. Others are requests that start with “good, but…”: faster delivery, fewer notifications, better search. Part of this group is model error, but the picture is still clear.
The biggest complaint is the app itself
| Topic | Hepsiburada | Trendyol | n11 |
|---|---|---|---|
| App errors | 28.3% | 24.9% | 30.2% |
| Product & seller | 17.8% | 23.0% | 21.0% |
| Shipping & delivery | 14.8% | 18.4% | 17.1% |
| Price & promotions | 11.1% | 11.4% | 20.0% |
| Returns & cancellations | 16.1% | 12.2% | 13.7% |
| Customer service | 10.3% | 12.6% | 14.1% |
| Payment & membership | 10.6% | 4.3% | 9.3% |
| Ads & notifications | 4.6% | 8.7% | 8.3% |
| Trust & fraud | 6.2% | 6.9% | 7.3% |
| Search & availability | 5.1% | 5.9% | 6.3% |
App errors (not opening, freezing, failing to log in) rank first for all three brands and appear in more than a quarter of negative reviews. That is partly expected from store reviews, but it is notable that they outrank shipping and returns.
The differences between brands are more interesting:
- Price at n11. 20% of negative reviews mention price, against 11% for the other two apps.
- Returns and payment at Hepsiburada. Returns and cancellations 16%, payment and membership 11%. At Trendyol, payment and membership is only 4%.
- Products, sellers and shipping at Trendyol. Product and seller 23%, shipping 18%. Complaints about ads and notifications are also about twice Hepsiburada's.
Week by week
- Hepsiburada
- Trendyol
- n11
By week start; at least 110 reviews per brand each week
Show weekly values as a table
| Week | Hepsiburada | Trendyol | n11 |
|---|---|---|---|
| Aug 3 | 42.4% | 17.0% | 21.6% |
| Aug 10 | 51.3% | 19.9% | 19.1% |
| Aug 17 | 46.0% | 22.2% | 19.6% |
| Aug 24 | 40.3% | 28.0% | 9.9% |
| Aug 31 | 42.5% | 22.1% | 12.8% |
| Sep 7 | 42.7% | 26.0% | 18.8% |
| Sep 14 | 44.6% | 25.5% | 8.3% |
| Sep 21 | 48.2% | 32.2% | 17.1% |
| Sep 28 | 39.7% | 25.1% | 12.1% |
Hepsiburada stayed between 40% and 51% throughout. Trendyol's negative share climbed from 17% in early August to as high as 32% in late September. n11 is low but volatile; with around 150 reviews a week, single-week swings should not be over-read.
Limitations
Conclusion
Review text showed two things the star rating does not: what the complaint is about, and the dissatisfaction hiding behind high ratings. Across all three apps the most discussed problem is not delivery or price but the app failing to work. By brand, priorities diverge: price perception for n11, returns and payment for Hepsiburada, seller quality and shipping for Trendyol.
On the method side, the lesson is not to use a ready-made sentiment model without validating it. Before adaptation, the same model was labeling five-star “I'm satisfied” reviews as negative.
This is an independent study. It has no affiliation with the brands named and relies only on publicly available reviews.