Bir YouTube Yorumu Kimin Hakkında? Havayolu Videolarında Hedefli Duygu AnaliziWho Is a YouTube Comment About? Target-Aware Sentiment Analysis on Airline Videos
THY, Pegasus ve AJet videolarının altındaki 6.659 Türkçe yorumu duygu, hedef ve konuya göre etiketledim. Yorumların yalnızca üçte biri markayla ilgili çıktı.I labelled 6,659 Turkish comments under THY, Pegasus and AJet videos by sentiment, target and topic. Only a third of them turned out to be about the brand.
- Python
- NLP
- LLM
- Scikit-learn
Duygu analizinde genellikle tek soru sorulur: yorum olumlu mu, olumsuz mu? Bu çalışmada ikinci bir soru ekledim: yorum kimin hakkında? THY, Pegasus ve AJet ile ilgili YouTube videolarının altındaki 6.659 Türkçe yorumu duygu, hedef ve konuya göre etiketledim. İkinci soru, ilk sorunun cevabını baştan sona değiştirdi.
Veri
Yorumları YouTube'dan, yazdığım bir Python scriptiyle topladım. Her marka için aynı beş nötr arama kalıbını kullandım (“… uçuş deneyimi”, “… inceleme” gibi). “Rötar” ya da “şikâyet” gibi kelimelerle arasaydım sonucu baştan olumsuza çekmiş olurdum. Başlığında markanın adı geçen videoları aldım, her videodan en yeni 300 yoruma kadar. Yorum yazanların adlarını toplamadım.
- Topla
- Ayıkla
- Etiketle
- Analiz et
- Ucuz modelle
karşılaştır
| Adım | Kalan yorum |
|---|---|
| Toplanan yorum | 13.151 |
| Tekrarlar atıldı | 13.126 |
| Türkçe olanlar | 9.614 |
| Markayla ilgisiz videolar çıkarıldı | 8.760 |
| Yanıtlar çıkarıldı (analiz kümesi) | 6.659 |
Yorumların %27'si Türkçe değildi; dili otomatik bir dil tanıyıcıyla belirledim. Yanıtları çıkardım, çünkü yanıt başka bir yorumcuya yazılır ve üstündeki yorum olmadan doğru okunamaz.
Videolar dört türe ayrılıyor
Aramalardan 115 video geldi. Başlıklarına bakarak her birini elle bir türe atadım; 8 tanesi markayla ilgisizdi (aynı adı taşıyan bir şarkı, şirketin voleybol takımının bir maçı gibi) ve çıkarıldı.
| Marka | Reklam ve marka filmi | Haber ve analiz | Uçuş videosu | Kaza anlatısı | Toplam | Video |
|---|---|---|---|---|---|---|
| THY | 2.901 | 380 | 209 | 502 | 3.992 | 36 |
| Pegasus | 375 | 596 | 98 | 974 | 2.043 | 31 |
| AJet | 82 | 361 | 181 | – | 624 | 27 |
Dağılım dengesiz: THY yorumlarının %73'ü reklam filmlerinin, Pegasus yorumlarının %48'i kaza anlatılarının altında. Bu dengesizliği ben yaratmadım, arama sonuçları böyle geldi. Aşağıdaki en önemli bulgu da buradan çıkıyor.
Etiketleme
Her yoruma üç etiket verildi:
- Hedef: Yorum markayla mı, videonun kendisiyle mi (reklam, müzik, anlatıcı), yoksa başka bir şeyle mi (siyaset, dua ve taziye, başka şirketler, sohbet) ilgili?
- Duygu: Yazanın o hedefe karşı tutumu olumlu mu, olumsuz mu, nötr mü?
- Konu: Hedef markaysa hangi konuda? Güvenlik, yönetim, fiyat, operasyon, hizmet, konfor ya da genel.
Etiketleri bir dil modeline (Claude) verdirdim. Model her yorumu, altına yazıldığı videonun başlığı ve türüyle birlikte okudu ve yazılı bir kılavuza göre karar verdi. Kelime listesi ya da kural kullanılmadı; “harika, yine üç saat rötar” gibi bir yorumun olumsuz sayılması buna bağlı.
Etiketlere ne kadar güvenilebileceğini görmek için rastgele 400 yorumu, ilk etiketleri görmeyen ayrı bir geçişte yeniden etiketlettim.
| Etiket | İki geçişin uyumu | Cohen kappa |
|---|---|---|
| Duygu (3 sınıf) | %96,5 | 0,95 |
| Hedef (3 sınıf) | %95,5 | 0,93 |
| Konu (8 sınıf) | %96,5 | 0,93 |
| Üçü birden aynı | %91,2 | – |
Yorum kimin hakkında?
Yorumların %33'ü markayla, %35'i videonun kendisiyle, %32'si başka şeylerle ilgili.
- %27%44%30
- %55%27%18
- %29%29%42
- %31%25%45
Reklam filmlerinin altında yorumların %44'ü reklamı konuşuyor: müziği, oyuncuyu, uyandırdığı duyguyu. Kaza anlatılarında yorumların %45'i ne markayla ne videoyla ilgili: dua, taziye, genel havacılık bilgisi. Markanın en çok konuşulduğu yer haber ve analiz videoları (%55).
Hedef, duyguyu da belirliyor:
| Yorumun hedefi | Yorum | Olumlu | Nötr | Olumsuz |
|---|---|---|---|---|
| Marka | 2.222 | %34,9 | %15,4 | %49,6 |
| Video | 2.335 | %64,7 | %23,0 | %12,2 |
| Diğer | 2.102 | %16,0 | %61,5 | %22,5 |
Videoya yazılan yorumların %65'i olumlu, markaya yazılanların %50'si olumsuz. İkisini ayırmayan bir analiz, “reklam çok güzel olmuş” yorumunu markaya övgü sayar.
Hedefi yok sayınca ne oluyor?
Hedefe bakmadan hesaplanan olumsuz payı üç markada da düşük kalıyor: THY'de %20 yerine %36, Pegasus'ta %38 yerine %60. Fark en çok THY'de, çünkü yorumlarının çoğu reklam filmlerinin altında ve reklamı övüyor.
Marka mı, video türü mü?
Yukarıdaki grafiğe bakıp “THY'ye tepki Pegasus'a olandan çok daha az” demek kolay. Önce olumsuz payının video türüne göre nasıl değiştiğine bakmak gerekiyor.
Nokta oranı, çizgi %95 aralığı gösteriyor.
Aynı markalar, aynı platform; reklam filminin altında olumsuz payı %26, kaza anlatısının altında %83. Tür içinde markalara ayırınca tablo değişiyor:
| Marka | Reklam ve marka filmi | Haber ve analiz | Uçuş videosu | Kaza anlatısı |
|---|---|---|---|---|
| THY | %2116–28 · n=649 | %5325–76 · n=165 | %2611–80 · n=38 | %8882–93 · n=146 |
| Pegasus | %2318–54 · n=191 | %6350–74 · n=308 | az verin=15 | %8073–84 · n=309 |
| AJet | %8772–95 · n=54 | %6647–79 · n=258 | %3519–56 · n=89 | – |
Reklam filmlerinde THY %21, Pegasus %23: fark yok. Haber videolarında %53 ve %63; aralıklar geniş ve üst üste biniyor. İki markanın toplamdaki farkı, yorumlarının hangi tür videolardan geldiğine bağlı. Bunu tek rakamla görmek için Pegasus'un oranlarını THY'nin video karışımına göre yeniden ağırlıklandırdım:
| Karşılaştırma | Ham olumsuz payı | THY'nin video karışımıyla | THY (aynı türlerde) |
|---|---|---|---|
| Pegasusreklam, haber, kaza | %60 | %3934–59 | %37 |
| AJetreklam, haber, uçuş | %62 | %8169–87 | %27 |
- Pegasus ile THY arasındaki fark büyük ölçüde örneklemden geliyor. Pegasus'un yorumları THY'ninkilerle aynı tür karışımına sahip olsaydı olumsuz payı %60 değil %39 olurdu; THY'de aynı türlerde %37. Aralık geniş (34–59), yani gerçek bir fark olmadığı da söylenemez; söylenebilecek olan, ham farkın buna kanıt olmadığı.
- AJet'te fark kapanmıyor. Kendi resmi videolarının altında bile markaya yönelik yorumların %87'si olumsuz. Ama bu rakam 4 videodaki 54 yoruma dayanıyor ve AJet'in markaya yönelik yorumlarının %55'i yalnızca üç videodan geliyor.
Olumsuz yorumlar hangi konuda?
| Konu | THY | Pegasus | AJet |
|---|---|---|---|
| Güvenlikkaza, pilotaj, bakım | %38 | %45 | %10 |
| Yönetimstrateji, sponsorluk, isim değişikliği | %36 | %15 | %27 |
| Genelkonu belirtmeden övgü ya da yergi | %5 | %12 | %19 |
| Operasyonrötar, iptal, bagaj | %3 | %5 | %21 |
| Fiyatbilet fiyatı, ek ücret | %12 | %10 | %6 |
| Hizmetekip, ikram, müşteri hizmetleri | %4 | %7 | %13 |
| Konforkoltuk, kabin, uçak tipi | %2 | %7 | %3 |
Her sütun kendi içinde %100. Koyu yazılan, o markada en büyük pay.
- Pegasus: Olumsuz yorumların %45'i güvenlikle ilgili ve bunların %85'i kaza videolarının altında yazılmış.
- THY: Olumsuzlar ikiye bölünüyor: güvenlik (%38, eski kazaların anlatıldığı videolar) ve yönetim (%36, sponsorluk anlaşmaları ve şirketin yönetilme biçimi). Markaya yönelik yorumların %52'si ise konu belirtmeyen, çoğu gurur ve övgü içeren yorumlar.
- AJet: En büyük pay yönetimde (%27); çoğu isim değişikliği ve şirketin geleceğiyle ilgili. Operasyon ve hizmet şikâyetlerinin payı diğer iki markadan yüksek.
Fiyat, operasyon, hizmet ve konfor, yani yolcunun günlük deneyimi, markaya yönelik yorumların yalnızca %23'ü. YouTube yorumları marka algısı ve gündem için iyi bir kaynak; müşteri deneyimi şikâyetleri için değil.
Ucuz bir model aynı işi yapar mı?
6.659 yorumu dil modeliyle etiketlemek mümkün. Her hafta yüz binlerce yorum geliyorsa pahalı. Bu yüzden dil modelinin etiketlerini öğretmen olarak kullanıp küçük bir model eğittim: kelime ve harf n-gramlarından TF-IDF, üstüne lojistik regresyon. Modeli her seferinde hiç görmediği videoların yorumlarıyla sınadım.
| Yaklaşım | Duygu: doğruluk | Duygu: makro F1 | Hedef: doğruluk | Hedef: makro F1 |
|---|---|---|---|---|
| Hep en sık sınıfı söyle | %39 | 0,19 | %35 | 0,17 |
| Sözlük: olumlu ve olumsuz kelime say | %51 | 0,48 | – | – |
| TF-IDF + lojistik regresyongörmediği videolarda | %71 | 0,70 | %70 | 0,69 |
| Aynı model, rastgele bölmeaynı videonun yorumları iki tarafta | %74 | 0,73 | %74 | 0,73 |
- Kelime saymak yetmiyor. Sözlük yaklaşımı yorumların yalnızca %51'inde dil modeliyle aynı kararı veriyor ve olumsuzları kaçırıyor: THY yorumlarında olumsuz payını %10 buluyor, etiketlerde %20.
- Küçük model %71'de kalıyor. Her on yorumun üçünde dil modelinden farklı karar veriyor. Tek tek yorumlara bakılacak bir iş için yeterli değil.
- Rastgele bölme burada da iyimser. Aynı videonun yorumları eğitimde ve testte birlikte olunca doğruluk %74'e çıkıyor.
Pratikte çoğu zaman tek tek yorumlar değil, bir oran izlenir. Küçük modelin tahminlerinden markaya yönelik olumsuz payını hesaplayıp etiketlerle karşılaştırdım:
| Marka | Dil modeli etiketleri | Hafif model | Fark |
|---|---|---|---|
| THY | %36,4 | %35,5 | −0,9 puan |
| Pegasus | %59,7 | %64,5 | +4,8 puan |
| AJet | %62,1 | %62,8 | +0,7 puan |
Yorumların %29'unda dil modelinden farklı karar veren model, marka düzeyindeki oranı 1–5 puan farkla tutturuyor; hatalar büyük ölçüde birbirini götürüyor. Haftalık bir göstergeyi izlemek için küçük model, tek bir yorumu sınıflandırmak için dil modeli daha doğru araç.
Kaç etiket yetiyor?
| Eğitimdeki etiketli yorum | Duygu: makro F1 | Hedef: makro F1 |
|---|---|---|
| 250 | 0,57 | 0,57 |
| 500 | 0,62 | 0,59 |
| 1.000 | 0,65 | 0,63 |
| 2.000 | 0,67 | 0,65 |
| 4.000 | 0,69 | 0,69 |
| 5.327 | 0,70 | 0,69 |
Bin etiketle model büyük kısmını öğreniyor; sonrası yavaş ama hâlâ artıyor. Eğri düzleşmediği için, daha fazla etiketin yorum düzeyindeki doğruluğu biraz daha artıracağını söylemek mümkün.
Sınırlar
Sonuç
Bir markanın videosunun altına yazılan yorum, o marka hakkında olmak zorunda değil. Bu veride üç yorumdan yalnızca biri öyleydi; geri kalanı sayıldığında marka olduğundan olumlu görünüyordu.
İkinci ders örneklemle ilgili. THY ile Pegasus arasındaki 23 puanlık farkın büyük kısmı, hangi tür videoların toplandığından geliyordu. Bir duygu skorunu iki marka ya da iki dönem arasında karşılaştırmadan önce, kaynak karışımının aynı olup olmadığına bakmak gerekiyor.
Veri: YouTube yorumları, 7 Ekim 2026'da toplandı. Yorum metinleri ve yazar bilgisi paylaşılmıyor. Oranlar bu örnekleme aittir; şirketlerin hizmet kalitesi hakkında bir değerlendirme değildir.
Sentiment analysis usually asks one question: is the comment positive or negative? In this study I added a second one: who is the comment about? I labelled 6,659 Turkish comments under YouTube videos about THY (Turkish Airlines), Pegasus and AJet by sentiment, target and topic. The second question changed the answer to the first.
Data
I collected the comments from YouTube with a Python script I wrote. For each brand I used the same five neutral search patterns (such as “… flight experience” and “… review”). Searching with words like “delay” or “complaint” would have pulled the result towards negative from the start. I kept videos whose title names the brand and took up to the 300 newest comments from each. I did not collect commenters' names.
- Collect
- Filter
- Label
- Analyse
- Compare with
a cheap model
| Step | Comments left |
|---|---|
| Comments collected | 13,151 |
| Duplicates removed | 13,126 |
| Turkish comments | 9,614 |
| Videos unrelated to the brand removed | 8,760 |
| Replies removed (analysis set) | 6,659 |
27% of the comments were not in Turkish; I identified the language with an automatic language detector. I removed replies, because a reply is written to another commenter and cannot be read correctly without the comment above it.
The videos fall into four types
The searches returned 115 videos. I assigned each one to a type by hand from its title; 8 were unrelated to the brand (a song with the same name, a match played by the airline's volleyball team) and were removed.
| Brand | Ads and brand films | News and analysis | Flight videos | Crash stories | Total | Videos |
|---|---|---|---|---|---|---|
| THY | 2,901 | 380 | 209 | 502 | 3,992 | 36 |
| Pegasus | 375 | 596 | 98 | 974 | 2,043 | 31 |
| AJet | 82 | 361 | 181 | – | 624 | 27 |
The distribution is uneven: 73% of THY's comments sit under ad films and 48% of Pegasus's under crash stories. I did not create this imbalance; it is how the search results came back. The most important finding below comes from it.
Labelling
Each comment received three labels:
- Target: Is the comment about the brand, about the video itself (the ad, the music, the narrator), or about something else (politics, prayers and condolences, other companies, chat)?
- Sentiment: Is the writer's attitude towards that target positive, negative or neutral?
- Topic: If the target is the brand, which topic? Safety, management, price, operations, service, comfort or general.
The labels were assigned by a language model (Claude). The model read each comment together with the title and type of the video it was written under and decided according to a written guideline. No word lists or rules were used; that is what lets a comment like “great, a three-hour delay again” count as negative.
To see how far the labels can be trusted, I had a random 400 comments labelled again in a separate pass that did not see the first labels.
| Label | Agreement between passes | Cohen's kappa |
|---|---|---|
| Sentiment (3 classes) | 96.5% | 0.95 |
| Target (3 classes) | 95.5% | 0.93 |
| Topic (8 classes) | 96.5% | 0.93 |
| All three identical | 91.2% | – |
Who is the comment about?
33% of the comments are about the brand, 35% about the video itself and 32% about something else.
- 27%44%30%
- 55%27%18%
- 29%29%42%
- 31%25%45%
Under ad films, 44% of the comments talk about the ad: its music, its actors, the feeling it evokes. Under crash stories, 45% are about neither the brand nor the video: prayers, condolences, general aviation talk. The brand is discussed most under news and analysis videos (55%).
The target also shapes the sentiment:
| Target of the comment | Comments | Positive | Neutral | Negative |
|---|---|---|---|---|
| Brand | 2,222 | 34.9% | 15.4% | 49.6% |
| Video | 2,335 | 64.7% | 23.0% | 12.2% |
| Other | 2,102 | 16.0% | 61.5% | 22.5% |
65% of the comments about the video are positive, while 50% of those about the brand are negative. An analysis that does not separate the two counts “what a beautiful ad” as praise for the brand.
What happens when the target is ignored?
Computed without the target, the negative share comes out low for all three brands: 20% instead of 36% for THY, 38% instead of 60% for Pegasus. The gap is largest for THY, because most of its comments sit under ad films and praise the ad.
Brand or video type?
Looking at the chart above, it is easy to say “THY draws far less criticism than Pegasus”. First it is worth seeing how the negative share changes with the type of video.
The dot is the share, the line the 95% interval.
Same brands, same platform; the negative share is 26% under ad films and 83% under crash stories. Splitting each type by brand changes the picture:
| Brand | Ads and brand films | News and analysis | Flight videos | Crash stories |
|---|---|---|---|---|
| THY | 21%16–28 · n=649 | 53%25–76 · n=165 | 26%11–80 · n=38 | 88%82–93 · n=146 |
| Pegasus | 23%18–54 · n=191 | 63%50–74 · n=308 | too fewn=15 | 80%73–84 · n=309 |
| AJet | 87%72–95 · n=54 | 66%47–79 · n=258 | 35%19–56 · n=89 | – |
Under ad films THY is at 21% and Pegasus at 23%: no difference. Under news videos they are at 53% and 63%; the intervals are wide and overlap. The overall gap between the two brands depends on which types of video their comments come from. To put that into one number, I reweighted Pegasus's shares to THY's mix of video types:
| Comparison | Raw negative share | With THY's video mix | THY (same types) |
|---|---|---|---|
| Pegasusads, news, crash | 60% | 39%34–59 | 37% |
| AJetads, news, flight | 62% | 81%69–87 | 27% |
- The gap between Pegasus and THY comes largely from the sample. If Pegasus's comments had the same mix of video types as THY's, its negative share would be 39% rather than 60%; THY is at 37% on the same types. The interval is wide (34–59), so a real difference cannot be ruled out; what can be said is that the raw gap is not evidence of one.
- For AJet the gap does not close. Even under its own official videos, 87% of the comments about the brand are negative. But that figure rests on 54 comments under 4 videos, and 55% of AJet's brand comments come from just three videos.
What are the negative comments about?
| Topic | THY | Pegasus | AJet |
|---|---|---|---|
| Safetycrashes, piloting, maintenance | 38% | 45% | 10% |
| Managementstrategy, sponsorship, rebranding | 36% | 15% | 27% |
| Generalpraise or criticism with no stated topic | 5% | 12% | 19% |
| Operationsdelays, cancellations, baggage | 3% | 5% | 21% |
| Pricefares, extra fees | 12% | 10% | 6% |
| Servicecrew, catering, customer service | 4% | 7% | 13% |
| Comfortseats, cabin, aircraft type | 2% | 7% | 3% |
Each column sums to 100%. Bold marks the largest share for that brand.
- Pegasus: 45% of the negative comments are about safety, and 85% of those were written under crash videos.
- THY: The negatives split in two: safety (38%, videos retelling old crashes) and management (36%, sponsorship deals and how the company is run). Meanwhile 52% of all comments about the brand name no topic and are mostly pride and praise.
- AJet: The largest share is management (27%), mostly about the rebranding and the company's future. Operations and service complaints take a larger share than for the other two.
Price, operations, service and comfort, the passenger's everyday experience, make up only 23% of the comments about the brand. YouTube comments are a good source for brand perception and the news agenda, not for customer-experience complaints.
Can a cheap model do the same job?
Labelling 6,659 comments with a language model is feasible. With hundreds of thousands of comments every week it gets expensive. So I used the language model's labels as a teacher and trained a small model: TF-IDF on word and character n-grams with logistic regression on top. Each time, the model was tested on comments from videos it had never seen.
| Approach | Sentiment: accuracy | Sentiment: macro F1 | Target: accuracy | Target: macro F1 |
|---|---|---|---|---|
| Always predict the most common class | 39% | 0.19 | 35% | 0.17 |
| Lexicon: count positive and negative words | 51% | 0.48 | – | – |
| TF-IDF + logistic regressionon unseen videos | 71% | 0.70 | 70% | 0.69 |
| Same model, random splitcomments from one video on both sides | 74% | 0.73 | 74% | 0.73 |
- Counting words is not enough. The lexicon approach agrees with the language model on only 51% of the comments and misses negatives: it puts the negative share of THY's comments at 10%, against 20% in the labels.
- The small model stops at 71%. It disagrees with the language model on three comments out of ten. That is not enough for work that looks at individual comments.
- A random split is optimistic here too. With comments from the same video in both training and test, accuracy rises to 74%.
In practice one usually tracks a share, not individual comments. I computed the negative share of comments about the brand from the small model's predictions and compared it with the labels:
| Brand | Language-model labels | Light model | Difference |
|---|---|---|---|
| THY | 36.4% | 35.5% | −0.9 pts |
| Pegasus | 59.7% | 64.5% | +4.8 pts |
| AJet | 62.1% | 62.8% | +0.7 pts |
A model that disagrees with the language model on 29% of the comments gets the brand-level share within 1–5 points; the errors largely cancel out. For tracking a weekly indicator the small model is the right tool, for classifying a single comment the language model is.
How many labels are enough?
| Labelled comments in training | Sentiment: macro F1 | Target: macro F1 |
|---|---|---|
| 250 | 0.57 | 0.57 |
| 500 | 0.62 | 0.59 |
| 1,000 | 0.65 | 0.63 |
| 2,000 | 0.67 | 0.65 |
| 4,000 | 0.69 | 0.69 |
| 5,327 | 0.70 | 0.69 |
With a thousand labels the model learns most of what it will learn; after that the gain is slow but still there. Since the curve has not flattened, more labels would raise comment-level accuracy a little further.
Limitations
Conclusion
A comment written under a brand's video does not have to be about that brand. In this data only one in three was; counting the rest made the brand look more positive than it was.
The second lesson is about sampling. Most of the 23-point gap between THY and Pegasus came from which types of video were collected. Before comparing a sentiment score between two brands or two periods, it is worth checking whether the mix of sources is the same.
Data: YouTube comments, collected on October 7, 2026. Comment texts and author details are not shared. The shares describe this sample; they are not an assessment of the companies' service quality.