Bir YouTube Yorumu Kimin Hakkında? Havayolu Videolarında Hedefli Duygu AnaliziWho Is a YouTube Comment About? Target-Aware Sentiment Analysis on Airline Videos

THY, Pegasus ve AJet videolarının altındaki 6.659 Türkçe yorumu duygu, hedef ve konuya göre etiketledim. Yorumların yalnızca üçte biri markayla ilgili çıktı.I labelled 6,659 Turkish comments under THY, Pegasus and AJet videos by sentiment, target and topic. Only a third of them turned out to be about the brand.

  • Python
  • NLP
  • LLM
  • Scikit-learn

Duygu analizinde genellikle tek soru sorulur: yorum olumlu mu, olumsuz mu? Bu çalışmada ikinci bir soru ekledim: yorum kimin hakkında? THY, Pegasus ve AJet ile ilgili YouTube videolarının altındaki 6.659 Türkçe yorumu duygu, hedef ve konuya göre etiketledim. İkinci soru, ilk sorunun cevabını baştan sona değiştirdi.

Etiketlenen yorum6.65994 video, üç havayolu
Markayla ilgili yorum%33kalanı videoyu ya da başka bir şeyi konuşuyor
Olumsuz payı, video türüne göre%26 – %83reklam filminden kaza anlatısına
Okurken: Buradaki oranlar müşteri memnuniyetini ölçmüyor. YouTube'da bu videoların altına yorum yazanların ne dediğini gösteriyor; çalışmanın konusu da zaten bu ikisinin neden aynı şey olmadığı.

Veri

Yorumları YouTube'dan, yazdığım bir Python scriptiyle topladım. Her marka için aynı beş nötr arama kalıbını kullandım (“… uçuş deneyimi”, “… inceleme” gibi). “Rötar” ya da “şikâyet” gibi kelimelerle arasaydım sonucu baştan olumsuza çekmiş olurdum. Başlığında markanın adı geçen videoları aldım, her videodan en yeni 300 yoruma kadar. Yorum yazanların adlarını toplamadım.

Çalışmanın adımları
  1. Topla
  2. Ayıkla
  3. Etiketle
  4. Analiz et
  5. Ucuz modelle
    karşılaştır
AdımKalan yorum
Toplanan yorum13.151
Tekrarlar atıldı13.126
Türkçe olanlar9.614
Markayla ilgisiz videolar çıkarıldı8.760
Yanıtlar çıkarıldı (analiz kümesi)6.659

Yorumların %27'si Türkçe değildi; dili otomatik bir dil tanıyıcıyla belirledim. Yanıtları çıkardım, çünkü yanıt başka bir yorumcuya yazılır ve üstündeki yorum olmadan doğru okunamaz.

Videolar dört türe ayrılıyor

Aramalardan 115 video geldi. Başlıklarına bakarak her birini elle bir türe atadım; 8 tanesi markayla ilgisizdi (aynı adı taşıyan bir şarkı, şirketin voleybol takımının bir maçı gibi) ve çıkarıldı.

MarkaReklam ve marka filmiHaber ve analizUçuş videosuKaza anlatısıToplamVideo
THY2.9013802095023.99236
Pegasus375596989742.04331
AJet82361181–62427

Dağılım dengesiz: THY yorumlarının %73'ü reklam filmlerinin, Pegasus yorumlarının %48'i kaza anlatılarının altında. Bu dengesizliği ben yaratmadım, arama sonuçları böyle geldi. Aşağıdaki en önemli bulgu da buradan çıkıyor.

Etiketleme

Her yoruma üç etiket verildi:

  • Hedef: Yorum markayla mı, videonun kendisiyle mi (reklam, müzik, anlatıcı), yoksa başka bir şeyle mi (siyaset, dua ve taziye, başka şirketler, sohbet) ilgili?
  • Duygu: Yazanın o hedefe karşı tutumu olumlu mu, olumsuz mu, nötr mü?
  • Konu: Hedef markaysa hangi konuda? Güvenlik, yönetim, fiyat, operasyon, hizmet, konfor ya da genel.

Etiketleri bir dil modeline (Claude) verdirdim. Model her yorumu, altına yazıldığı videonun başlığı ve türüyle birlikte okudu ve yazılı bir kılavuza göre karar verdi. Kelime listesi ya da kural kullanılmadı; “harika, yine üç saat rötar” gibi bir yorumun olumsuz sayılması buna bağlı.

Etiketlere ne kadar güvenilebileceğini görmek için rastgele 400 yorumu, ilk etiketleri görmeyen ayrı bir geçişte yeniden etiketlettim.

Etiketİki geçişin uyumuCohen kappa
Duygu (3 sınıf)%96,50,95
Hedef (3 sınıf)%95,50,93
Konu (8 sınıf)%96,50,93
Üçü birden aynı%91,2–
Bu tutarlılık, doğruluk değil. İki geçiş de aynı modelden geliyor; ikisi aynı yönde yanılıyor olabilir. Etiketler insan etiketiyle karşılaştırılmadı. Aşağıdaki oranları “dil modelinin kılavuza göre okuması” olarak okumak gerekir.

Yorum kimin hakkında?

Yorumların %33'ü markayla, %35'i videonun kendisiyle, %32'si başka şeylerle ilgili.

Video türüne göre yorumun hedefi
  • Reklam ve marka filmi3.358 yorum%27%44%30
  • Haber ve analiz1.337 yorum%55%27%18
  • Uçuş videosu488 yorum%29%29%42
  • Kaza anlatısı1.476 yorum%31%25%45

Reklam filmlerinin altında yorumların %44'ü reklamı konuşuyor: müziği, oyuncuyu, uyandırdığı duyguyu. Kaza anlatılarında yorumların %45'i ne markayla ne videoyla ilgili: dua, taziye, genel havacılık bilgisi. Markanın en çok konuşulduğu yer haber ve analiz videoları (%55).

Hedef, duyguyu da belirliyor:

Yorumun hedefiYorumOlumluNötrOlumsuz
Marka2.222%34,9%15,4%49,6
Video2.335%64,7%23,0%12,2
Diğer2.102%16,0%61,5%22,5

Videoya yazılan yorumların %65'i olumlu, markaya yazılanların %50'si olumsuz. İkisini ayırmayan bir analiz, “reklam çok güzel olmuş” yorumunu markaya övgü sayar.

Hedefi yok sayınca ne oluyor?

Olumsuz yorum payı: bütün yorumlar ve yalnızca markaya yönelik olanlar
  • THYbütün yorumlar%20
  • THYyalnızca markaya yönelik%36
  • Pegasusbütün yorumlar%38
  • Pegasusyalnızca markaya yönelik%60
  • AJetbütün yorumlar%47
  • AJetyalnızca markaya yönelik%62

Hedefe bakmadan hesaplanan olumsuz payı üç markada da düşük kalıyor: THY'de %20 yerine %36, Pegasus'ta %38 yerine %60. Fark en çok THY'de, çünkü yorumlarının çoğu reklam filmlerinin altında ve reklamı övüyor.

Marka mı, video türü mü?

Yukarıdaki grafiğe bakıp “THY'ye tepki Pegasus'a olandan çok daha az” demek kolay. Önce olumsuz payının video türüne göre nasıl değiştiğine bakmak gerekiyor.

Markaya yönelik yorumlarda olumsuz payı, video türüne göre
  • Reklam ve marka filmi894 yorum, 34 video%2619–34
  • Uçuş videosu142 yorum, 21 video%3320–51
  • Haber ve analiz731 yorum, 22 video%6251–70
  • Kaza anlatısı455 yorum, 11 video%8378–87

Nokta oranı, çizgi %95 aralığı gösteriyor.

Aynı markalar, aynı platform; reklam filminin altında olumsuz payı %26, kaza anlatısının altında %83. Tür içinde markalara ayırınca tablo değişiyor:

MarkaReklam ve marka filmiHaber ve analizUçuş videosuKaza anlatısı
THY%2116–28 · n=649%5325–76 · n=165%2611–80 · n=38%8882–93 · n=146
Pegasus%2318–54 · n=191%6350–74 · n=308az verin=15%8073–84 · n=309
AJet%8772–95 · n=54%6647–79 · n=258%3519–56 · n=89–

Reklam filmlerinde THY %21, Pegasus %23: fark yok. Haber videolarında %53 ve %63; aralıklar geniş ve üst üste biniyor. İki markanın toplamdaki farkı, yorumlarının hangi tür videolardan geldiğine bağlı. Bunu tek rakamla görmek için Pegasus'un oranlarını THY'nin video karışımına göre yeniden ağırlıklandırdım:

KarşılaştırmaHam olumsuz payıTHY'nin video karışımıylaTHY (aynı türlerde)
Pegasusreklam, haber, kaza%60%3934–59%37
AJetreklam, haber, uçuş%62%8169–87%27
  • Pegasus ile THY arasındaki fark büyük ölçüde örneklemden geliyor. Pegasus'un yorumları THY'ninkilerle aynı tür karışımına sahip olsaydı olumsuz payı %60 değil %39 olurdu; THY'de aynı türlerde %37. Aralık geniş (34–59), yani gerçek bir fark olmadığı da söylenemez; söylenebilecek olan, ham farkın buna kanıt olmadığı.
  • AJet'te fark kapanmıyor. Kendi resmi videolarının altında bile markaya yönelik yorumların %87'si olumsuz. Ama bu rakam 4 videodaki 54 yoruma dayanıyor ve AJet'in markaya yönelik yorumlarının %55'i yalnızca üç videodan geliyor.
Aralıklar neden bu kadar geniş? Aynı videonun altındaki yorumlar birbirine benzer; biri tartışma başlatır, diğerleri ona katılır. Belirsizliği hesaplarken yorumları değil videoları yeniden örnekledim. Pegasus için 823 yorum var gibi görünüyor, ama asıl örneklem 28 video.

Olumsuz yorumlar hangi konuda?

Markaya yönelik olumsuz yorumların konulara dağılımı
KonuTHYPegasusAJet
Güvenlikkaza, pilotaj, bakım%38%45%10
Yönetimstrateji, sponsorluk, isim değişikliği%36%15%27
Genelkonu belirtmeden övgü ya da yergi%5%12%19
Operasyonrötar, iptal, bagaj%3%5%21
Fiyatbilet fiyatı, ek ücret%12%10%6
Hizmetekip, ikram, müşteri hizmetleri%4%7%13
Konforkoltuk, kabin, uçak tipi%2%7%3

Her sütun kendi içinde %100. Koyu yazılan, o markada en büyük pay.

  • Pegasus: Olumsuz yorumların %45'i güvenlikle ilgili ve bunların %85'i kaza videolarının altında yazılmış.
  • THY: Olumsuzlar ikiye bölünüyor: güvenlik (%38, eski kazaların anlatıldığı videolar) ve yönetim (%36, sponsorluk anlaşmaları ve şirketin yönetilme biçimi). Markaya yönelik yorumların %52'si ise konu belirtmeyen, çoğu gurur ve övgü içeren yorumlar.
  • AJet: En büyük pay yönetimde (%27); çoğu isim değişikliği ve şirketin geleceğiyle ilgili. Operasyon ve hizmet şikâyetlerinin payı diğer iki markadan yüksek.

Fiyat, operasyon, hizmet ve konfor, yani yolcunun günlük deneyimi, markaya yönelik yorumların yalnızca %23'ü. YouTube yorumları marka algısı ve gündem için iyi bir kaynak; müşteri deneyimi şikâyetleri için değil.

Ucuz bir model aynı işi yapar mı?

6.659 yorumu dil modeliyle etiketlemek mümkün. Her hafta yüz binlerce yorum geliyorsa pahalı. Bu yüzden dil modelinin etiketlerini öğretmen olarak kullanıp küçük bir model eğittim: kelime ve harf n-gramlarından TF-IDF, üstüne lojistik regresyon. Modeli her seferinde hiç görmediği videoların yorumlarıyla sınadım.

YaklaşımDuygu: doğrulukDuygu: makro F1Hedef: doğrulukHedef: makro F1
Hep en sık sınıfı söyle%390,19%350,17
Sözlük: olumlu ve olumsuz kelime say%510,48––
TF-IDF + lojistik regresyongörmediği videolarda%710,70%700,69
Aynı model, rastgele bölmeaynı videonun yorumları iki tarafta%740,73%740,73
  • Kelime saymak yetmiyor. Sözlük yaklaşımı yorumların yalnızca %51'inde dil modeliyle aynı kararı veriyor ve olumsuzları kaçırıyor: THY yorumlarında olumsuz payını %10 buluyor, etiketlerde %20.
  • Küçük model %71'de kalıyor. Her on yorumun üçünde dil modelinden farklı karar veriyor. Tek tek yorumlara bakılacak bir iş için yeterli değil.
  • Rastgele bölme burada da iyimser. Aynı videonun yorumları eğitimde ve testte birlikte olunca doğruluk %74'e çıkıyor.

Pratikte çoğu zaman tek tek yorumlar değil, bir oran izlenir. Küçük modelin tahminlerinden markaya yönelik olumsuz payını hesaplayıp etiketlerle karşılaştırdım:

MarkaDil modeli etiketleriHafif modelFark
THY%36,4%35,5−0,9 puan
Pegasus%59,7%64,5+4,8 puan
AJet%62,1%62,8+0,7 puan

Yorumların %29'unda dil modelinden farklı karar veren model, marka düzeyindeki oranı 1–5 puan farkla tutturuyor; hatalar büyük ölçüde birbirini götürüyor. Haftalık bir göstergeyi izlemek için küçük model, tek bir yorumu sınıflandırmak için dil modeli daha doğru araç.

Kaç etiket yetiyor?

Eğitimdeki etiketli yorumDuygu: makro F1Hedef: makro F1
2500,570,57
5000,620,59
1.0000,650,63
2.0000,670,65
4.0000,690,69
5.3270,700,69

Bin etiketle model büyük kısmını öğreniyor; sonrası yavaş ama hâlâ artıyor. Eğri düzleşmediği için, daha fazla etiketin yorum düzeyindeki doğruluğu biraz daha artıracağını söylemek mümkün.

Sınırlar

Örneklemi YouTube seçtiHangi videoların çıkacağını arama sıralaması belirledi. Başka aramalar başka bir karışım, dolayısıyla başka oranlar verir.
Yorumcu, müşteri değilYorum yazanların o havayoluyla uçup uçmadığı bilinmiyor.
Etiketler dil modelindenİki geçiş tutarlı, ama insan etiketiyle karşılaştırma yapılmadı.
Video türünü ben atadımBaşlığa bakarak ve tek kişi olarak; bazı videolar iki türe de girebilir.
Az video, geniş aralıkMarka ve tür kırılımında hücre başına 4–21 video var.
Zaman yokYorumların %48'i 2025 ve sonrasından, ama tarihler yaklaşık olduğu için trend analizi yapmadım.

Sonuç

Bir markanın videosunun altına yazılan yorum, o marka hakkında olmak zorunda değil. Bu veride üç yorumdan yalnızca biri öyleydi; geri kalanı sayıldığında marka olduğundan olumlu görünüyordu.

İkinci ders örneklemle ilgili. THY ile Pegasus arasındaki 23 puanlık farkın büyük kısmı, hangi tür videoların toplandığından geliyordu. Bir duygu skorunu iki marka ya da iki dönem arasında karşılaştırmadan önce, kaynak karışımının aynı olup olmadığına bakmak gerekiyor.

Veri: YouTube yorumları, 7 Ekim 2026'da toplandı. Yorum metinleri ve yazar bilgisi paylaşılmıyor. Oranlar bu örnekleme aittir; şirketlerin hizmet kalitesi hakkında bir değerlendirme değildir.

Sentiment analysis usually asks one question: is the comment positive or negative? In this study I added a second one: who is the comment about? I labelled 6,659 Turkish comments under YouTube videos about THY (Turkish Airlines), Pegasus and AJet by sentiment, target and topic. The second question changed the answer to the first.

Comments labelled6,65994 videos, three airlines
Comments about the brand33%the rest talk about the video or something else
Negative share by video type26% – 83%from ad films to crash stories
Reading note: These shares do not measure customer satisfaction. They show what people who comment under these videos say; why those two are not the same thing is the subject of the study.

Data

I collected the comments from YouTube with a Python script I wrote. For each brand I used the same five neutral search patterns (such as “… flight experience” and “… review”). Searching with words like “delay” or “complaint” would have pulled the result towards negative from the start. I kept videos whose title names the brand and took up to the 300 newest comments from each. I did not collect commenters' names.

Steps of the study
  1. Collect
  2. Filter
  3. Label
  4. Analyse
  5. Compare with
    a cheap model
StepComments left
Comments collected13,151
Duplicates removed13,126
Turkish comments9,614
Videos unrelated to the brand removed8,760
Replies removed (analysis set)6,659

27% of the comments were not in Turkish; I identified the language with an automatic language detector. I removed replies, because a reply is written to another commenter and cannot be read correctly without the comment above it.

The videos fall into four types

The searches returned 115 videos. I assigned each one to a type by hand from its title; 8 were unrelated to the brand (a song with the same name, a match played by the airline's volleyball team) and were removed.

BrandAds and brand filmsNews and analysisFlight videosCrash storiesTotalVideos
THY2,9013802095023,99236
Pegasus375596989742,04331
AJet82361181–62427

The distribution is uneven: 73% of THY's comments sit under ad films and 48% of Pegasus's under crash stories. I did not create this imbalance; it is how the search results came back. The most important finding below comes from it.

Labelling

Each comment received three labels:

  • Target: Is the comment about the brand, about the video itself (the ad, the music, the narrator), or about something else (politics, prayers and condolences, other companies, chat)?
  • Sentiment: Is the writer's attitude towards that target positive, negative or neutral?
  • Topic: If the target is the brand, which topic? Safety, management, price, operations, service, comfort or general.

The labels were assigned by a language model (Claude). The model read each comment together with the title and type of the video it was written under and decided according to a written guideline. No word lists or rules were used; that is what lets a comment like “great, a three-hour delay again” count as negative.

To see how far the labels can be trusted, I had a random 400 comments labelled again in a separate pass that did not see the first labels.

LabelAgreement between passesCohen's kappa
Sentiment (3 classes)96.5%0.95
Target (3 classes)95.5%0.93
Topic (8 classes)96.5%0.93
All three identical91.2%–
This is consistency, not accuracy. Both passes come from the same model and may be wrong in the same direction. The labels were not compared with human labels. The shares below should be read as “the language model's reading under the guideline”.

Who is the comment about?

33% of the comments are about the brand, 35% about the video itself and 32% about something else.

Target of the comment by video type
  • Ads and brand films3,358 comments27%44%30%
  • News and analysis1,337 comments55%27%18%
  • Flight videos488 comments29%29%42%
  • Crash stories1,476 comments31%25%45%

Under ad films, 44% of the comments talk about the ad: its music, its actors, the feeling it evokes. Under crash stories, 45% are about neither the brand nor the video: prayers, condolences, general aviation talk. The brand is discussed most under news and analysis videos (55%).

The target also shapes the sentiment:

Target of the commentCommentsPositiveNeutralNegative
Brand2,22234.9%15.4%49.6%
Video2,33564.7%23.0%12.2%
Other2,10216.0%61.5%22.5%

65% of the comments about the video are positive, while 50% of those about the brand are negative. An analysis that does not separate the two counts “what a beautiful ad” as praise for the brand.

What happens when the target is ignored?

Share of negative comments: all comments versus comments about the brand only
  • THYall comments20%
  • THYcomments about the brand only36%
  • Pegasusall comments38%
  • Pegasuscomments about the brand only60%
  • AJetall comments47%
  • AJetcomments about the brand only62%

Computed without the target, the negative share comes out low for all three brands: 20% instead of 36% for THY, 38% instead of 60% for Pegasus. The gap is largest for THY, because most of its comments sit under ad films and praise the ad.

Brand or video type?

Looking at the chart above, it is easy to say “THY draws far less criticism than Pegasus”. First it is worth seeing how the negative share changes with the type of video.

Negative share among comments about the brand, by video type
  • Ads and brand films894 comments, 34 videos26%19–34
  • Flight videos142 comments, 21 videos33%20–51
  • News and analysis731 comments, 22 videos62%51–70
  • Crash stories455 comments, 11 videos83%78–87

The dot is the share, the line the 95% interval.

Same brands, same platform; the negative share is 26% under ad films and 83% under crash stories. Splitting each type by brand changes the picture:

BrandAds and brand filmsNews and analysisFlight videosCrash stories
THY21%16–28 · n=64953%25–76 · n=16526%11–80 · n=3888%82–93 · n=146
Pegasus23%18–54 · n=19163%50–74 · n=308too fewn=1580%73–84 · n=309
AJet87%72–95 · n=5466%47–79 · n=25835%19–56 · n=89–

Under ad films THY is at 21% and Pegasus at 23%: no difference. Under news videos they are at 53% and 63%; the intervals are wide and overlap. The overall gap between the two brands depends on which types of video their comments come from. To put that into one number, I reweighted Pegasus's shares to THY's mix of video types:

ComparisonRaw negative shareWith THY's video mixTHY (same types)
Pegasusads, news, crash60%39%34–5937%
AJetads, news, flight62%81%69–8727%
  • The gap between Pegasus and THY comes largely from the sample. If Pegasus's comments had the same mix of video types as THY's, its negative share would be 39% rather than 60%; THY is at 37% on the same types. The interval is wide (34–59), so a real difference cannot be ruled out; what can be said is that the raw gap is not evidence of one.
  • For AJet the gap does not close. Even under its own official videos, 87% of the comments about the brand are negative. But that figure rests on 54 comments under 4 videos, and 55% of AJet's brand comments come from just three videos.
Why are the intervals so wide? Comments under the same video resemble each other; one starts an argument and the others join in. To estimate uncertainty I resampled videos, not comments. Pegasus appears to have 823 comments, but the real sample is 28 videos.

What are the negative comments about?

Topics of the negative comments about each brand
TopicTHYPegasusAJet
Safetycrashes, piloting, maintenance38%45%10%
Managementstrategy, sponsorship, rebranding36%15%27%
Generalpraise or criticism with no stated topic5%12%19%
Operationsdelays, cancellations, baggage3%5%21%
Pricefares, extra fees12%10%6%
Servicecrew, catering, customer service4%7%13%
Comfortseats, cabin, aircraft type2%7%3%

Each column sums to 100%. Bold marks the largest share for that brand.

  • Pegasus: 45% of the negative comments are about safety, and 85% of those were written under crash videos.
  • THY: The negatives split in two: safety (38%, videos retelling old crashes) and management (36%, sponsorship deals and how the company is run). Meanwhile 52% of all comments about the brand name no topic and are mostly pride and praise.
  • AJet: The largest share is management (27%), mostly about the rebranding and the company's future. Operations and service complaints take a larger share than for the other two.

Price, operations, service and comfort, the passenger's everyday experience, make up only 23% of the comments about the brand. YouTube comments are a good source for brand perception and the news agenda, not for customer-experience complaints.

Can a cheap model do the same job?

Labelling 6,659 comments with a language model is feasible. With hundreds of thousands of comments every week it gets expensive. So I used the language model's labels as a teacher and trained a small model: TF-IDF on word and character n-grams with logistic regression on top. Each time, the model was tested on comments from videos it had never seen.

ApproachSentiment: accuracySentiment: macro F1Target: accuracyTarget: macro F1
Always predict the most common class39%0.1935%0.17
Lexicon: count positive and negative words51%0.48––
TF-IDF + logistic regressionon unseen videos71%0.7070%0.69
Same model, random splitcomments from one video on both sides74%0.7374%0.73
  • Counting words is not enough. The lexicon approach agrees with the language model on only 51% of the comments and misses negatives: it puts the negative share of THY's comments at 10%, against 20% in the labels.
  • The small model stops at 71%. It disagrees with the language model on three comments out of ten. That is not enough for work that looks at individual comments.
  • A random split is optimistic here too. With comments from the same video in both training and test, accuracy rises to 74%.

In practice one usually tracks a share, not individual comments. I computed the negative share of comments about the brand from the small model's predictions and compared it with the labels:

BrandLanguage-model labelsLight modelDifference
THY36.4%35.5%−0.9 pts
Pegasus59.7%64.5%+4.8 pts
AJet62.1%62.8%+0.7 pts

A model that disagrees with the language model on 29% of the comments gets the brand-level share within 1–5 points; the errors largely cancel out. For tracking a weekly indicator the small model is the right tool, for classifying a single comment the language model is.

How many labels are enough?

Labelled comments in trainingSentiment: macro F1Target: macro F1
2500.570.57
5000.620.59
1,0000.650.63
2,0000.670.65
4,0000.690.69
5,3270.700.69

With a thousand labels the model learns most of what it will learn; after that the gain is slow but still there. Since the curve has not flattened, more labels would raise comment-level accuracy a little further.

Limitations

YouTube chose the sampleSearch ranking decided which videos came up. Other searches would give a different mix and therefore different shares.
Commenters, not customersWhether the people commenting have flown with the airline is unknown.
Labels come from a language modelThe two passes are consistent, but there is no comparison with human labels.
I assigned the video typesFrom the title and alone; some videos could fit two types.
Few videos, wide intervalsIn the brand-by-type breakdown there are 4–21 videos per cell.
No time dimension48% of the comments are from 2025 or later, but dates are approximate, so I did no trend analysis.

Conclusion

A comment written under a brand's video does not have to be about that brand. In this data only one in three was; counting the rest made the brand look more positive than it was.

The second lesson is about sampling. Most of the 23-point gap between THY and Pegasus came from which types of video were collected. Before comparing a sentiment score between two brands or two periods, it is worth checking whether the mix of sources is the same.

Data: YouTube comments, collected on October 7, 2026. Comment texts and author details are not shared. The shares describe this sample; they are not an assessment of the companies' service quality.