NLP Nedir? Doğal Dil İşleme ile Metin Analizi
Tokenization'dan TF-IDF'e, metin sınıflandırmadan müşteri yorumu analizine kadar doğal dil işlemenin temel kavramlarını örneklerle inceliyorum.
- Yapay Zeka
- Sosyal Medya

İnsanların yazdığı metinleri bilgisayarların analiz edebileceği yapılara dönüştürmek — NLP'nin özü bu. Müşteri yorumlarından sosyal medya paylaşımlarına, çağrı merkezi kayıtlarından arama sorgularına kadar büyük miktardaki metni anlamlandırmak için kullanılan temel alanlardan biri.
Bu yazıda tokenization, stop words, stemming, lemmatization ve TF-IDF gibi temel kavramları; metin sınıflandırmayı ve bütün bunların gerçek bir müşteri geri bildirimi analizinde nasıl birleştiğini ele alıyorum.
“NLP'nin amacı yalnızca kelimeleri işlemek değil; metnin içindeki sinyalleri ortaya çıkararak daha hızlı ve tutarlı kararlar almaya yardımcı olmaktır.”
1. NLP Nedir?
Natural Language Processing (NLP), Türkçesiyle doğal dil işleme, insanların kullandığı doğal dili bilgisayarların işleyebileceği, analiz edebileceği ve gerektiğinde anlamlandırabileceği yapılara dönüştürmeye odaklanan alandır.
İnsan dili yapılandırılmış bir veri tablosu gibi değil. Aynı duygu farklı kelimelerle anlatılabilir, cümleler bağlama göre farklı anlamlar taşır; yazım hataları, kısaltmalar, emojiler ve internet dili analizi zorlaştırır. NLP teknikleri bu karmaşık metinleri makine tarafından işlenebilir hale getiriyor.
Metni yapılandır
Ham cümleleri token, özellik veya sayısal temsil gibi işlenebilir yapılara dönüştür.
Örüntüleri bul
Kelime sıklıkları, temalar, duygular ve kullanıcı davranışları gibi örüntüleri keşfet.
Karar üret
Metinleri sınıflandır, önceliklendir veya bir iş kararını destekleyen içgörülere dönüştür.
2. Metin Verisi Neden Bu Kadar Değerli?
Şirketlerin önemli bir bölümü yapılandırılmış verinin yanında çok büyük miktarda yapılandırılmamış metin üretiyor: sosyal medya yorumları, müşteri hizmetleri kayıtları, e-postalar, değerlendirmeler, forum içerikleri, arama sorguları ve açık uçlu anket cevapları.
Yorum ve paylaşımlar
Marka hakkında konuşulanların büyük bölümü serbest metin halinde gelir.
Müşteri geri bildirimleri
Destek kayıtları ve talepler, deneyimin en doğrudan kaynağıdır.
Arama metinleri
Kullanıcıların ne aradığı, ihtiyacın ham halidir.
Açık uçlu cevaplar
Sayısal skorların arkasındaki nedeni açıklar.
Buradaki asıl problem veri miktarından çok anlamın dağınık olması. “Kargo çok geç geldi”, “teslimat süresi berbattı” ve “ürün geç ulaştı” farklı cümlelerdir; ancak iş açısından aynı probleme işaret eder. NLP bu ifadeler arasındaki benzerlik ve farklılıkları analiz etmeye yardımcı oluyor.
3. Bir NLP Projesinin Temel Akışı
NLP projesi tek bir algoritmadan oluşmaz. Başarılı sonuç için veri hazırlama, metin ön işleme, özellik çıkarımı ve modelleme adımlarının birlikte ele alınması gerekir.
- Veriyi
Topla - Temizle
- Tokenize
Et - Özellik
Çıkar - Modelle
- İçgörü
Üret
4. Tokenization Nedir?
Tokenization, bir metni daha küçük parçalara, yani token'lara ayırma işlemi. En basit örnekte token'lar kelimelerdir; daha gelişmiş sistemlerde alt kelimeler de token olabilir.
"Uygulama çok hızlı çalışıyor ama giriş sorunu yaşıyorum."
# tokenization sonrası
["Uygulama", "çok", "hızlı", "çalışıyor",
"ama", "giriş", "sorunu", "yaşıyorum"]Tokenization, NLP akışının temel yapı taşlarından biri; sonraki adımların büyük bölümü bu token'lar üzerinden ilerliyor. Türkçe gibi eklemeli dillerde kelimelerin çok sayıda farklı biçimde ortaya çıkabilmesi, tokenization ve dil ön işleme adımlarını ayrıca önemli hale getiriyor.
5. Stop Words Nedir?
Stop words, belirli bir görev açısından bilgi katkısı düşük olabilen ve analizden çıkarılması düşünülen kelimeler. “ve”, “ile”, “bir”, “bu” gibi kelimeler bazı metin madenciliği görevlerinde bu kapsamda ele alınır.
6. Stemming Nedir?
Stemming, kelimeleri daha basit bir kök formuna indirgemeye çalışan yöntem. Amaç, aynı kökten türeyen farklı kelime biçimlerini benzer bir temsil altında toplamak.
koşuyor → koş
koşacak → koş
koştum → koşStemming kurallı ve hızlı olabilir; ancak ortaya çıkan kök her zaman dilbilgisel olarak gerçek bir kelime olmak zorunda değil. Bu nedenle kullanılan algoritmanın dile uygunluğu ve problemin amacı önemli.
7. Lemmatization Nedir?
Lemmatization, kelimeleri dilbilgisel ve sözlüksel bilgi kullanarak temel sözlük formuna, yani lemma'ya indirmeyi amaçlıyor. Bu yüzden stemming'e göre daha dilbilgisi odaklı bir yaklaşım.
Hızlı ve kaba
Basitleştirilmiş kökler üretir; sonuç dilde gerçek bir kelime olmayabilir.
Dilbilgisi odaklı
Sözlük bilgisini dikkate alarak anlamlı temel formu bulmaya çalışır.
Hangi yöntemin daha iyi olduğu tek bir kurala bağlı değil. Küçük bir metin sınıflandırma probleminde hız önemli olabilirken, dilsel anlamın kritik olduğu bir projede daha gelişmiş normalizasyon yöntemleri tercih edilebilir.
8. TF-IDF Nedir?
Bilgisayar için “metnin anlamı” doğrudan sayısal bir değer değil. TF-IDF, kelimelerin dokümanlar arasındaki ayırt ediciliğini ölçmek için kullanılan klasik ve çok yaygın bir metin temsil yöntemi.
Temel fikir şu: Bir kelime belirli bir dokümanda sık geçiyorsa önemli olabilir; fakat bütün dokümanlarda sürekli geçen bir kelime ayırt edici değildir. TF-IDF bu iki bilgiyi birlikte kullanarak daha seyrek ve ayırt edici kelimelere daha yüksek ağırlık verir.
Term frequency
Kelimenin belirli dokümanda ne sıklıkta geçtiğini dikkate alır.
Inverse document frequency
Kelimenin tüm dokümanlarda ne kadar yaygın olduğunu ölçerek ayırt ediciliğini belirler.
Basit müşteri yorumları örneği
| Yorum | Öne çıkabilecek terimler | Olası tema |
|---|---|---|
| “Kargo çok hızlı geldi.” | kargo, hızlı | Teslimat |
| “Kargo üç günde geldi, çok geç.” | kargo, geç | Teslimat problemi |
| “Uygulama hızlı ama giriş yapılamıyor.” | uygulama, hızlı, giriş | Teknik problem |
TF-IDF doğrudan “bu yorum olumlu” demez. Bunun yerine makine öğrenmesi modeline, metni sınıflandırabilmesi için sayısal özellikler sağlar.
9. Metin Sınıflandırma Nedir?
Metin sınıflandırma, metinleri önceden tanımlanmış kategorilere otomatik atama problemi. Örneğin bir müşteri yorumunu “olumlu”, “olumsuz” veya “nötr” olarak sınıflandırabiliriz.
Duygu sınıfı
Olumlu, olumsuz veya nötr duygu ayrımı.
Konu sınıfı
Yorumun fiyat, kargo, ürün veya destek konularından hangisine ait olduğu.
İstenmeyen içerik
Mesajların spam veya normal içerik olarak ayrılması.
- Ham
Metin - Ön
İşleme - TF-IDF
- Model
- Tahmin
Yaygın kombinasyonlar: TF-IDF + lojistik regresyon, TF-IDF + Naive Bayes, TF-IDF + doğrusal SVM. Burada başarı yalnızca algoritmaya bağlı değil; eğitim verisinin kalitesi, sınıfların dengesi, etiketlerin tutarlılığı ve veri sızıntısının önlenmesi de en az model seçimi kadar önemli.
10. Gerçek Hayat Örneği: Müşteri Geri Bildirimlerinin Analizi
Bir e-ticaret markasının her ay on binlerce müşteri yorumu aldığını düşünelim. Yönetim, “Müşteriler en çok hangi konulardan şikâyet ediyor?” ve “Hangi konular olumlu, hangileri olumsuz?” sorularına hızlı cevap vermek istiyor.
1. Veriyi hazırlama
Yorumlar toplanır; tekrar eden kayıtlar, gereksiz karakterler ve problem için uygun olmayan içerikler temizlenir. Dil, platform ve tarih bilgileri sonraki aşamalarda kullanılmak üzere saklanır.
2. Tokenization ve normalizasyon
Cümleler token'lara ayrılır; gerekiyorsa stop words, yazım varyasyonları ve kök/lemma düzeltmeleri uygulanır.
3. Tema ve duygu analizi
Yorumlar “kargo”, “fiyat”, “ürün kalitesi” ve “müşteri hizmetleri” gibi temalara atanır. Aynı anda sentiment modeli yorumun olumlu veya olumsuz yönünü tahmin eder.
Teslimat
“geç geldi”, “hızlı teslimat”, “kurye” gibi ifadeler.
Kalite
“kaliteli”, “bozuk”, “beklediğimden farklı” gibi ifadeler.
Maliyet algısı
“pahalı”, “indirim”, “fiyat performans” gibi ifadeler.
Müşteri hizmetleri
“cevap alamadım”, “ilgili ekip” gibi ifadeler.
11. Python ile Temel NLP Örneği
Klasik bir örnekte metinler TF-IDF ile sayısallaştırılıp bir sınıflandırma modeliyle kullanılabilir.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
texts = [
"Ürün çok kaliteli ve hızlı geldi",
"Kargo çok geç geldi",
"Uygulama kullanımı oldukça kolay",
"Müşteri hizmetleri sorunu çözmedi"
]
y = ["olumlu", "olumsuz", "olumlu", "olumsuz"]
vectorizer = TfidfVectorizer()
X = vectorizer.fit_transform(texts)
model = LogisticRegression()
model.fit(X, y)
new_text = ["Ürün hızlı geldi"]
new_X = vectorizer.transform(new_text)
print(model.predict(new_X))Bu örneğin amacı üretim ortamında hazır bir NLP sistemi kurmak değil, temel mantığı göstermek: metin → sayısal temsil → model → tahmin. Gerçek bir projede veri setinin büyüklüğüne göre daha kapsamlı ön işleme, model doğrulama ve dile özel yöntemler gerekir.
12. NLP Hangi Alanlarda Kullanılır?
| Alan | NLP kullanımı | İş değeri |
|---|---|---|
| Sosyal medya | Sentiment, konu ve trend analizi | Marka algısını ve konuşulan başlıkları takip etmek |
| Müşteri hizmetleri | Talep sınıflandırma, konu tespiti | Talepleri doğru ekibe daha hızlı yönlendirmek |
| E-ticaret | Yorum analizi | Ürün sorunlarını ve memnuniyeti izlemek |
| Pazarlama | Reklam ve metin analizi | Mesajların hangi temalarla ilişkilendiğini anlamak |
| Arama sistemleri | Sorgu analizi | Kullanıcının aradığını daha iyi eşleştirmek |
| Risk ve uyum | Doküman ve iletişim taraması | Belirli ifadeleri veya risk sinyallerini tespit etmek |
Sosyal medya analitiğinde NLP özellikle değerli; çünkü mention hacmi “ne kadar konuşulduğunu” gösterirken, NLP bu konuşmanın ne hakkında olduğunu ve nasıl bir dil kullanıldığını anlamaya yardımcı oluyor.
13. Klasik Teknikler ve Modern Yaklaşımlar
Klasik temsil
Hızlı, açıklanabilir ve küçük/orta ölçekli sınıflandırma görevleri için güçlü bir başlangıç.
Vektör temsili
Kelimeleri ve metinleri anlamsal ilişkileri taşıyan yoğun vektörlerle temsil eder.
Bağlam odaklı
Bağlamı güçlü yakalayan modern mimariler; sınıflandırma, özetleme ve bilgi çıkarımında kullanılır.
14. İyi Bir NLP Projesi İçin Kontrol Listesi
Net soru
Hangi soruya cevap aranıyor, çıktı ne işe yarayacak?
Kalite
Metinler doğru, güncel ve analiz amacına uygun mu?
Dil uyumu
Türkçedeki ekler, yazım biçimleri ve internet dili ele alındı mı?
Doğru temizlik
Her kelimeyi otomatik silmek yerine probleme göre işlem yapılıyor mu?
Tutarlılık
Sınıflandırma eğitim verisinde etiketler tutarlı mı?
Doğru ölçü
Model yalnızca accuracy ile değil precision, recall ve F1 ile de değerlendiriliyor mu?
Sonuç: Metni Veriye, Veriyi İçgörüye Dönüştürmek
NLP, insanların doğal dille ürettiği büyük miktardaki metni analiz edilebilir hale getirmenin temel yollarından biri. Tokenization, stop words, stemming ve lemmatization metni hazırlarken; TF-IDF gibi yöntemler metni sayısal özelliklere dönüştürüyor. Metin sınıflandırma ise bu temsilleri kullanarak yorumları, mesajları veya dokümanları kategorilere ayırıyor.
Gerçek değer tek tek tekniklerin ötesinde ortaya çıkıyor: sosyal medya yorumlarının duygu ve konu bazında analizi, müşteri geri bildirimlerinin otomatik sınıflandırılması veya binlerce açık uçlu anket cevabının özetlenmesi NLP'yi doğrudan iş kararlarına bağlayabiliyor.
“Metni veriye, veriyi içgörüye dönüştürmek — NLP'nin analitikteki asıl işlevi bu.”
Turning the text people write into structures computers can analyse — that is the essence of NLP. From customer reviews to social media posts, from call centre records to search queries, it is one of the core fields used to make sense of large volumes of text.
In this post I go through the basics — tokenization, stop words, stemming, lemmatization and TF-IDF — along with text classification, and how all of it comes together in a real customer feedback analysis.
“The aim of NLP is not only to process words; it is to surface the signals inside the text so decisions can be made faster and more consistently.”
1. What Is NLP?
Natural Language Processing (NLP) is the field focused on turning the natural language people use into structures that computers can process, analyse and, where needed, interpret.
Human language is not like a structured data table. The same feeling can be expressed with different words, sentences carry different meanings depending on context, and typos, abbreviations, emojis and internet slang make analysis harder. NLP techniques make this messy text machine-processable.
Structure the text
Turn raw sentences into processable structures such as tokens, features or numeric representations.
Find the patterns
Discover word frequencies, themes, sentiment and user behaviour patterns.
Produce decisions
Classify and prioritise texts, or turn them into insights that support a business decision.
2. Why Is Text Data So Valuable?
Alongside structured data, most companies produce a huge amount of unstructured text: social media comments, customer service records, emails, reviews, forum content, search queries and open-ended survey answers.
Comments and posts
Most of what is said about a brand arrives as free text.
Customer feedback
Support records and tickets are the most direct source of experience.
Search text
What users search for is the raw form of their need.
Open-ended answers
They explain the reason behind the numeric scores.
The real problem is less the volume and more that the meaning is scattered. “The cargo came very late”, “the delivery time was terrible” and “the product arrived late” are different sentences that point to the same business problem. NLP helps analyse the similarities and differences between such expressions.
3. The Basic Flow of an NLP Project
An NLP project is not a single algorithm. For a good result, data preparation, text preprocessing, feature extraction and modeling have to be handled together.
- Collect
Data - Clean
- Tokenize
- Extract
Features - Model
- Produce
Insight
4. What Is Tokenization?
Tokenization is splitting a text into smaller pieces, the tokens. In the simplest case the tokens are words; in more advanced systems sub-words can be tokens too.
"The app runs very fast but I have a login problem."
# after tokenization
["The", "app", "runs", "very", "fast",
"but", "I", "have", "a", "login", "problem"]Tokenization is one of the building blocks of the NLP flow; most of the following steps work on these tokens. In agglutinative languages such as Turkish, where words appear in many different forms, tokenization and language preprocessing matter even more.
5. What Are Stop Words?
Stop words are words whose information value may be low for a given task and which are considered for removal from the analysis. Words like “and”, “with”, “a” and “this” are treated this way in some text mining tasks.
6. What Is Stemming?
Stemming reduces words to a simpler root form. The aim is to gather different word forms derived from the same root under a similar representation.
running → run
runs → run
ran → runStemming can be rule-based and fast, but the root it produces does not have to be a real word grammatically. That is why the fit between the algorithm and the language, and the purpose of the problem, both matter.
7. What Is Lemmatization?
Lemmatization aims to reduce words to their dictionary form — the lemma — using grammatical and lexical information. It is therefore a more grammar-aware approach than stemming.
Fast and rough
Produces simplified roots; the result may not be a real word in the language.
Grammar-aware
Uses dictionary knowledge to find a meaningful base form.
Which one is better is not a single rule. Speed can matter in a small classification problem, while a project where linguistic meaning is critical may call for more advanced normalisation.
8. What Is TF-IDF?
For a computer, “the meaning of a text” is not a number by itself. TF-IDF is a classic and very common representation method used to measure how distinctive words are across documents.
The core idea: a word that appears often in a particular document may be important, but a word that appears constantly in every document is not distinctive. TF-IDF combines both pieces of information and gives higher weight to rarer, more distinctive words.
Term frequency
Considers how often the word appears in a given document.
Inverse document frequency
Measures how widespread the word is across all documents, and so how distinctive it is.
A simple customer review example
| Review | Terms that stand out | Possible theme |
|---|---|---|
| “The cargo arrived very fast.” | cargo, fast | Delivery |
| “The cargo took three days, very late.” | cargo, late | Delivery problem |
| “The app is fast but I cannot log in.” | app, fast, login | Technical problem |
TF-IDF does not say “this review is positive”. Instead it gives the machine learning model the numeric features it needs to classify the text.
9. What Is Text Classification?
Text classification is the problem of automatically assigning texts to predefined categories. A customer review, for example, can be classified as “positive”, “negative” or “neutral”.
Emotional class
Positive, negative or neutral sentiment.
Topic class
Whether the review is about price, delivery, the product or support.
Unwanted content
Separating messages into spam and normal content.
- Raw
Text - Pre-
processing - TF-IDF
- Model
- Prediction
Common combinations are TF-IDF + logistic regression, TF-IDF + Naive Bayes and TF-IDF + linear SVM. Success here does not depend on the algorithm alone; the quality of the training data, class balance, label consistency and avoiding data leakage matter at least as much as model choice.
10. A Real-Life Example: Analysing Customer Feedback
Imagine an e-commerce brand receiving tens of thousands of customer reviews every month. Management wants quick answers to “what do customers complain about most?” and “which topics are positive and which are negative?”.
1. Preparing the data
The reviews are collected; duplicates, unnecessary characters and content unsuitable for the problem are cleaned. Language, platform and date information is kept for later stages.
2. Tokenization and normalisation
Sentences are split into tokens; stop words, spelling variations and stem/lemma corrections are applied where needed.
3. Theme and sentiment analysis
Reviews are assigned to themes such as “delivery”, “price”, “product quality” and “customer service”, while a sentiment model predicts whether each one is positive or negative.
Cargo
Phrases such as “arrived late”, “fast delivery”, “courier”.
Quality
Phrases such as “good quality”, “broken”, “not what I expected”.
Cost perception
Phrases such as “expensive”, “discount”, “value for money”.
Customer service
Phrases such as “no reply”, “the relevant team”.
11. A Basic NLP Example in Python
In a classic example, texts are turned into numbers with TF-IDF and used with a classification model.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
texts = [
"The product is great quality and arrived fast",
"The cargo arrived very late",
"The app is quite easy to use",
"Customer service did not solve the problem"
]
y = ["positive", "negative", "positive", "negative"]
vectorizer = TfidfVectorizer()
X = vectorizer.fit_transform(texts)
model = LogisticRegression()
model.fit(X, y)
new_text = ["The product arrived fast"]
new_X = vectorizer.transform(new_text)
print(model.predict(new_X))The point of this example is not to build a production NLP system but to show the basic logic: text → numeric representation → model → prediction. A real project needs broader preprocessing, model validation and language-specific methods depending on the size of the data set.
12. Where Is NLP Used?
| Area | NLP use | Business value |
|---|---|---|
| Social media | Sentiment, topic and trend analysis | Tracking brand perception and the topics being discussed |
| Customer service | Ticket classification, topic detection | Routing tickets to the right team faster |
| E-commerce | Review analysis | Monitoring product issues and satisfaction |
| Marketing | Ad and copy analysis | Understanding which themes the messages connect to |
| Search systems | Query analysis | Matching what the user is looking for more closely |
| Risk and compliance | Document and communication scanning | Detecting certain expressions or risk signals |
NLP is especially valuable in social media analytics: mention volume shows “how much is being said”, while NLP helps understand what the conversation is about and what kind of language is used.
13. Classic Techniques and Modern Approaches
Classic representation
Fast, explainable and a strong starting point for small to mid-sized classification tasks.
Vector representation
Represents words and texts with dense vectors that carry semantic relationships.
Context-aware
Modern architectures that capture context strongly; used in classification, summarisation and information extraction.
14. A Checklist for a Good NLP Project
A clear question
Which question is being answered, and what will the output be used for?
Quality
Are the texts accurate, current and fit for the purpose of the analysis?
Language fit
Have suffixes, spelling variants and internet slang been handled?
The right cleaning
Is the processing shaped by the problem rather than deleting every word automatically?
Consistency
Are the labels in the classification training data consistent?
The right measure
Is the model judged with precision, recall and F1 as well as accuracy?
Conclusion: Turning Text into Data, and Data into Insight
NLP is one of the core ways of making the huge amount of natural-language text people produce analysable. Tokenization, stop words, stemming and lemmatization prepare the text, while methods such as TF-IDF turn it into numeric features. Text classification then uses those representations to sort reviews, messages or documents into categories.
The real value appears beyond the individual techniques: analysing social media comments by sentiment and topic, classifying customer feedback automatically or summarising thousands of open-ended survey answers connects NLP directly to business decisions.
“Turning text into data and data into insight — that is what NLP really does in analytics.”