Python text processing NLTK and pandas
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.
Introduction
NLTK and pandas are complementary tools for text processing in Python. NLTK handles natural language operations — tokenization, stemming, lemmatization, stopword removal, POS tagging, and sentiment analysis. pandas handles the data structure — loading text datasets, applying NLP functions across columns, filtering, grouping, and exporting results. The typical workflow is: load text data into a DataFrame with pandas, apply NLTK transformations using .apply() or .map(), and analyze the results with pandas aggregation.
Setup
Tokenization
Stopword Removal
Stemming and Lemmatization
Sentiment Analysis with VADER
Full Text Processing Pipeline
N-grams and Frequency Analysis
TF-IDF with pandas
Common Pitfalls
- Not downloading NLTK data before using it: Functions like
word_tokenizeandstopwords.words()require downloaded data files. Callnltk.download('punkt')andnltk.download('stopwords')before use, or you getLookupError. - Applying NLTK functions to NaN values in a DataFrame: If a text column contains
NaN,word_tokenize(NaN)raisesTypeError. Filter or fill NaN values first:df["text"].fillna("").apply(word_tokenize). - Using stemming when lemmatization is more appropriate: Stemming produces non-words (
"studies"becomes"studi"). For tasks where readable output matters (search, display), use lemmatization. For bag-of-words classification where exact form does not matter, stemming is faster. - Not specifying the POS tag for lemmatization:
WordNetLemmatizer.lemmatize("better")returns"better"by default (assumes noun). Passpos="a"(adjective) to get"good". For best results, POS-tag words first withnltk.pos_tag()and map tags to WordNet POS. - Loading an entire large CSV into memory before processing: For large datasets, use
pd.read_csv(chunksize=10000)to process in chunks. Applying NLTK to millions of rows at once can exhaust memory.
Summary
- Use NLTK for tokenization, stopword removal, stemming/lemmatization, and sentiment analysis
- Use pandas to load, structure, and aggregate text data with
.apply()and.explode() - Preprocess text by lowercasing, tokenizing, removing stopwords, and lemmatizing
- Use VADER (
SentimentIntensityAnalyzer) for quick sentiment scoring without training data - Combine with scikit-learn's
TfidfVectorizerfor feature extraction in ML pipelines
Related reading
- Python tf-idf-cosine to find document similarity
- Rasa NLU Confidence \`Score\` Computation
- Recommended way to embed PDF in HTML?
- Regular expression matching a multiline block of text
- python tsne.transform does not exist?
- Python weighted median algorithm with pandas
- Python threading. How do I lock a thread?
- Python threads all executing on a single core
.png&w=3840&q=75)
Tackling System Design Interview Problems
A short course that equips you with the skills to approach system design interviews methodically.
Start the free courseNo course covers this one yet
Basis Lab writes one for you from a sentence about what you want to be able to do, then teaches it and checks you understood. Your first course is free.
Build my courseTrack what you have practised
A free account saves your progress, solutions and study plan across every problem on Codemia.
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.