Introduction
Some questions about words are counting questions: how often do support tickets mention pricing, what share of reviews raise reliability, which features dominate competitor messaging? Content analysis is the method for answering them: the systematic coding of communication (text, transcripts, media) into categories that can be counted, compared, and tracked. It is where qualitative material meets quantitative discipline. This article covers how content analysis works, the manifest-latent divide, and the reliability machinery that separates measurement from impression.
What is Content Analysis?
Content analysis is a research method for systematically categorising the content of communication (documents, transcripts, open-ended answers, reviews, media) using explicit coding rules, so that the material's features can be quantified and compared. Its classic formulation comes from mid-century communication research (Bernard Berelson's 1952 definition emphasised objective, systematic, quantitative description of manifest content), born in studies of newspapers and wartime propaganda, and its defining trade is stated up front: reduce rich material to defined categories, and in exchange gain counts that support comparison across sources, groups, and time, at scales close reading cannot reach.
Manifest and Latent Content
The method's central distinction. Manifest content is what's literally present and countable with minimal judgment: the word appears, the feature is named, the ticket mentions a price. Latent content is the meaning beneath the surface (frustration, sarcasm, an implied comparison), requiring interpretation that coders can reasonably disagree about. Manifest coding is more reliable and shallower; latent coding is richer and riskier; rigorous studies state which they're doing, and build the reliability checks to match, because "systematic" is precisely what separates content analysis from reading with a highlighter.
The Procedure
1. Define the question, the universe, and the unit.
What's being asked, of which body of material (all Q3 tickets? a defined sample of reviews?), coded at what unit (word, sentence, ticket, review)? Sampling from large corpora follows the same logic as sampling people.
2. Build the codebook.
Categories with names, definitions, inclusion and exclusion rules, and example cases: the instrument itself. Codes can come from theory and prior research (deductive) or from a first inductive pass over a subsample, then get frozen for the systematic run.
3. Pilot, train, and measure agreement.
Independent coders on a shared subsample, disagreements resolved into codebook refinements, and inter-coder reliability reported (Cohen's kappa and relatives): the step that converts opinion into measurement, and the first thing to check in anyone else's content analysis.
4. Code the corpus, then count and compare.
Frequencies, proportions, co-occurrences, and trends over time or across segments, analysed with the standard categorical machinery where inference is claimed.
Content Analysis in Product Research
The product-side corpora are everywhere: open-ended survey answers coded into issue categories and tracked wave over wave; support tickets and app-store reviews quantified into a pain-point league table; competitor sites and onboarding flows coded for claims and features; and open-text or transcribed video answers from studies (the qualitative half of a mixed Ballpark study) converted into countable categories alongside their task metrics. Automation extends reach: keyword rules and machine-assisted classification (including sentiment analysis, essentially automated latent coding of tone) can process volumes no human team can, with the standing requirement that automated codes be validated against human judgment on samples before anyone trusts the dashboard.
Against its neighbours: thematic analysis pursues interpreted meaning and pattern, comfortable with themes that resist counting; content analysis pursues systematic enumeration, comfortable with the reduction that counting requires. Mature teams use both, often on the same material: content analysis for the "how much", thematic work for the "what's really going on".
The Takeaway
Content analysis is measurement applied to words: a defined corpus, a disciplined codebook, agreement you can quantify, and counts you can compare. Declare the manifest-latent line, earn the reliability statistic, validate any automation, and let the enumerable questions be answered at scale, while remembering what was traded for the counts: everything about the material that a category can't hold.
Further reading
For the method's procedure and standards:
Articles:
1. Content Analysis: Guide, Methods and Examples - Scribbr
The full procedure (units, codebooks, reliability, analysis) with worked examples across material types.
2. How to Analyze Qualitative Data: Thematic Analysis - Nielsen Norman Group
The interpretive neighbour, for choosing between counting and meaning-making, or sequencing both.