Job titles are written to attract candidates and to satisfy internal grading. Neither purpose produces a consistent vocabulary. Any analysis that groups by title is making assumptions, and it is worth knowing which.
Four kinds of mess
- Synonyms. Software Engineer, Developer, Programmer, SDE, Member of Technical Staff.
- Seniority baked in. Senior, Staff, Principal, Lead, II, III — inconsistently and not comparably between employers.
- Invented titles. Growth Ninja, Customer Happiness Hero. Rare but never zero.
- Composite roles. "Data Engineer / Analytics Engineer" is genuinely two roles in one requisition.
The three approaches, and what each costs
| Approach | Strength | Failure mode |
|---|---|---|
| Keyword rules | Predictable, debuggable, cheap | Brittle; misses anything unanticipated |
| Embeddings and clustering | Catches synonyms without a list | Merges adjacent-but-different roles; hard to explain |
| Standard taxonomy (ESCO, O*NET, SOC) | Comparable to official statistics | Coarse; lags new roles by years |
There is no fourth option that avoids the trade-off. Anyone who tells you their title normalisation is solved has chosen one of these and stopped mentioning the cost.
The seniority trap
Stripping seniority to group titles is usually right for demand analysis and usually wrong for salary analysis. "Engineer" and "Staff Engineer" belong in the same demand bucket and in very different pay buckets. Any system that uses one grouping for both will produce a median that describes nobody.
We keep seniority as a separate dimension so that a cohort can be grouped one way and split the other.
A practical recommendation
Use rules for the roles you care about and be honest that everything else is a long tail. A curated list of eighty roles you can defend beats an automatic clustering of eight thousand you cannot explain to a customer who asks why two things were merged.