Keyword Research Tools Compared: The Seed List Problem
Founder, OrganicRank. Builds the crawler, the scoring and the fix pipeline described in these posts. Every claim here is one you can reproduce on your own domain.
The keyword tools are excellent at the job they do. The job they do is expansion: give them a seed, they return everything adjacent to it. Which quietly means the ceiling on your keyword research is whatever you happened to type into the box.
The short version
- Semrush and Ahrefs both have enormous keyword databases and both start from seeds you provide.
- Seed-based expansion inherits your blind spots by design - an audience you never considered produces no seed, so it returns no keywords.
- Companies describe themselves in industry language; customers search in problem language. Seeds written by the company over-index on the first.
- The alternative is derivation: read the site to establish what it sells and who buys it, then find the language those buyers actually use.
How do keyword research tools actually work?
You supply a seed keyword and the tool returns related terms from its index, ranked by volume and difficulty. Every major platform works this way - the differences are database size, freshness and how the difficulty score is modelled.
None of this is a criticism of the databases. Semrush's index in particular is a serious asset that took a decade and a lot of money to build, and if you want to know what a term's volume trend looked like in 2019, nothing we do replaces it.
The limitation is structural rather than technical, and it sits before the database is consulted at all.
| Tool | Entry price | Approach | Needs a seed? |
|---|---|---|---|
| Semrush | $117-140/mo | Largest keyword database, deep historical data | Yes |
| Ahrefs | $29-129/mo | Strong index, clicks data, credit-metered | Yes |
| Moz Pro | $39-49/mo | Smaller index, clearest explanations | Yes |
| Google Keyword Planner | Free with Ads | Advertiser-oriented, banded volumes | Yes |
| OrganicRank | Free audit, $149/mo | Derived from your site and live search | No |
What is the seed list problem?
A seed list is a record of what you already thought of. Expansion tools return variations of your seeds, so any audience, phrasing or market you did not consider cannot appear in the output - it produced no seed to expand from.
The bias is systematic, not random, which is what makes it expensive. Companies write about themselves in the vocabulary of their industry. Customers search in the vocabulary of their problem, often before they know your category exists. A seed list written internally therefore over-represents category jargon and under-represents the symptom-level queries where buying journeys actually begin.
The failure is invisible, too. Your keyword report looks comprehensive - thousands of terms, neatly clustered. Nothing in it indicates that an entire audience is missing, because absence does not show up in a spreadsheet of what is present.
What does derived keyword research look like?
Read the site to establish what it sells and what a customer is worth, infer the distinct audiences from that evidence, map each audience's journey from unaware to ready-to-buy, and only then generate queries - validated against what competitors already rank for.
The output differs in shape as well as content. Expansion gives you a long list ranked by volume. Derivation gives you a sequence: what to target now at your authority, and what becomes reachable once the authority arrives.
- 1Establish the offering and the realistic value of one converted customer.
- 2Infer the audiences - groups with different problems, not demographic slices.
- 3For each, map the questions asked at each stage of the buying journey.
- 4Turn those into queries, including the symptom-level phrasing used before anyone knows your category.
- 5Validate against reality: what do competitors rank for that you missed entirely?
- 6Filter by whether you could realistically rank given your current authority.
Should you rank keywords by volume?
No. Rank by expected revenue - buying intent, times conversion likelihood, times what a customer is worth, times your realistic probability of ranking. Volume is the most visible number and the most misleading one.
A 10,000-search definitional query attracts people who want a definition. A 200-search query with your category, a qualifier and a buying signal attracts people with a budget. Sorting by volume systematically directs effort at the first kind, which is how teams end up with traffic charts that go up and pipelines that do not.
The ranking-probability term is what keeps the plan honest. A valuable keyword you have no chance of ranking for this year is an aspiration, not an opportunity. It belongs in the plan, but not at the top of it.
Sources
Every figure above comes from the vendor's own published material. Check them.
Frequently asked questions
Do I still need Semrush or Ahrefs for keyword data?
For historical trends, precise volume bands and large-scale position tracking, yes. Derivation is better at finding demand you had not considered; a licensed index is better at sizing demand you already know about.
Can a tool really work out my audiences from my website?
It can infer them from what you sell, how you price, who your copy addresses and what your competitors target. Those inferences should be shown with their evidence so you can correct them - which is different from asking you to supply them upfront.
How many keywords should a plan target?
Fewer than most plans list. Thirty commercially meaningful queries ranked properly beats a thousand tracked and none owned.