← All posts

AI keyword research: how to do it without junk data

By Ben Bowler

AI keyword research: how to do it without junk data

Ask a chatbot for keyword volumes and it will give you a beautiful table. Neat numbers, plausible spread, a difficulty score. Roughly none of it is real.

AI keyword research works when the model does judgement work and a real data source does the counting. Language models are good at grouping keywords into topics, classifying search intent, spotting the questions your list is missing, and drafting briefs. They cannot know how many people searched a term last month, because that number isn’t a fact they were trained on. So the workflow that survives contact with a content calendar is a split: pull volume and difficulty from a keyword tool, hand the list to the model for clustering and intent, and never let the model invent a number you’ll later plan around.

Here’s what that looks like in practice, including where it still goes wrong.

Can AI give you accurate search volume?

No, and it’s worth understanding why rather than just being told.

A keyword tool’s volume figure comes from somewhere specific: clickstream panels, Google Ads data, or a modelled blend of both. It’s an estimate, but it’s an estimate with a method attached. A language model has none of that. When you ask for volume, it produces a number shaped like the numbers it has seen, which is exactly the failure mode that made the SEO community distrust chatbot keyword research the first time round.

The problem got quieter rather than solved. Modern assistants will often refuse, or hedge, or go and search the web instead of answering from memory. That’s better, but the output can still be a confident number sourced from a random blog post from 2021.

The same caution applies to the newer metric everyone wants, which is how often people ask an AI assistant about something. It pays to read how those figures are built. DataForSEO, whose data sits behind a lot of the tools selling AI visibility, documents its method openly: the AI Overviews figure is derived from Google search volume, and the ChatGPT figure comes from an algorithm counting People Also Ask questions that contain your keyword.

Read that again, because it’s the useful bit. The ChatGPT number is a Google-derived proxy. It’s a defensible one, questions people ask Google in natural language do correlate with questions people ask an assistant, but it is not a count of anything typed into ChatGPT. Treat it as directional, quote it as an estimate, and don’t build a business case on the decimal places.

The rule that avoids all of this: if a number will influence a decision, it comes from a tool with a documented methodology. If it’s a judgement about meaning, structure, or intent, the model is genuinely good at it.

What is AI actually good at in keyword research?

Four things, all of which used to eat an afternoon.

Clustering. Give it 400 keywords and it will group them into coherent topics far faster than you’d do it by hand, and better than a pure string-matching approach because it understands that “google ads budget” and “how much to spend on ppc” are the same intent wearing different clothes. For lists under a few hundred keywords, where you’re clustering to plan content rather than to map exact pages, this is the strongest use case there is.

Intent classification. Label each cluster informational, commercial, or transactional, and the content decision falls out of it. Informational gets a guide, commercial gets a comparison, transactional gets a product or service page. Getting this wrong is why well-written posts fail to rank: you wrote a guide for a query that wanted a pricing page.

Gap-finding. Paste your existing titles and the cluster list, ask what a person researching this topic would ask next that you haven’t covered. The answers are usually obvious in hindsight, which is the point.

Brief-writing. Turning a cluster plus the top-ranking pages into an outline with the questions to answer. Not the draft, the brief.

Notice that none of these require the model to know a fact about the world. They’re all reasoning over data you supplied, which is the shape of task where models are reliable.

The workflow

Five steps, and the order matters, because the whole point is that real data enters before judgement does.

  1. Seed and expand in a tool. Start from your product terms and your competitors’ ranking terms. Export with volume, difficulty, and CPC attached. CPC is the underrated column here: it’s the closest free signal you have to commercial intent, because someone is paying that much for a click for a reason.

  2. Filter before the model sees it. Drop the branded terms you can’t win, the volumes too small to matter for your market, and the obviously irrelevant. Feeding 4,000 raw keywords into a chat window buys you a slower, vaguer answer.

  3. Cluster and classify. Hand over the filtered list, keeping the volume column so the model can weight clusters by opportunity, with an explicit instruction not to alter, invent, or extrapolate any figure. Ask for clusters, a primary keyword per cluster, intent, and the page type you’d build.

  4. Sanity-check the clusters against reality. Search the primary keyword for each cluster you’re planning to act on and look at what actually ranks. If the SERP shows product pages and your plan says blog post, the plan is wrong and the model won’t tell you. This step takes ten minutes and saves a quarter’s worth of misaimed writing.

  5. Prioritise on your numbers, not the model’s. Volume from the tool, difficulty from the tool, commercial value from your own CPA maths. The model’s opinion about which cluster is “high potential” is a guess dressed as a recommendation.

If you’re wiring this to an agent rather than doing it in a chat window, the same discipline applies in a more useful form. Several keyword tools now expose MCP servers, so an agent can pull the data itself instead of you exporting a CSV and pasting it. Search Engine Land’s guide to MCP servers for SEO research, published in July 2026, makes the point that matters: you see the same numbers through the MCP server as you do in the tool’s own interface. The provenance survives the trip. The model does the reasoning, the tool does the counting, and nothing gets invented in between.

That’s the version of this workflow worth building, and it’s the same architectural argument as wiring an agent to your ad accounts rather than screenshotting the dashboard at it. Reading is the safe half of both jobs: a keyword pull can be wrong and cost you an afternoon, where a budget write can be wrong and cost you a month.

Where it still goes wrong

Confident clusters built on nothing. If you ask a model to “do keyword research for a project management tool” with no data, you’ll get thirty terms that sound right. Some will be real, some will have no search volume at all, and the two are indistinguishable in the output. Always supply the list.

Intent labels that ignore the SERP. A model classifies from the words. Google classifies from behaviour. “Best crm for small business” reads commercial and mostly returns listicles, but “crm” reads informational and returns product pages. When they disagree, the SERP is right.

Homogenised output. Ask for briefs across twelve clusters in one session and they start rhyming: the same five H2 shapes, the same intro move. Rankings are not the main casualty here, your readers are. Break the pattern deliberately, or run briefs in separate sessions.

Losing the audit trail. Six weeks later, “why are we writing this post?” should have an answer better than “the AI suggested it”. Keep the source export, the cluster output, and the date. If an agent is doing the pulling, this comes free, provided you’re using something that logs its tool calls rather than leaving the record inside a chat history you’ll close.

Optimising for a metric nobody can see. The current temptation is to plan content around AI-assistant visibility. It’s a real distribution channel and worth writing for, but the measurement is immature. Write for the question, keep the answer quotable and self-contained, and treat the tracking numbers as weather rather than instruments.

A prompt worth stealing

Nothing clever, but it’s the one that keeps the numbers honest:

Here are 180 keywords with monthly volume and difficulty from Ahrefs. Group them into topic clusters. For each cluster give me: a primary keyword, the total volume of the cluster, the dominant search intent, and the page type I should build. Do not change, estimate, or add any volume figure. If a keyword doesn’t fit a cluster, put it in “unsorted” rather than forcing it.

The last two sentences do the work. Without them you get tidy clusters, an invented number or two, and nothing in “unsorted” — because a model asked to sort will always find somewhere to put things.


The same rule applies when an agent runs your ads: judgement from the model, numbers from the platform. FlyWheel gives your AI agent one MCP surface across Reddit, Google Ads, Meta, and X, with every tool call logged, and new campaigns shipped paused by default. Get started with FlyWheel.