Ask ChatGPT for the monthly search volume of a keyword and it will give you a number. The number will have a plausible shape. It might be 2,400, or 8,100, or 590, which are exactly the kinds of figures a real keyword tool returns.
It is not real. It was not looked up. The model generated a figure that looks like a search volume, in the same way it generates a sentence that looks like a sentence.
This is the single most expensive misunderstanding in AI SEO right now, and it is easy to miss, because a made-up number and a real one are visually identical. There is no error message. Nothing goes red.
So here is the honest answer to whether ChatGPT can do keyword research. It can do half of it brilliantly and the other half not at all, and the half it cannot do is the half people are asking it for. If you want to do keyword research inside ChatGPT and actually trust the output, that missing half is a connection rather than a prompt. ContextBolt SEO is one URL you paste in once, and from then on the chat pulls real search volume, real difficulty scores and today’s live top ten, so the model reasons over numbers it looked up instead of numbers it produced. The rest of this post is what goes wrong without that, how to spot it in ten seconds, and what the workflow looks like once it is fixed.
- Ideas, yes. Numbers, no. ChatGPT is strong at generating and grouping keywords. It has no live volume or difficulty data.
- It will not tell you it is guessing. Ask for volumes and you get confident, well-formatted figures that were invented.
- Better prompts do not fix it. A missing data source is not a phrasing problem.
- Connect real data instead. A custom connector gives ChatGPT live volume, difficulty, SERPs, and your Search Console.
- The data is Ahrefs-grade, decision-useful and directional. ContextBolt SEO is $35/month, with a 7-day free trial.
What ChatGPT is genuinely good at
Start here, because the answer is not “nothing” and the skeptical takes overshoot.
Language work is where it earns its place. Give it a product and it will expand a topic into phrasings you would not have written down, including the awkward ones real people type. Give it a messy 400-row keyword export and it will cluster it by intent faster than you will, and the clustering is usually good. Ask it what question sits behind a term and it is genuinely insightful, because that is a reading-comprehension job.
It is also strong at the step after the decision. Once you know the target, it turns that into an outline, a set of headings, and a brief. That is a writing task and it is what the model is for.
None of these need live data. Every one of them is about words, meaning and structure, which is exactly what a language model has. If you stopped reading here and used ChatGPT only for those jobs, you would get real value and never hit the problem.
Where it breaks, and why
The break is specific. It is not that the model is dumb, and it is not that your prompt was bad.
Three numbers decide whether a keyword is worth writing about. Monthly search volume, which comes from clickstream and ad-platform data. Ranking difficulty, which is computed from the backlink profiles of whoever currently ranks. And the live top ten, which changes weekly and is different in every country.
A language model predicts the next token from patterns in text it was trained on. Search volume was never in that text as a lookup table. Even where a number did appear in some blog post, it was one figure, for one country, at one moment, years before you asked. There is no table inside the model to consult.
So when you ask, the model does the only thing it can. It predicts what a plausible answer looks like. You get a number shaped like a volume figure, with no lookup behind it. Chris Long put it bluntly on LinkedIn. Search Engine Land states it plainly too, in a piece that is otherwise enthusiastic about the tool. ChatGPT does not have access to search volume the way Keyword Planner, Semrush and Ahrefs do.
It is not a fringe complaint either. The top result on Google for “chatgpt keyword research” today is an r/SEO thread asking how reliable it actually is. When the highest-ranking page for a technique is a thread questioning whether it works, that tells you something about the technique.
Difficulty is worse than volume, because it sounds more like an opinion. A model will happily tell you a term is “low competition” based on the words in it. Real difficulty depends on who is ranking and how many domains link to them, which is a fact about the internet today, not a property of the phrase.
Browsing does not fix it
This is where most people land next, and it is a reasonable guess that turns out to be wrong.
Turning on web browsing lets the model open Google and read a results page. That helps a little. It can now see who ranks, at least in the country the request resolved to, at least at that moment.
But Google does not print search volume on the results page. It does not print difficulty. It does not show you the 900 related terms you should have considered, or the 40 domains linking to position three. Reading one page of results is not keyword research, it is one screenshot of one SERP.
Browsing also makes the failure harder to spot, not easier, because now some of the answer is real. The competitor list came from a live page. The volume next to it did not. A half-sourced answer is more convincing than an unsourced one and just as wrong where it matters.
Better prompts do not help
The most-shared advice on this topic is a prompt. There are listicles of four supercharged prompts, and Search Engine Land has a long piece full of prompt patterns that is genuinely well written, and honest enough to say up front that the volume data is not there.
Here is the opinionated part. For the idea-generation half, those prompts are useful and you should steal them. For the numbers half, they are a better way to get made-up figures. No arrangement of words creates a data source. A prompt that says “only use real search volumes” changes the tone of the answer and nothing about where the answer comes from.
You can watch this happen. Ask the same volume question in three separate chats and compare. If the figures move between sessions, nothing was looked up, because a lookup returns the same value twice.
That test takes two minutes and it is the fastest way to convince a skeptical colleague.
How to tell whether a number was looked up
You do not need to take my word for any of this. Four checks will tell you what you are dealing with, in any chat client, in about five minutes.
Ask the same question twice in fresh chats. A lookup returns the same value every time. A generated figure drifts. If “project management software” is 40,500 in one chat and 33,100 in another, nothing was looked up. This is the fastest and most convincing test.
Ask for the country. Real search volume is per country and per language. “Volume for X” with no market attached is not a real figure, and a connected tool will either ask you or state which market it used.
Ask what it called. A connected agent can tell you which tool it used and show the raw result. An unconnected one will describe its reasoning instead, in fluent language, without ever naming a source. Watch for the shift from “I queried” to “based on typical patterns”.
Invent a keyword and ask for its volume. Take a phrase nobody could be searching, something like “purple stapler subscription for cats”. If you get a confident monthly figure back rather than a zero or an admission that there is no data, you have your answer, and it took thirty seconds.
That last one is worth doing in front of anyone who is about to make a content plan out of a chat transcript. It is much harder to argue with than a paragraph of explanation.
How to give it real numbers
The fix is not a prompt. It is a data source, and the plumbing for that is now standardized.
The Model Context Protocol, or MCP, is an open standard for letting an agent call outside tools. Anthropic published it in late 2024, and ChatGPT supports custom connectors that speak it. An SEO MCP server is that protocol pointed at search data.
Once one is connected, the shape of the conversation changes. You ask which of five terms is worth writing about. The model calls a keyword tool, gets real volumes and difficulty scores back, pulls the live top ten, and reasons over numbers it did not invent. The language strengths are still there. They are now sitting on top of facts.
The step-by-step for turning that on is in SEO inside ChatGPT. It is a connector and a URL, not an integration project.
What good looks like once it is connected
The workflow that works is boring and it is the same one a competent SEO has always run. The agent just does the fetching.
You bring a topic, not a title. The agent expands it into candidate terms, which is the part it was always good at. Then it scores the promising ones with real difficulty, pulls the live top ten for the best two or three, and stops.
That stop matters. It shows you a table and waits. You pick the target, because you know things the agent does not, like which SERP your site has any business entering and which client would hate the angle.
Only then does it draft. The draft can be wrong and you will fix it. The target cannot be wrong, because a well-written page aimed at a phrase nobody searches is a wasted month.
The order matters more than it looks. Most people let the agent draft first and check the numbers afterwards, because drafting is the satisfying part. By then the piece has an angle, a title and a shape, and nobody rewrites all three because the volume came back at 20. The check has to happen while changing your mind is still cheap.
It is also worth asking for the live top ten before you commit, not just the volume. A term with decent volume and a first page owned entirely by Amazon, Reddit and three publishers with thousands of referring domains is not an opportunity, whatever the difficulty score says. Seeing the actual ten domains is the fastest way to know whether you have any business being there.
Copy this prompt
Expand this topic into 20 candidate keywords. For the 5 with the most
promise, pull real volume and difficulty, then show me the live top 10
for the best 3. Put it in one table and stop there. Do not draft
anything until I pick the target.
That prompt only works when a data source is connected. Without one it produces the same confident table with invented columns, which is a good demonstration of the whole problem.
The honest limits
Connecting real data solves the invention problem. It does not make the numbers perfect, and I am not going to pretend otherwise.
Search volume from any tool is an estimate. Ahrefs, Semrush and everyone else are modeling from clickstream and ad data, and they disagree with each other on the same keyword. ContextBolt SEO is Ahrefs-grade, not Ahrefs. The figures are decision-useful and directionally right, which is enough to pick a target and walk away from a losing SERP. They are not precise enough to argue a decimal place in a client report.
Your own Search Console is the exception, and it is free. That is real measured data about your site, not an estimate about the market. Connect it and the agent can reason over what actually happened rather than what a model thinks happens.
There is also a cost discipline that catches people in week two. Every real lookup costs something, so an agent told to research forty topics will happily spend a month’s allowance in an afternoon. Give it a shortlist, not a category, and let it expand ideas freely while it only scores the handful you actually care about. Ideas are free. Numbers are not.
The other limit is judgment. A connected agent will still suggest a keyword that is technically winnable and strategically pointless for your business. It does not know your margins, your sales cycle, or that the last three posts on that theme brought traffic that never converted. That part is still your job, and connecting a data source makes it more obvious, not less.
What changes is the floor. You stop making decisions on numbers that were generated to look like numbers. If you are running this inside an always-on agent rather than a chat window, the same server works there, and SEO for agents covers that setup.
Ideas from the model. Numbers from a source. The decision from you.