Most AI Visibility Programs Start With the Wrong Questions
· Signal Harbor · Originally published in Signal Harbor Weekly
AI visibility tools can track hundreds or even thousands of prompts.
That sounds rigorous. But before companies celebrate a visibility score, there is a more basic question they need to answer:
Are these prompts connected to anything a real buyer would actually ask?
A company can perform well across a weak set of prompts and still be missing from the questions that shape real buying decisions. More measurement does not solve that problem. It can simply make the wrong measurement look more precise.
Recognition is not the same as discovery
There is a major difference between asking an AI platform about a company by name and asking it to recommend options within a category.
A branded prompt might ask:
What does this company do?
The company is already inside the question. The platform only needs to recognize and describe it.
An unbranded prompt might ask:
What are the best options for a growing company that needs to solve a specific operational problem?
Now the platform has to decide which companies deserve to enter the conversation.
Research shows that AI systems were likely to recognize products when they were named directly but surfaced them far less frequently in unbranded category questions.
That means a visibility program built mostly around branded prompts may produce reassuring results while missing the harder commercial question:
Does the company appear before the buyer already knows its name?
Real buyers do not always ask simple questions
Traditional search-trained marketers think in short keywords.
AI research is different. People can describe their situation, add constraints, ask follow-up questions, and refine the answer through a conversation.
Recent survey findings suggest that buyers are already changing how they conduct research. Many are using more specific, question-based, and conversational language. More than half of the surveyed buyers included constraints such as budget, required features, compatibility, or a particular use case. Others refined their questions through multiple exchanges.
A vague prompt might ask:
What is the best project-management software?
A more realistic prompt might ask:
What project-management tools are best for a midsized service company that needs scheduling, client communication, mobile access, and simple financial reporting?
The second question is harder to measure, but it is also more valuable.
It reveals whether the AI system understands who the product is for, which problems it solves, and how it compares when a buyer introduces real constraints.
The prompt list should begin with the customer
Companies should not build their entire measurement program with generic questions.
The strongest prompt sets should come from the language buyers already use.
That may include:
questions raised during sales calls;
common objections;
customer interviews;
support conversations;
competitor comparisons;
product-marketing research;
questions asked before someone books a demo or requests a proposal.
The evidence points in a clear direction: if buyers use detailed, contextual questions to research and compare companies, then a useful prompt bank has to reflect that same behavior.
Not every prompt serves the same purpose
A useful prompt bank should cover different parts of the buying process.
Discovery questions test whether the company appears when someone is looking for possible solutions.
Comparison questions test how the company is evaluated against alternatives.
Validation questions test whether AI accurately explains capabilities, limitations, use cases, and trade-offs.
Purchase-stage questions test the facts that may influence a final decision, such as pricing structure, implementation requirements, integrations, availability, or contract terms.
This matters because not every error carries the same risk.
Being absent from a broad awareness question may reduce discovery.
Being inaccurately described during a final comparison may remove a company from the final consideration.
A single visibility score hides all of those differences.
The right questions still need repeated measurement
Choosing commercially meaningful prompts is only the first step.
AI responses can change even when the prompt stays the same. The companies mentioned, the order in which they appear, and the sources used to support the answer may vary across runs and over time.
A recent study found substantial variation in the sources cited when the same prompts were tested repeatedly. This highlights the need for multiple runs and longer measurement windows to produce more statistically reliable results.
This means one screenshot is not enough.
But repeatedly testing weak prompts is not enough either.
A credible measurement program needs both:
questions tied to real buying decisions;
enough repeated observations to interpret the results responsibly.
Without the first, the program measures something commercially weak.
Without the second, it may mistake randomness for progress.
The prompt bank should become a business asset
A strong prompt bank is more than a list of phrases entered into an AI platform.
Over time, it can become a structured view of:
what buyers care about;
which use cases drive consideration;
where competitors control the conversation;
which facts are unclear or outdated;
where the company is misunderstood;
which questions create the greatest commercial risk.
That is much more useful than knowing that a brand appeared in a certain percentage of responses.
AI visibility measurement should not begin with the dashboard.
It should begin with customer language, buying decisions, and the questions that determine who enters consideration.
How Signal Harbor approaches this
At Signal Harbor, we do not begin with a generic list of prompts.
We start by identifying the questions that reflect how buyers actually research, compare, and evaluate companies. We then test those questions repeatedly across major AI platforms to understand where a company appears, how it is described, which competitors are being recommended, and what sources are shaping the answers.
The goal is not to produce another visibility score.
It is to build a measurement system around the questions that carry real commercial weight, then determine whether changes in marketing, content, communications, or public information lead to a sustained improvement.
That is the difference between tracking AI activity and measuring how a company is actually entering buyer consideration.
Sebastian Miller
Co-Founder, Signal Harbor
signalharborconsulting.com
sebastian.miller@signalharborconsulting.com
402-306-2213
References
Sharma, A. P. (2026). The discovery gap: How Product Hunt startups vanish in LLM organic discovery queries. arXiv preprint. https://arxiv.org/abs/2601.00912
Gartner. (2026, January 20). Gartner survey finds only one-third of consumers say GenAI rivals search engines; marketers must optimize for both AI-driven and traditional search. Gartner Newsroom. https://www.gartner.com/en/newsroom/press-releases/gartner-survey-finds-only-one-third-of-consumers-say-genai-rivals-search-engines-marketers-must-optimize-for-both-ai-driven-and-traditional-search
Loktionova, M. (2026, March 5). How AI tools influence the modern buyer journey: A survey of 1,000+ U.S. consumers. Semrush. https://www.semrush.com/blog/ai-tools-the-modern-buyer-journey-study/
Schulte, J., Bleeker, M., & Kaufmann, P. (2026). Don’t measure once: Measuring visibility in AI search (GEO). arXiv preprint. https://arxiv.org/abs/2604.07585
This article first appeared in Signal Harbor Weekly, the Signal Harbor newsletter on AI visibility.
