What AI Gets Wrong About Companies
· Signal Harbor · Originally published in Signal Harbor Weekly
Most AI-visibility reporting starts with one question:
Did the company appear?
That matters, but it is not enough.
A company can appear in an AI-generated answer and still be represented incorrectly. Its location may be outdated. Its services may be misunderstood. Its strongest differentiator may be missing. It may be confused with another company or recommended using evidence that does not actually support the recommendation.
Recent research on large language models shows why accuracy cannot be treated as a simple yes-or-no measurement. An answer can contain correct information while also leaving out important facts, relying on outdated sources, or attaching citations that do not fully support its claims.
This is the problem Signal Harbor is being built to address.
We measure more than whether a company is mentioned. We look at how companies are represented, what information may be shaping the answer, which competitors are recommended instead, and where the public evidence may need to be corrected or strengthened.
Several company-level errors deserve separate attention.
1. Incorrect company information
The clearest problem is when an AI system states something that is simply wrong.
It might list:
a location that does not exist;
a service the company does not provide;
an executive who no longer works there;
an inaccurate certification;
the wrong service area;
or a product feature that was never offered.
These errors are often described broadly as hallucinations.
A 2026 study published in Nature argues that hallucinations are partly an incentive problem. When models are evaluated primarily on whether they produce the expected answer, attempting an answer can be rewarded more than acknowledging uncertainty. The researchers do not claim this explains every hallucination, but their findings help explain why a model may present an uncertain fact with more confidence than the evidence warrants.
For a business, the practical question is not only whether a mistake occurred.
It is:
Where could the mistake be coming from?
An incorrect claim may reflect:
an outdated business directory;
an old page that remains indexed;
inconsistent descriptions across websites;
a third-party article;
or information about a similarly named company.
Those are problems a business can investigate.
Signal Harbor’s role is to identify the incorrect claim, compare it with verified company information, and examine the sources surrounding the answer. That gives the company a clearer path toward correcting the public information it controls and addressing inaccurate third-party listings where possible.
2. Outdated information
Some incorrect AI answers were once accurate.
Companies regularly change their:
pricing;
products;
locations;
leadership;
policies;
hours;
services;
and geographic coverage.
The old information does not always disappear when the business updates its website.
Researchers behind the 2025 HoH benchmark studied what happens when current and outdated information coexist in the sources available to a retrieval-augmented generation system. They found that stale information can reduce answer accuracy and distract models from newer, correct information.
Consider a software company that ended its free plan two years ago. Its current pricing page may be accurate, but older reviews, comparison articles, and directory pages may still describe the discontinued offer.
An AI system could then recommend the product based on a benefit that no longer exists.
This creates customer friction, but it also creates a measurement opportunity.
Signal Harbor is working to identify when company claims appear outdated, trace the public sources that may be reinforcing the old information, and monitor whether the company’s representation changes after those sources are improved.
The relevant question is not simply:
Is the information accurate?
It is:
Is it accurate now?
3. Missing services and differentiators
An answer does not need to contain a false statement to represent a company poorly.
Imagine that a home-care provider offers dementia care, overnight support, and specialized caregiver matching. An AI system may accurately describe it as a home-care company while leaving out all three differentiators.
The description is not necessarily wrong.
But it may still remove the company from consideration when someone asks for specialized dementia or overnight care.
Recent factuality research distinguishes between precision and recall.
Precision asks whether the facts included in an answer are correct. Recall asks whether the answer included the important facts it should have covered.
The 2025 VeriFact research found that strong factual precision does not always correspond with strong factual recall. In other words, an answer may be mostly correct while still leaving out meaningful information.
For companies, this can appear as:
an omitted service;
a missing geographic limitation;
an ignored qualification;
a generic description that leaves out the company’s strongest advantage;
or an accurate company mention that never reaches the final recommendation.
At Signal Harbor, we think of these as representation gaps.
A representation gap may show that an important differentiator is:
mentioned only once;
buried deep inside the company website;
described inconsistently;
absent from trusted third-party sources;
or unsupported by enough clear public evidence.
Those are areas a company can improve.
The company may need clearer service pages, stronger supporting evidence, more consistent descriptions, or credible third-party coverage that confirms what makes it different.
4. Category and company confusion
An AI system may also misunderstand what category a company belongs in or confuse it with another organization.
This is more likely when:
the company has a generic name;
several organizations use similar names;
the company has changed its positioning;
its website describes too many unrelated services;
or different sources classify the company in different ways.
A consulting firm may be described as a software company. A logistics broker may be confused with a shipping carrier. A local business may be mixed with a national company that uses the same name.
This matters because recommendation systems depend on category understanding.
A company cannot consistently appear for the right questions if the system does not clearly understand:
what it offers;
who it serves;
where it operates;
and how it differs from nearby alternatives.
Signal Harbor measures how consistently companies are categorized and described across prompts, models, and repeated observations.
When confusion appears repeatedly, the problem may point toward unclear homepage language, inconsistent business listings, weak category-specific content, or conflicting third-party descriptions.
Improving those signals does not guarantee that every AI platform will immediately change its answers. It does give the company a clearer and more consistent public evidence layer for those systems to interpret.
5. Unsupported or mismatched claims
A citation does not automatically make an AI answer trustworthy.
An AI system may cite a real source that is related to the topic but does not fully support the claim being made.
For example, it might call a product “the fastest in its category” while citing a company announcement that contains no comparison with competing products.
The source exists.
The evidence still does not justify the conclusion.
The 2025 LongCite research evaluated how well language models provide fine-grained citations for long-form answers. The authors found considerable room for improvement and developed methods for connecting individual statements with more precise sentence-level evidence.
For companies, a source mismatch might involve:
using an old article to support current pricing;
treating a company homepage as independent proof of quality;
using a review page to justify a broad superiority claim;
relying on a competitor’s website to describe the company;
or attaching a general industry article to a company-specific conclusion.
This is why Signal Harbor does not treat citation count as the final measurement.
The more important question is:
Does the source actually support what the AI answer says?
When it does not, the company may need clearer documentation, stronger independent validation, more precise public claims, or corrections to the source being relied upon.
A practical framework for measuring company representation
Based on these research areas, Signal Harbor is developing a practical taxonomy for evaluating how companies appear in AI-generated answers.
This is a working measurement framework informed by current research. It is not yet an established academic standard.
Factual error
The answer states something objectively incorrect.
Temporal error
The answer relies on information that is no longer current.
Service or product confusion
The answer misunderstands what the company offers or who it serves.
Geographic error
The answer gives the wrong location, service area, availability, or market coverage.
Pricing or policy error
The answer misstates prices, terms, warranties, eligibility requirements, or policies.
Entity confusion
The system mixes the company with another organization, location, product, or similarly named brand.
Material omission
The answer leaves out a service, differentiator, qualification, limitation, or condition that could affect whether the company is considered.
Unsupported claim
The answer makes a stronger claim than the available evidence supports.
Source mismatch
The source cited or retrieved does not fully support the claim attached to it.
These errors should not be compressed into one generic accuracy score.
An outdated address, an omitted service, and an unsupported recommendation create different business risks. They also require different responses.
An incorrect location may require directory updates.
An outdated product description may require older pages and third-party articles to be corrected.
A missing differentiator may show that the company has not built enough public evidence around an important capability.
Entity confusion may require clearer positioning and more consistent company descriptions.
A source mismatch may reveal that AI systems are relying on weak or inappropriate evidence.
Visibility is only the starting point
A more complete AI-visibility measurement system should answer several questions:
Inclusion: Was the company mentioned?
Accuracy: Were the claims correct?
Freshness: Was the information current?
Completeness: Were important services or differentiators missing?
Category fit: Did the system understand what the company does?
Source support: Did the evidence support the claims?
Competitive position: How was the company framed relative to the alternatives?
This is what Signal Harbor is working to measure.
We do not believe a company should be told only that it appeared in 40% of AI answers.
It should also know:
what those answers said;
what they left out;
whether the information was current;
which competitors appeared instead;
what sources may have shaped the response;
and whether changes to the company’s public evidence improve the result over time.
Being visible is not the same as being understood.
Signal Harbor is building the measurement layer that helps companies see the difference.
Sebastian Miller
Co-Founder, Signal Harbor
signalharborconsulting.com
sebastian.miller@signalharborconsulting.com
402-306-2213
References
Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2026). Evaluating large language models for accuracy incentivizes hallucinations. Nature.
Liu, X., Zhang, L., Munir, S., Gu, Y., & Wang, L. (2025). VeriFact: Enhancing long-form factuality evaluation with refined fact extraction and reference facts. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing.
Ouyang, J., Pan, T., Cheng, M., Yan, R., Luo, Y., Lin, J., & Liu, Q. (2025). HoH: A dynamic benchmark for evaluating the impact of outdated information on retrieval-augmented generation. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics.
Zhang, J., Bai, Y., Lv, X., Gu, W., Liu, D., Zou, M., Cao, S., Hou, L., Dong, Y., Feng, L., & Li, J. (2025). LongCite: Enabling large language models to generate fine-grained citations in long-context question answering. Findings of the Association for Computational Linguistics: ACL 2025.
This article first appeared in Signal Harbor Weekly, the Signal Harbor newsletter on AI visibility.
