The Risk of Synthetic Data in AI Discovery
· Signal Harbor · Originally published in Signal Harbor Weekly
AI platforms are quickly becoming a new front door for business discovery.
Buyers are no longer just searching Google. They are asking ChatGPT, Gemini, Claude, Perplexity, and other AI systems to summarize markets, compare companies, and build shortlists.
But those answers are only as trustworthy as the public information they are built on.
So what happens when the public evidence shaping your brand’s AI reputation is fake, manipulated, or synthetic?
That question is becoming more important because the same public web that feeds AI discovery is also being flooded with fake reviews, bot-generated content, AI-written spam, and manipulated reputation signals.
The risk is no longer theoretical. Major platforms are already building enforcement systems to detect and punish this behavior.
The Illusion of Authenticity
One of the biggest problems in today’s digital market is that synthetic content is becoming harder to separate from authentic customer feedback.
Recent research shows that people struggle to identify AI-generated product reviews, performing close to random guessing. In one study, human readers averaged only 50.8% accuracy when trying to distinguish real reviews from AI-generated ones. Even more concerning, large language models were not reliable at detecting fake reviews either. [1]
Humans and AI systems fail for different reasons.
People often become suspicious of reviews that sound too polished, too positive, or too perfectly written, even when those reviews are real. AI systems, on the other hand, may assume content is authentic if it follows the expected structure of a normal review.
That creates a serious problem.
If fake reviews, bot-generated comments, or manipulated third-party content look real to both people and machines, they can begin to shape how AI systems understand a company, category, or competitive market.
Why This Matters for AI Discovery
AI search and recommendation systems rely heavily on public evidence.
That evidence can include company websites, review platforms, comparison sites, forums, business directories, news articles, social content, and other third-party sources.
If those sources are incomplete, outdated, or manipulated, the AI’s answer can be distorted.
This matters because many AI systems use retrieval-based methods to gather information before generating a response. In simple terms, the system searches for relevant public information, pulls from those sources, and then uses that information to form an answer.
That creates a new kind of reputation risk.
If synthetic or manipulated content becomes part of the public evidence layer, it can influence how AI systems describe companies, compare competitors, and recommend options to buyers.
Security research on retrieval-augmented generation has shown that manipulated text can be inserted into public or external information sources in ways that influence the answer a model produces. One recent paper on covert poisoning attacks found that these attacks can manipulate RAG systems with high success under certain retrieval settings. [2]
In a commercial setting, this means fake or low-quality public signals could affect whether a company is recommended, overlooked, or misrepresented inside AI-generated answers.
The Platform Pushback Is Already Here
The temptation for some companies will be obvious: create fake reviews, mass-produce AI content, flood forums, or manipulate public signals in an attempt to influence AI recommendations.
That is a dangerous strategy.
Major platforms are already moving in the opposite direction.
Google has made clear that its spam policies apply not only to traditional search rankings but also to attempts to manipulate generative AI responses in Google Search. This matters because AI overviews and AI-generated search experiences are increasingly connected to the same public information layer businesses have relied on for years. [3]
Google specifically targets scaled content abuse, including the use of generative AI tools to produce large amounts of low-value or unoriginal content designed primarily to manipulate visibility. [3]
The penalties can be serious.
Sites that violate these policies may rank lower, lose eligibility for search features, or be removed from search results entirely. [3]
That means the risk of synthetic content is not just that it may fail. The bigger risk is that it can damage the same visibility infrastructure a company depends on.
Fake Reviews Are Also Being Policed
The same issue exists on review platforms.
Google Maps prohibits fake engagement, including reviews that are not based on genuine experiences, reviews posted from multiple accounts, and content created through methods designed to mimic genuine engagement. It also prohibits businesses from offering incentives in exchange for positive reviews, review revisions, or the removal of negative reviews. [4]
The penalties go beyond deleting suspicious reviews.
Platforms can suspend account privileges, terminate accounts, and, in serious cases, take further enforcement action. [4]
For businesses, this should be a clear warning.
Fake reviews may appear to create short-term credibility, but they can create long-term visibility and trust problems. If a platform removes reviews, flags suspicious activity, or penalizes the account behind them, the business may lose more than the fake signal. It may lose credibility with both the platform and the buyer.
Reddit and Public Sentiment Are Part of the Same Problem
Reddit is another important part of this conversation because it has become a major source of public sentiment.
Buyers use Reddit to research companies, products, software, services, and local recommendations. AI systems also use public discussions across the web to understand how people talk about brands and categories.
But Reddit is not an open field for manipulation.
Research on Reddit moderation found that content from throwaway accounts is more likely to violate rules and be removed by moderation than content from standard pseudonymous accounts. The same research also notes that Reddit moderation involves both human moderators and automated tools. [5]
That matters because synthetic discussion is not just spam. It can become part of the public reputation layer that shapes buyer perception and AI interpretation.
A fake Reddit thread, a manipulated review profile, or a network of low-quality AI-written posts may seem small in isolation. But when AI systems begin summarizing public sentiment at scale, those signals can become part of how a market is explained.
The Pushback Is Bigger Than Tech Platforms
The pushback against fake reviews is also moving into law and regulation.
In the United States, the Federal Trade Commission finalized a rule targeting fake reviews and testimonials, including reviews by people who do not exist, people who did not have real experience with a product or service, and reviews generated through artificial intelligence. [6]
The United Kingdom has also moved to ban fake online reviews through consumer protection reforms. The broader signal is clear: fake reviews are no longer just a platform policy issue. They are becoming a legal and regulatory issue as well. [7]
That should change how businesses think about AI visibility.
The goal should not be to create more synthetic signals. The goal should be to build a cleaner and more trustworthy public evidence layer.
AI Search Systems Are Building Technical Defenses
There is also a technical side to the pushback.
AI search systems that use retrieval-augmented generation are being tested against poisoning attacks. These attacks attempt to insert manipulated or adversarial content into the information layer so the AI system retrieves it and produces a distorted answer. [2]
To defend against this, researchers and engineers are exploring filters that detect unnatural writing patterns, repetitive content, duplicate text, and other signs of synthetic manipulation.
This does not mean the problem is solved.
It means the ecosystem is moving toward greater scrutiny.
As AI search becomes more important, the platforms and systems behind it will likely become more aggressive about identifying low-quality, synthetic, or manipulative content.
That creates a clear strategic choice for businesses.
They can try to manipulate the system and risk penalties, distrust, and inaccurate AI representation.
Or they can build durable visibility through authentic public evidence.
What This Means for Your Business
Fake reviews, bot-generated content, and synthetic reputation signals may create short-term visibility, but they introduce serious long-term risk.
The danger is not only that platforms may penalize the behavior.
The deeper risk is that AI systems may absorb manipulated or low-quality public information and use it to shape how buyers understand your company.
That can affect whether your business is recommended, ignored, or misrepresented at the exact moment a buyer is forming a shortlist.
This is why synthetic data is not just a content problem.
It is a discovery problem.
It is a reputation problem.
And eventually, it can become a revenue problem.
The Takeaway
You cannot build durable AI visibility on fake signals.
Trying to hack AI systems with synthetic reviews, spammy content, mass-produced AI pages, or manipulated public mentions creates risk across three fronts: platform penalties, buyer distrust, and inaccurate AI representation.
The better strategy is authentic public evidence.
Businesses should focus on:
Clear website messaging
Real customer reviews
Legitimate case studies
Accurate third-party mentions
Consistent business information
Content that reflects real buyer questions and language
Before optimizing for AI visibility, companies need to measure how AI systems currently perceive them.
Are they being recommended?
Are competitors showing up instead?
Are AI systems describing them accurately?
What public sources are shaping the answer?
In the new discovery layer, authentic evidence is the strongest defense against synthetic noise.
The companies that understand this early will have an advantage. Not because they trick AI systems, but because they give them better, cleaner, and more trustworthy information to work with.
Sebastian Miller
Co-Founder, Signal Harbor
signalharborconsulting.com
sebastian.miller@signalharborconsulting.com
402-306-2213
Research Note
The ideas in this issue are informed by peer-reviewed research, public studies, platform policies, and Signal Harbor’s ongoing analysis across major AI platforms. As this field develops, Signal Harbor will continue evaluating new research and refining how AI recommendation performance should be measured.
Signal Harbor Weekly helps business leaders understand how AI systems represent companies, influence discovery, and shape buying decisions.
Sources & References
[1] Meng, W., Harvey, J., Goulding, J., Carter, C. J., Lukinova, E., Smith, A., Frobisher, P., Forrest, M., & Nica-Avram, G. (2025). Large Language Models as “Hidden Persuaders”: Fake Product Reviews are Indistinguishable to Humans and Machines. arXiv.
[2] Li, C., Zhang, J., Cheng, A., Ma, Z., Li, X., & Ma, J. (2025). CPA-RAG: Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models. arXiv.
[3] Google Search Central. Spam Policies for Google Web Search.
[4] Google. Prohibited & Restricted Content: Maps User-Generated Content Policy.
[5] Guo, C., & Caine, K. (2025). Throwaway Accounts and Moderation on Reddit. arXiv / Symposium on Usable Security and Privacy.
[6] Federal Trade Commission. Final Rule on the Use of Consumer Reviews and Testimonials.
[7] UK Government / Digital Markets, Competition, and Consumers Act 2024. Consumer protection reforms addressing fake reviews and online deception.
This article first appeared in Signal Harbor Weekly, the Signal Harbor newsletter on AI visibility.
