Analyzing Text of Unknown Origin

Picture of Christopher DeFelice
Christopher DeFelice
Author

The question we are getting more often, from partners, usually the data-focused teams, is how we account for AI-generated content. They are asking whether we have a filter, something that screens it out before it reaches the analysis.

Start with the number everyone quotes. Half of new articles on the open web are now classified as primarily AI-generated.¹ It measures articles, not conversation.

Conversation is where patient experience lives, and it behaves differently. The first large-scale study of it, covering 51 subreddits and nine million comments, found synthetic text, meaning posts written by AI, present but rare, peaking at 9% in some communities in some months. It could only examine posts long enough to test, roughly 170 words and up.²

Which communities were included matters.

The study included mental health and suicide support forums. Those showed minimal synthetic text.² It did not include rare or chronic disease communities. No published study I can find has looked at them.

Prevalence is the wrong measure. Placement is the right one.

Synthetic posts have a signature. Encouraging and kind, thin on hard detail, thin on what happened to me. And they draw as many upvotes as human comments, sometimes more.²

What makes a patient community worth listening to is the opposite. Symptoms nobody wrote down. One person’s course set against another’s. Detail that is specific, unglamorous, and lived. That is the same tail that erodes first when models train on their own output, a failure known as model collapse.³ I wrote about that dynamic in Authentic Data Is Becoming an Asset Class.

Scale decides whether that matters. In a community producing a few thousand posts a year, whatever the condition, a few percent can be a theme.

So counting stopped being safe on its own. Tally mentions, weight terms, group themes, or rank by engagement across text you cannot trace, and synthetic posts can enter each of those steps unannounced.

So, the filter. The honest answer is that no reliable one exists for text this short. Detection works on long articles, which is why the Reddit study set its floor at roughly 170 words.¹² Three quarters of the comments in that dataset fell below it.

Provenance was a differentiator. It is becoming a floor.

Here is where I might be wrong. Drift monitoring, which watches for statistical shifts in what we ingest, may catch in aggregate what per-post detection cannot. Ours has not flagged that shift, which is worth something and less than it sounds: drift is measured against a baseline, so contamination already in the baseline never trips an alarm. When one does trip we go back and review the documents behind it, which is possible because each one points somewhere. And synthetic text stays rare in health support communities because community norms appear to hold it off, not because anything prevents it. Norms are not a controlled variable.

At web scale, a few percent is noise. At community scale, it is a finding.

At TREND Community, we partner with patient communities and we capture experience from conversations happening online, across therapeutic areas. Origin is documented either way. How we work

References

1. Graphite. AI Now Writes as Many Online Articles as Humans. May 2026. Industry research. https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do

2. La Cava, L., Aiello, L.M., and Tagarelli, A. Machines in the Crowd? Measuring the Footprint of Machine-Generated Text on Reddit. arXiv:2510.07226, October 2025. Preprint. Analysis restricted to texts of at least 250 tokens, roughly 170 words, which cut the corpus from 38.1 million comments to 9.0 million. https://arxiv.org/abs/2510.07226

3. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y. AI models collapse when trained on recursively generated data. Nature 631, 755–759 (2024). https://doi.org/10.1038/s41586-024-07566-y