What a strange correlation can and can't tell you
One of the systems I'm most proud of started with a pile of moderation reports.
About 15 years ago, I was part of building online communities. One of them eventually had close to a million monthly active users, and by active I mean participants, not just readers (or lurkers, as we called them). It was a competitive environment. The tone could get sharp, and the volume of reported content kept growing.
The obvious response was to hire more moderators. We kept the same number of staff as the community grew from 200,000 users to almost a million.
We could do that because we stopped treating every report as equally urgent. We went back through the reports from when we had 200,000 users and looked for details that appeared more often in the troublesome threads. Formatting, word choices, posting history and time of day all gave us small signals.
I've written before about the odd details we found, including the punctuation in a title. Each correlation was weak on its own. We combined hundreds of them into a system that gave moderators a prioritized queue. They could reach a thread while it was still manageable, often before anyone had reported it.
Years later, at another job, I was digging through hundreds of thousands of financial profiles spanning 30 years. The platform scored profiles against partner criteria, and we were searching the data for new signals. I noticed that people who entered their names in lowercase reported, on average, 10% less income than people who used capitals.
The correlation was there. We chose not to use it. It offered no causal explanation, and capitalization would have become a proxy for something much larger than the way a person completed a form. Whatever predictive value it had, we couldn't treat formatting as a fact about the person.
Those two experiences changed how I think about data. In moderation, a collection of strange correlations helped humans decide where to look first. In finance, another strange correlation gave us something we decided to leave alone. Finding a pattern is only the beginning. You still have to decide whether it belongs in a product, and whether you can defend what happens when you use it.
I still wonder how many useful signals sit unnoticed in ordinary data, and how many are better left there.