Eli Goodman, Co-founder and CEO of Datos, a Semrush Company, on where AI training data really comes from, what bad data quietly costs you, and why data quality always comes back to people.
Here’s a number that should change how you read every AI visibility report. According to Eli Goodman, over 90% of LLM searches around the world right now are about productivity. People are asking ChatGPT to fix their PowerPoint, not to recommend running shoes. But the bots that most tools use to measure AI search are mostly asking about brands and shopping.
That gap between what bots do and what humans do was the thread running through this whole conversation.
Eli has been in the data business for 25 years. He started at Gartner, spent almost a decade as Corporate Evangelist at Comscore, ran strategic sales at Jumpshot, and then co-founded Datos in 2020 with two partners in Germany he met over LinkedIn. They didn’t meet in person for 18 months. Datos provides clickstream data from a panel of about 20 million people worldwide, and Semrush acquired a majority stake in late 2023.
“I love data as ingredients. I’m in the sugar business and I like to sell data to people that bake cakes.”
— Eli Goodman, Co-founder and CEO, Datos, a Semrush Company
How does dirty data undermine AI model accuracy?
Eli starts further up the funnel than most people. Before you ask whether data is accurate, ask how it became data in the first place.
“The number one question should always be: where does your data come from and why is it trustworthy?”
— Eli Goodman, Co-founder and CEO, Datos, a Semrush Company
Is it cookie based? Human based? Bot based? Scraped? There are plenty of compliant ways to collect data, and they produce very different results.
His AI search example makes the point. To track brand visibility in LLMs at scale, most companies send bots to run hundreds of millions of queries against ChatGPT, Perplexity, and the rest. That data is useful for building tools. But bots ask about the things advertisers care about, and real people are mostly using LLMs to get work done. If you only look at bot data, you miss what humans actually want.
“There’s the quantitative, like I count things and I just want to make sure I’m showing up. But there’s the qualitative, or almost the philosophical, which is what is the spirit of what somebody’s after. The intent. And that can get lost if you’re just using tools that are purely bot-based.”
— Eli Goodman, Co-founder and CEO, Datos, a Semrush Company
Why do data buyers need to understand how a dataset was collected?
Eli gave a great example of how a small technical quirk turns into a big wrong conclusion. In some LLM interfaces, like Perplexity, the URL refreshes every time you respond in a conversation. Look at the raw clickstream without knowing that, and it looks like one person asked the same question 20 times. Your model decides that’s huge intent. It isn’t.
That’s why Datos tries to be upfront about what its data is good for, what it isn’t, and where people are likely to misread it. If you sell data, you’d better know every wrinkle in it.
This one hit home for me because we see it in identity data constantly. Companies come to BDEX asking why their campaigns are underperforming. We look at what they bought and find one person linked to 15 device IDs. When we filter out bots and click farms, two of those IDs are real. His phone and his iPad. The rest are fraud. The buyer assumed their provider’s data was good because they paid for it.
Eli also made a distinction I really liked.
“Data can be dirty and not be dirty data. Some people are very good at cleansing and prepping and transforming, so maybe that’s how you want it, as raw as possible. I wouldn’t mistake dirty data and bad data.”
— Eli Goodman, Co-founder and CEO, Datos, a Semrush Company
His practical advice is to test the people on the other side of the deal. Datos delivers through cloud storage like AWS S3. If a buyer receives S3 credentials and has no idea what to do with them, that tells you a lot about whether they’re really building an AI model.
What are the hidden costs of bad data?
Eli broke the hidden costs into layers, and the first one surprised me. It’s personnel.
AI is new enough that it’s hard to judge whether the person across the interview table knows what they’re talking about. If you can’t check their work, you might not have anyone who can tell good data from bad. That problem exists before you even buy a dataset.
The second layer is cloud computing. Eli said the monthly cloud bill still blows his hair back. Datos keeps five years of data and keeps growing, and if you buy bad data you still pay to store and process it.
The third is fragility. Build on bad data and you end up with something held together by toothpicks and duct tape. Swapping out a data source later isn’t like swapping commodities. It can mean taking the whole thing offline.
“Do you want it fast? Do you want it right? It generally tends to not be both. So you better pick which one you’re going to go for. And I would rather be right than fast.”
— Eli Goodman, Co-founder and CEO, Datos, a Semrush Company
He also talks to clients about the long game. It’s not the first contract he’s worried about. It’s the next 10. If your data provider only seems focused on getting a signature before year end, that’s a warning sign. We feel the same way at BDEX. The best relationships we have are with customers we’ve grown alongside, not ones we signed and forgot.
How does data quality protect revenue and customer trust?
Eli’s answer here was all about people. If you’re in the big data business, you can’t just turn on a fire hose and walk away. Terabytes of complex data sit on a shelf if nobody helps the customer use it, and then the customer doesn’t renew because they think they got taken.
He has a framework for the people who do that work. Everyone falls into one of two buckets. “Stereo instructions” people thrive with a clear playbook: pull this lever, push this button, follow this script. “Mumblecore” people, named after the improvised film style, do their best work when you give them the high notes and let them riff. In a data business, where AWS goes down and data changes constantly, you need both.
“Somebody does sign off. Somebody has to make a decision. That is somebody’s job. That is somebody’s mortgage. These things don’t exist without the people.”
— Eli Goodman, Co-founder and CEO, Datos, a Semrush Company
His favorite tool reflects that too. Eli uses Chorus AI for meeting notes because it gives everyone a video recording, an audio recording, and a written summary. Some people need to hear it, others need to read it. The notes cover everybody.
The Bigger Picture
Jessie said it best near the end of the episode. We sat down to talk about data and ended up talking about people.
That’s the honest truth of this business. Bots can generate billions of queries, but they don’t tell you what humans actually want. A dataset can look perfect until you learn how it was collected. And the most expensive data problem is often the person who can’t tell the difference. At BDEX, we spend a lot of time removing bots and bad IDs from identity data because we’ve seen what happens when companies skip that step. Eli’s advice applies to every buyer out there: ask where the data comes from, ask what it isn’t good for, and pick right over fast.
Connect with Eli Goodman on LinkedIn or learn more about Datos.
This article was adapted from an episode of Deconstructing Data, BDEX’s weekly podcast on data-driven marketing. Tune in live every Thursday at 4:15 PM Eastern on LinkedIn.
About BDEX: For companies that need clean identity data to power their products, BDEX offers unmatched quality and execution. Visit bdex.com and click “Talk to an Expert” to get started.
Video
Watch the full episode on YouTube: The True Cost of Dirty Data in AI, How Bad Data Undermines Business Success