Why 90% of Your Data Is Actually Working Against You

Neil Sahota, CEO of ACSILabs and United Nations AI Advisor, explains why hoarding data is costing companies more than it’s helping, and why synthetic data is quietly becoming the next big IP battleground.

A very large international retail company Neil Sahota has worked with has never deleted a single product record. They’re sitting on more than 7 billion of them, including products that haven’t been manufactured in 30 years. Ask them why they keep it all, and the answer is always some version of “we might need it one day.” That instinct, the idea that more data automatically means more value, is exactly what Neil says is quietly wrecking companies’ AI efforts right now.

Neil Sahota has been in the AI space for over 20 years. He’s the CEO of ACSILabs, where the focus is cognitive science and using AI and metaverse technology to clone expertise, essentially training people up to the level of your best salesperson or analyst. He also helped launch the United Nations’ AI for Good initiative about a decade ago, which has completed over 500 projects and positively impacted roughly 1.4 billion people. He joined us on Deconstructing Data to talk about why most organizations don’t actually have a data problem. They have a belief problem.

Why does more data make AI worse, not better?

This is the core idea Neil kept coming back to, and it runs against pretty much everything companies have been told for the last decade.

“Most organizations say they don’t really have a data problem, but they have a belief problem. We are used to getting correct answers from computers, and so with AI, the belief is the data is telling the truth when in reality what’s really happening is it’s telling a comforting story.”
— Neil Sahota, CEO, ACSILabs

Here’s the thing. We’ve been told for years that data is the new oil, so companies collected everything and hoarded it, mistaking volume for intelligence. Neil pointed out that a human child only needs to hear about 10 million words to become fluent in a language, while a large language model needs closer to 100 million. It’s not the volume that matters, it’s whether the data is actually the right data.

“Most data collected today is actually creating drag, not value. It’s not that it’s unused, what it’s actually doing is disordering our decision-making. We fall in love with old behavioral signals from customers that may not exist anymore. We look at stale attributes and treat them as facts when they’re really just probabilities.”
— Neil Sahota, CEO, ACSILabs

That’s why Neil says the best AI teams he works with today aren’t obsessed with collecting more. They’re doing what he calls data subtraction.

“The honest truth is the best AI teams I see and work with today, they spend just as much time killing data as they do collecting it.”
— Neil Sahota, CEO, ACSILabs

Why is synthetic data becoming the next IP war zone?

Synthetic data, meaning data that’s generated to look real but isn’t tied to an actual transaction or event, started out solving a real problem. Banks don’t have enough real examples of money laundering to train fraud models on, so they generate synthetic transactions that mimic the patterns. It’s spread from financial services into medical imaging, training diagnostic models on synthetic X-rays and MRIs when real patient data is scarce or locked up.

But Neil’s warning is that synthetic data isn’t neutral, and most companies are treating it like it is.

“Synthetic data is not created equal. It’s not neutral. While it looks like real data, it reflects a lot of the assumptions of whoever is actually creating it. Whatever assumptions, whatever biases we’re baking into it, we could have an IP risk where we’re training models on proprietary behavior patterns. It could leak a strategic advantage.”
— Neil Sahota, CEO, ACSILabs

He compared it to an idea floated years ago about building drones to replace declining bee populations. Researchers knew maybe 10 or 12 things bees do and could replicate those, but bees probably do a thousand things, and nobody knows what happens when you’re only replicating a fraction of the behavior. Synthetic data has the same blind spot baked in.

“I really believe that in about five years, lawsuits won’t be about stolen data anymore. It’ll be about stolen patterns. The way we actually create the synthetic data will become the key IP point out there.”
— Neil Sahota, CEO, ACSILabs

One real example he gave: banks building anti-money-laundering models assumed laundering always involved large transactions. In reality, plenty of laundering happens in transactions as small as $5. Nobody launders money in single wire transfers labeled “I need a million dollars.” That flawed assumption, baked into synthetic training data, would have skewed entire fraud models if it hadn’t been caught.

How does data bias sneak into your models without anyone noticing?

This one hit close to home for me. I brought up something I’d actually run into that week, completely unrelated to AI. Near where I live in Florida, there’s been a push to cull coyotes because they’re raiding sea turtle nests, and a local group’s data showed coyotes were destroying 40% of nests on our beach. Sounds alarming until you zoom out. That beach represents about 20% of the nests on our island, which itself accounts for roughly 1% of Florida’s total turtle nests statewide. Once you look at the whole state, that 40% becomes 0.2% of total nests, a completely different conclusion from the same underlying facts, just because the data set was too narrow.

Neil connected this directly to what happens with synthetic data and, honestly, all data.

“If humans define the rules for synthetic data, bias doesn’t disappear. Humans are flawed, so your synthetic data is going to be flawed regardless.”
— Neil Sahota, CEO, ACSILabs

He backed that up with a genuinely wild historical example: it wasn’t until less than 20 years ago that the medical field realized women show different symptoms during a heart attack than men. Fifty-plus years of cardiac data had been built almost entirely around male symptom patterns, and nobody questioned it because the data looked complete.

What is the data half-life problem, and why does AI make it worse?

This might be the most practical topic we covered. Neil’s premise: most executives assume data depreciates slowly. In reality, some data expires before the meeting discussing it even ends.

“AI does not fix stale data, but it does amplify those errors faster. Once we’ve collected something, we see something, we think it’s always true.”
— Neil Sahota, CEO, ACSILabs

He told a story from his IBM Watson days about a retail industry lead who bought her nephew a tricycle on Amazon for his fifth birthday, a total one-off purchase. Amazon interpreted that as a signal she was in a “parent of young kids” demographic and buried her in toy ads and recommendations for months, even though she never looked at another toy again. The data captured a moment, not an ongoing pattern, and Amazon’s model never expired it.

This connects directly to identity data, which is literally what we do at BDEX, and Neil’s point landed hard here:

“Every piece of identity data, whether it be your email address, your IP address, your mobile ID, your postal address, they all have different half-lives. Your postal address, people don’t move nearly as often as your IP address could change. Your IP address could change weekly.”
— Neil Sahota, CEO, ACSILabs

Data captures a moment, not intent, and that moment has an expiration date that’s different for every data type. A purchase intent signal might be worth acting on for hours. A home-buying intent signal might stay relevant for a month. Treat them the same and you end up wasting money chasing signals that already died, like getting car-delivery marketing emails six months after you already moved the car.

The bigger picture

What struck me most in this conversation is how much it reframes what “good data” even means. For a decade, the industry chased volume: more records, more signals, more history. Neil’s argument, backed by two decades of actually building AI systems, is that volume without an expiration date and without honest interrogation of your own assumptions isn’t an asset. It’s drag. The companies getting real value out of AI right now aren’t the ones with the biggest data warehouses. They’re the ones disciplined enough to figure out what to delete.

That’s the whole game with identity data too. Quality and freshness beat raw volume every time, which is exactly why we built BDEX around solving for that instead of just accumulating more.

To connect with Neil, find him on LinkedIn or check out his work at neilsahota.com.

And if you’re thinking about how stale or bloated data might be working against your marketing right now, that’s exactly what we help fix at BDEX. Visit bdex.com and click “Talk to an Expert” to get started.

Video

Watch the full episode: The Data Lies You’re Living and the AI Secrets That Will Blow Them Up

This article was adapted from an episode of Deconstructing Data, BDEX’s weekly podcast on data-driven marketing. Tune in live every Thursday at 4:15 PM Eastern on LinkedIn.


About BDEX: For companies that need clean identity data to power their products, BDEX offers unmatched quality and execution. Visit bdex.com and click “Talk to an Expert” to get started.