A quick recap. Synthetic data are human-like responses created by AI. That means none of the seemingly human, open-ended responses in a synthetic dataset have actually been written by a real person. They have all been AI generated using LLMs.
The benefits of this are potentially huge: it could reduce the time it takes to collect and curate responses; its cheaper; and it could help you reach niche decision-making audiences. Whilst that all sounds very appealing, the potential problems with synthetic data – from its legality to its inherent biases – are also being pointed out.
Let’s dive into some of the questions at the forefront of the debate around this emerging new method.
Is synthetic data fake?
Synthetic data isn’t exactly fake. But it is artificial. That means it’s not simply invented or faked. If generated in the right way, it can accurately reflect the patterns and correlations of the source data. Companies like Evidenza are claiming that their synthetic data is 88% as similar as traditional market research. But, is 88% accuracy enough? When you’re making a business or brand decision that could cost millions, you want your data to be as accurate as possible.
Is synthetic data bias?
AI image generators like Stable Diffusion have been called out for not only reflecting the racial and gender disparities in the real world, but actually exaggerating and amplifying them. If the synthetic data has a bias towards a particular group, it could mean that your next marketing or brand campaign is missing opportunities by ignoring or misrepresenting certain people.
Is synthetic data legal?
There are well documented problems of the legality of the source data that AI pulls from. This has been particularly controversial with regards to AI image generators which take elements of existing photographs or paintings to create new images, giving rise to a debate around intellectual property infringement. In synthetic data, questions are being raised as to whether data that is not linked to a person can still be treated as personal data. On the flip side, synthetic data can anonymise datasets, ensuring privacy of the individual.
Can synthetic data replicate irrational human desires?
Synthetic data relies on the quality of the source data. If that data doesn’t accurately capture how humans really think or feel about something, then the synthesised market research will fall short. AI works by looking for patterns, creating generalisations – which could mean that individual quirks and unique answers might be ignored.
Is synthetic data right for you?
If synthetic data is used responsibly and can be said to accurately reflect the diversity and individuals of your customer segments, then it could be a good option – especially if your customers are hard to reach and the brand decision you’re making isn’t worth millions.
Kantar is a little more cautious, having carried out its own research: “Our conclusion is that right now, synthetic sample currently has biases, lacks variation and nuance in both qual and quant analysis. On its own, as it stands, it’s just not good enough to use as a supplement for human sample.”
As we’ve mentioned in our insight piece on AI in market research, the best and most accurate way forward might be a blended approach to your next piece of brand or market research – where traditional, human-sourced responses are supplemented by synthetic samples.
Here at saintnicks, we’re exploring different offers and methods with our research partners. If you’re interested in market research or brand tracking – get in touch with our Strategy Director, Mark to talk through how we can help.