{"id":28221,"date":"2024-07-31T14:17:07","date_gmt":"2024-07-31T18:17:07","guid":{"rendered":"https:\/\/www.forsta.com\/resources\/blog\/synthetic-data-what-you-need-to-know\/"},"modified":"2026-08-31T12:08:43","modified_gmt":"2026-08-31T16:08:43","slug":"synthetic-data-what-you-need-to-know","status":"publish","type":"post","link":"https:\/\/www.forsta.com\/resources\/blog\/synthetic-data-what-you-need-to-know\/","title":{"rendered":"Synthetic data: What you need to know"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">There\u2019s a major buzz around AI in the market research space, and synthetic data is on the tip of everyone\u2019s tongue. Imagine having a dataset that behaves like your target market but doesn\u2019t involve waiting for responses or asking personal questions. That\u2019s the magic of synthetic data.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Naturally, there are concerns and criticisms, namely whether synthetic data comes close to replicating organic human responses.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We\u2019re going to explore what\u2019s got everyone hyped up, and what you might need to keep in mind if you\u2019re going to explore synthetic data.&nbsp;&nbsp;<\/p>\n\n\n\n<h2 class=\"wp-block-heading has-h-3-font-size\">What is synthetic data?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">It\u2019s time to understand exactly what synthetic data is. This is a wide term that refers to information that is artificially created. This can cover many types of data but for research purposes there are a few specific use cases which we will explain later. For now, we can narrow the definition down by saying it\u2019s:&nbsp;<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Artificially generated&nbsp;<\/li>\n\n\n\n<li>Mimics the real world&nbsp;<\/li>\n\n\n\n<li>Can be customized&nbsp;<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">Synthetic data, by nature is manufactured or artificial data. Instead of being collected from real world events like surveys it\u2019s created algorithmically, by AI models. Typically, in the market research world it\u2019s used to augment participant responses from traditionally run primary research or create digital personas (more on these later).&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Tabular and structured synthetic data is the new frontier in market research AI. Computer-generated information has become indispensable in <a href=\"https:\/\/www.forsta.com\/resources\/the-new-era-for-research-agencies\/\" target=\"_blank\" rel=\"noopener\">thi<\/a>s<a href=\"https:\/\/www.forsta.com\/resources\/the-new-era-for-research-agencies\/\" target=\"_blank\" rel=\"noopener\"> new data-driven era<\/a>. It\u2019s cost effective, can be automatically annotated and analyzed, and gets around logistical, some ethical, and privacy issues associated with sensitive or hard to reach target audiences. It\u2019s so powerful that <a href=\"https:\/\/www.gartner.com\/en\/newsroom\/press-releases\/2022-06-22-is-synthetic-data-the-future-of-ai\" target=\"_blank\" rel=\"noopener\">Gartner estimates<\/a>\u202fsynthetic data will overshadow real data for training AI models by 2030.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An important distinction is that synthetic data doesn\u2019t come from nothing. It\u2019s informed by and supported by real-world, real human data. The algorithm you\u2019re using or developing must first learn the patterns, correlations and statistical properties of the training data. The synthetic part expands the original data set giving you options for advanced analysis. You can even go further and test new experiences or questions and see how your target audience would react.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Synthetic data models are more flexible than human users. Unlike a human researcher the models aren\u2019t going to get overwhelmed with a mass-delivery of data. You can create bigger, smaller, fairer or richer versions of the original data in an instant, producing new perspectives and decisions backed by evidence.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This flexibility also blurs the distinction between qualitative and quantitative data. AI&#8217;s language tools can create detailed, descriptive data and measure it accurately in real time.&nbsp;&nbsp;<\/p>\n\n\n\n<h2 class=\"wp-block-heading has-h-3-font-size\">Practical applications<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Digital personas&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The real hot topic is creating personas entirely or mostly from AI. Running these personas through traditional research methods produce new results to work with.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Ever thought \u201cugh, really should have included that in the survey\u201d? Well, it might just be possible. Using existing target audience data, virtual personas will react and respond like the real thing. Used for both qualitative and quantitative data, these digital personas give you an opportunity to gather insights without relying on survey responses, particularly useful for re-working collected data on sensitive topics. &nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">One of the most powerful use-cases of this application is in testing. Trialing hypothesis and checking research designs can slash costs and produce a ready-to-go approach you know will work to get the data you\u2019re after.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Expanding audiences&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Augmenting data is one of the things researchers are most excited about. It\u2019s been around a while, but developments and use have accelerated recently. The idea of supplementing traditional research to reduce research time and get more out of difficult to reach audiences is thrilling.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">AI learns the underlying probability distribution of your sample audience. By identifying these patterns, the models can then generate additional sample members that resemble the original audience. It\u2019s not just analysis, it\u2019s new data points that reflect the answers your target audience would give. This possibility is particularly useful in scenarios where traditional collection is limited or expensive. For example, a hard-to-reach audience like busy parents or collecting data on a sensitive topic like healthcare.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Privacy-safe versions of datasets for sharing&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Even when you have substantial human-gathered data sharing it can be a roadblock. Instead of masking or randomizing to anonymize the results, why not create meaningful copies of sensitive data that reflects all your findings? Synthetic customer datasets can be shared and collaborated on safely without fear of privacy breaches. Because generated data is made from scratch you don\u2019t risk identifying original subjects or losing utility by removing information. All the original patterns of correlation are present, avoiding the so-called privacy-utility trade off from traditional anonymization techniques. Typically, this means the more you anonymize your data, the less useful it becomes. You can avoid this completely with synthetic data.&nbsp;&nbsp;<\/p>\n\n\n\n<h2 class=\"wp-block-heading has-h-3-font-size\">Controversies and criticisms<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">There are some synthetic data cons.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Algorithm limitations&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Large Language Models (LLMs) like ChatGPT work with data to create statistical models of text, but they don\u2019t understand the meaning of the sample or their results. The ultimate question here is, are the results, correct? Tiny human nuances aren\u2019t picked up by AI and sometimes it just can\u2019t handle more complex issues or context.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It should be kept in mind that synthetic data models can only repeat patterns and likely results already found in the sample data. This isn\u2019t to say it can\u2019t find patterns you wouldn\u2019t have noticed yourself or couldn\u2019t extrapolate results into similar situations, but the output is only as good as the input, the original human-first market research results.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Bias amplification&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">AI ethics and bias is a vast topic by itself. In short, AI is likely to show bias, and we can\u2019t fix this. Human-sourced content will naturally contain some bias, it\u2019s an intrinsic part of being human. And AI learns from us. If these patterns are present in the source data, it can be repeated or amplified in AI generated results. For example, the famous case of <a href=\"https:\/\/www.bbc.co.uk\/news\/technology-45809919\" target=\"_blank\" rel=\"noopener\">Amazon\u2019s scrapped recruiting tool<\/a>. Despite not asking for a gender split in results, the AI accidentally learnt that male candidates were preferable. Because of the training data which represented human bias.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Confirmation bias can also result from generated data. After all, the model only has the provided training data to work with so <a href=\"https:\/\/medium.com\/inclusive-software\/insta-personas-synthetic-users-fc6e9cd1c301\" target=\"_blank\" rel=\"noopener\">unexpected results<\/a> or deeper meanings can be missed.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Trying to use AI itself to detect bias falls short because these models have no concept of \u2018right\u2019 and \u2018wrong\u2019, so no place to start and no bias-free human-made training data to build a new algorithm. Machines aren\u2019t great with ambiguity and bias can be subtle. The AI might not understand the data being fed into it, but you should. Inspect your data for bias or collection gaps and acknowledge the relevant issues that may occur in your research.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Reliability and validity&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Two vital words in research. How does AI stack up? Some critics go as far as saying that because AI models don\u2019t understand what they\u2019re saying \u2018synthetic users\u2019 are useless. AI models can only replicate patterns of emotion, not express true feelings when asked for more context. Ultimately, until <a href=\"https:\/\/www.kantar.com\/inspiration\/analytics\/what-is-synthetic-sample-and-is-it-all-its-cracked-up-to-be\" target=\"_blank\" rel=\"noopener\">more studies<\/a> comparing human to synthetic data are published we won\u2019t know how far we can push AI before it\u2019s unreliable and invalid.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Understanding that there\u2019s a chance you aren\u2019t necessarily getting the full width of human emotions or experiences can keep you from being too reliant on generative models or avoiding human-run research.&nbsp;&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Overreliance&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The era of synthetic data has clearly arrived, but some think we\u2019ve been too eager to embrace convenience. With fears that the alluring potential of synthetic data will wipe out traditional research, we need to stay realistic about what it can and can\u2019t do. If we become reliant and blind to potential errors or lack of evidence for synthesized results, decisions will fall flat, and the output could be potentially damaging.&nbsp;&nbsp;<\/p>\n\n\n\n<h2 class=\"wp-block-heading has-h-3-font-size\">Final thoughts<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">News of substituting humans makes for a great headline, but the experts using this tech understand the limitations and that traditional research isn\u2019t going anywhere any time soon. The landscape has changed, and AI can be a powerhouse when used properly. Of course, there is a time and place for synthetic panel studies, and it&#8217;s an exciting development in the world of market research. That&#8217;s why you could use any panel with <a href=\"https:\/\/www.forsta.com\/platform\/customer-experience\/digital-feedback\/\" target=\"_blank\" rel=\"noreferrer noopener\">Forsta Surveys<\/a>, even synthetic panels if required.\u00a0<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The option to augment existing data and wring every opportunity out of a sample using synthetic data is a game changer, cutting costs, time and allowing you to discover more than ever before. But the hard work is in the setup and working with the limitations.&nbsp; Synthetic data is an expansion, an expression of and supportive of traditional data capture, not a replacement. &nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>synthetic data<\/p>\n","protected":false},"author":53,"featured_media":26493,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","disable_featured_image":false,"temp_reference":0,"temp_parent":0,"pathfactory_url":"","webinar_link":"","webinar_link_text":"","start_date":"","start_time":"","end_date":"","end_time":"","pressganey_person_subtitle":"","pressganey_person_linkedin":"","pressganey_person_articles":"","pressganey_person_available":false,"url":"","menu_icon":"","is_mobile":false,"breadcrumb_base":"","author":42531,"co_author":0},"categories":[500],"tags":[506,507,508,509,510,511],"class_list":["post-28221","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-quantitative-research","tag-ai","tag-artificial-intelligence","tag-data-augmentation","tag-digital-personas","tag-market-research","tag-synthetic-data"],"_links":{"self":[{"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/posts\/28221","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/users\/53"}],"replies":[{"embeddable":true,"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/comments?post=28221"}],"version-history":[{"count":3,"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/posts\/28221\/revisions"}],"predecessor-version":[{"id":43647,"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/posts\/28221\/revisions\/43647"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/media\/26493"}],"wp:attachment":[{"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/media?parent=28221"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/categories?post=28221"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.forsta.com\/wp-json\/wp\/v2\/tags?post=28221"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}