Survey Data Quality in the Age of AI

You’ve probably asked ChatGPT a question. You may even have used it to write a text for you. Artificial intelligence (AI) is pretty good at these things – so, what if we used it to answer survey questions, or help us summarize questionnaire responses?
Researchers have started exploring whether AI can pose as a research assistant, interviewer, or even replace survey respondents. This could make polling much faster and cheaper. However, the inner workings of AI are complex, and researchers have many choices to make when designing AI-based studies. Both of these aspects could undermine the quality of the resulting survey data.
Wahrscheinlich haben Sie ChatGPT schon einmal eine Frage gestellt. Vielleicht haben Sie KI sogar schon einmal dazu genutzt, einen Text für Sie zu verfassen. Künstliche Intelligenz (KI) ist in solchen Dingen ziemlich gut – was wäre also, wenn wir sie dazu nutzen würden, Umfragen zu beantworten oder uns bei der Zusammenfassung von Umfrageantworten zu helfen?
Forschende haben inzwischen damit begonnen zu untersuchen, ob KI als Forschungsassistent*in oder Interviewer*in fungieren oder sogar Umfrageteilnehmende ersetzen kann. Dies könnte Umfragen deutlich schneller und kostengünstiger machen. Allerdings sind die inneren Abläufe der KI komplex, und Forschende müssen bei der Konzeption KI-basierter Studien viele Entscheidungen treffen. Beide Aspekte könnten die Qualität der daraus resultierenden Umfragedaten beeinträchtigen.
DOI: 10.34879/gesisblog.2026.124
What are LLMs?
When talking about AI, what many people really mean are large language models (LLMs), the technology powering chatbots such as ChatGPT. LLMs are a form of generative AI designed to process and generate human language based on relationships between words and statistical probability. For this purpose, LLMs are trained on a large, but selective corpus of text and image data from the Internet – such as book collections, Wikipedia entries, and social media data – that has been annotated (i.e., given descriptions). After initial training, but before they are published, LLMs are often provided with feedback on the responses they generate in an effort to align them with the intended human preferences, like being helpful, harmless, and honest.
When being queried, LLMs analyze the relationships between the words in a user’s input (prompt) based on their learned knowledge from the training data, and then iteratively predict the (most) likely next word. Users can influence the responses using so-called hyperparameters, controlling for example the maximum length of the responses or how creative or unexpected the answers should be.
Because LLMs can be used for a lot of different tasks in multiple languages and do not necessarily require a lot of technical skills, they are an attractive tool for a broad base of researchers. Holtdirk and colleagues (2025) provide an introduction to LLMs for social scientists.
Roles and tasks for LLMs in survey research
In general, three main potential roles for LLMs within the survey research process can be identified: research assistants, interviewers, and respondents. As such, LLMs can potentially be used in all stages of the survey research process – before, during, and after data collection – where they may help address errors impacting data quality while being more efficient than humans. This section lists some examples of possible use cases of LLMs in the survey research process. A recent systematic review of LLM use in survey research (von der Heyde et al., 2026) summarizes the empirical state of the art, including successes and failures.
Pre-data collection: Research assistant tasks
Before data collection, LLMs can act as research assistants by aiding the development of the research design, as well as the development, testing, refining, and translation of the materials used in a survey, including (parts of) the questionnaire and any other materials that are used in the recruitment of potential survey respondents.
Data collection: Interviewer and respondent tasks
During data collection, LLMs can act as interviewers by conducting fully standardized or adaptive, conversational interviews, where they dynamically administer questions based on previous responses, generate new follow-up questions on the fly, and provide examples to support respondent comprehension.
As respondents, LLMs can generate synthetic survey data, usually by simulating human respondents through persona-based prompts that reflect particular demographic or attitudinal profiles. For example, an LLM could be prompted to respond as a 75-year-old, high-school-educated woman who does not believe in climate change. These so-called silicon samples could be used for supplementing and partially replacing human data in pretesting, main data collection, and imputation.
Post-data collection: Research assistant tasks
After data collection, LLMs can assist survey researchers by (pre-)processing, analyzing, and summarizing numerical and text data, for example open-ended responses or social media data. LLMs could also help detect low-quality responses and divergences from interviewer scripts.
Data quality considerations when working with LLMs
Despite LLMs’ potential in developing instruments and collecting and processing survey data, LLM design and research design could introduce new errors and amplify existing biases regarding the understanding of different populations and constructs of interest. These biases can have direct impacts on the quality of survey data collected, generated, processed, and analyzed with the help of LLMs.
LLM design: Training data
LLMs are based on Internet data, which is not generated with the primary goal of producing generalizable estimates and correlates. This may jeopardize both representational and measurement aspects of survey data quality.
Training data coverage – Representation challenges
Representational errors might arise due to the unbalanced content of LLM training data, which likely does not feature the diversity of attitudes and behaviors present in human populations. In addition, bias can be introduced through the workers annotating LLM training data and aligning LLMs through their feedback, especially with regards to their personal backgrounds. As a result of these biases in LLM input, their output likely does not represent all groups or individuals equally, both in terms of scope and quality – outgroup stereotypes may be reproduced. In general, LLMs have been found to exhibit cultural and psychological biases, including a tendency towards reflecting or assuming WEIRD (Western, Educated, Industrialized, Rich, and Democratic) norms and traits, skewing politically left, and being biased against non-English-speaking subjects and marginalized subgroups. This lack of diversity and variance may raise issues when LLMs are tasked to mimic respondents, classify and analyze responses, or during questionnaire design and pre-testing.
Training data validity – Measurement challenges
From a measurement perspective, LLM training corpora do not necessarily reflect actual human preferences, as online data are not always valid indicators of actual attitudes or behaviors. Social media data depends on the platform algorithms and functions, and users might create content strategically or accidentally, expressing themselves in different ways from how they would offline.
Measurement challenges also arise when considering what LLM output technically is: the conditional probability of the previous words being followed by said output. Although the output resembles human-like text, it is unclear whether it also reflects (and therefore can approximate) human cognitive processes or theory-grounded constructs. This puts into question validity. Measurement quality is further complicated because the generated natural language output is not always the numerically most likely response. Thus, whether researchers use the text output at face value or whether they work with the underlying probabilities makes a difference – and the question remains which, if any, is a (more) valid measurement. Finally, the alignment guardrails put in place for general usage might run counter to the researcher’s specific goals, for example regarding designing or asking questions or classifying responses about sensitive topics.
Temporality
In addition, because LLM training data has a cutoff date, LLMs are not by default up to date with current developments, including changes in language use and global political, economic, and social realities. With human attitudes and behaviors fluctuating in response to such exogeneous events, this can lead to representational as well as measurement challenges when using LLMs in survey research, as LLMs may produce output based outdated understandings of attitudes and behaviors. Beyond LLM training data, the probabilistic nature of how LLMs generate text creates challenges for reliability and reproducibility.
Research design
LLM choice
Data quality of LLM-assisted survey research can also be impacted by researcher choices. The most obvious one is the specific choice of LLM. Each LLM is made up of a unique combination of training data and alignment processes, and LLMs vary in their optimization for specific languages or tasks. Therefore, different LLMs may perform differently given the same survey research process task – the question then is not only whether an LLM can perform a task, but which LLM can perform it. This poses a challenge for generalizability claims about which tasks can be augmented by LLMs, as well as for best practice recommendations.
LLM choice can also impact reliability and reproducibility. Especially proprietary LLMs (e.g., OpenAI’s GPT, Google’s Gemini, or Anthropic’s Claude) that cannot be downloaded locally are subject to deprecation, as providers may opt to retire older models, disabling researchers from using them in the long term or re-running their own or other’s previous analyses. Conversely, providers may decide to update still-active models. As a result, previously validated workflows may not be transferable to their updated versions or successors. Proprietary LLMs also raise concerns regarding transparency. Proprietary models tend to be closed “black boxes” with few to no insights into their building process. This makes it difficult to understand the data-generating process, attribute biases, and mitigate them or assess the LLM’s fitness for purpose.
Due to these concerns about extrinsic data quality, there have been calls for exclusively using open-source LLMs (e.g., DeepSeek, Llama, Mistral, Qwen, or OLMo) in research (see Palmer et al, 2023; Spirling, 2023). These models are freely downloadable and offer users greater control because these models can be installed and run on any server accessible to the user, ensuring privacy and reproducibility in results. However, not all researchers may have access to the required computing resources and expertise to run open models in contrast to the user-friendly chat interfaces often offered for proprietary models.
Hyperparameters and prompting
Further, the variability of model hyperparameters potentially inhibits data quality of LLM-augmented survey research. For example, there is a reliability – variability trade-off: while the amount of randomness in LLM outputs can be reduced, thereby increasing reliability, this also reduces variability to a level unlikely found in humans. The exact impact of these variables on the data-generating process within LLMs is opaque, challenging validity.
LLMs’ sensitivity to prompt wording poses another challenge for data quality in LLM-augmented survey research. The type, amount, and order of content in the prompt can impact the output. LLMs’ conversation “memory” is limited, and not all information is treated equally important. Therefore, the employed prompting approach becomes crucial for measurement error.
For a more detailed discussion of data quality considerations of LLMs in survey research, see von der Heyde (2025).
Conclusion
As the relevance of LLM strengths and weaknesses varies across tasks in the survey research process, any survey-related application of LLMs needs to be evaluated for the specific task and context it is to be employed in to safeguard data quality (fit-for-purpose). At GESIS, researchers from different departments work on systematically evaluating and addressing data quality issues in LLM-augmented survey research and provide training and consulting to social scientists for using LLMs in their research (for example at the Competence Center for Data Quality in the Social Sciences – KODAQS). With the necessary knowledge about the potentials and limitations of LLMs, as well as human supervision and validation, it is possible that they can be integrated with more tried and tested tools to provide a better picture of how societies think and act.
This post is based on the conference paper “Who Counts? The Potentials and Pitfalls of Using LLMs in Survey Research” by Leah von der Heyde, presented at the First Workshop on Bridging NLP and Public Opinion Research (NLPOR 2025).
Leave a Reply