August 1 2023
By Dr. Bill Luker (Synergy Data Science)
Here is a recent post, from the vast LinkedIn commentariat, that raises the often neglected issue of reliability and analytic validity (R&V) in survey data. Since survey data constitutes such a huge proportion of all data collected by business and academic research scientists, it’s important.
A potential problem is that people know less and less about how to conduct surveys well. Conducting a survey is easier than ever, but the same technologies that make surveys easier are also making response bias easier to creep into the results as well. I suspect that we are headed to a disaster of Literary Digest proportions (1936), for many of the same reasons. Of course, the data we have is very huge. But, at least for the problem that we want to analyze, the data is all wrong. Yet, there seems to be a big resistance to cleverly trying to address these problems instead of worshipping blindly at the altar of technology. —Sociologist and Data Scientist, LinkedIn commenter, September 2022.
To conduct surveys more complicated than one-day, one-question pop-ups, and preserve data reliability and validity[1] by ruling out various forms of bias in data collection, is still a painstaking process. My own inspection of Survey Monkey did not show a capacity for item analysis and other reliability tests for survey data.
This leads me to think in terms of the reliability of a measuring instrument, e.g., a questionnaire (survey instrument) administered to gauge attitudes toward brands of jeans. Many data scientists do not appreciate that reliability applies not just to the results of a measurement, but to the way in which a survey question is worded or phrased, i.e., the measurement instrument itself, that can determine data quality.
I’m taking a survey of whether employees wear jeans to work (yes or no), and what “kind” of jeans, as part of a marketing study. Responses are limited to a set of multiple choices. But the word “kind” can refer to a brand of jeans (Levi’s, Lee’s, etc.), or a style of jeans—skinny, boot cut, relaxed fit, and so on. I want to know what style, but I ask respondents to specify, again, the kind of jeans, and give them a choice between several brands.
My own confusion about what I say I want, and what I will get from the survey, enters the heads of some respondents. Some of them think “kind” means style, like you, and are puzzled that they are given brand names from which to choose. Some of them think it means brand names, and are just fine in specifying Levi’s or Lee’s or another brand. But if we were to poll the respondents on that question, asking whether they thought “kind” meant style or brand, there would be varying responses relative to what they believed they were being asked, rather than, necessarily, what I wanted to know. And I was not sure as well. The result? Increased noise in the survey data, and less reliability.
Another example measures reliability in terms of the principles of consistency and repeatability. It also illustrates the criticality of dealing with issues like this at the beginning of any organizational effort to collect and analyze data.
I once sat in a conference with 35 engineers of various stripes doing requirements specifications for a proposed space system. At the end of the first day of brainstorming, I scripted a questionnaire designed to capture data on the consistency and repeatability of the participants’ understandings of requirements they identified and named as critical to a successful system of systems.
We did not have time or space on a short instrument to ask each person to state what was meant by the terminology in Requirement 1, Requirement 2, and so on. But we got to the question of internal consistency by lowering the information requirements of the survey, and going through a logical back door: we asked the respondents to rank their top 20 requirements by importance (1 = highest importance, 20 = lowest).
This isn’t to say we expected each respondent to rank the requirements identically. But if there were consistent and repeatable understandings, i.e., reliable expressions/definitions of requirements and understandings of those expressions/definitions, respondents, on average, would all be ranking the same list. Said another way, everyone responding to the survey would be ranking the same definitions of each requirement.
To test for this, I used a statistic known as Cronbach’s Alpha (a ) that correlated the ranking of each requirement with every other one and averaged the correlations.
a is bounded by 1 and 0. It’s a correlation measure for all the data collected by the questionnaire, or any survey instrument. In this instance, if the same lists were being ranked, and thus the same requirement definitions, a would approach 0.5 or better, telling us that about 50 percent or more of the time that was the case.
In our tests of the requirements rankings, a averaged 0.15, indicating an absence of consistency and repeatability in participants’ understandings of expressions/definitions of the requirements.
The following situation likely prevailed: for Requirement 1, Respondent 1 had one understanding, Respondent 2 another, and Respondent 3 the same as Respondent 1, but different than Respondent 4, etc., and so on. In short, because the respondents had different understandings of each requirement, most of the time each person was ranking a different list. This generated much noise and not much correlation between and among rankings, in an effort demanding that every engineer associated with the project be on the same page throughout its execution.
The engineers had to redefine and refine their terms to eliminate the ambiguity in the wording or phrasing of requirements terminology. Failure to do this in requirements specification on system-engineered projects has (and did have, and does have) disastrous consequences for project and program execution. Every engineer must know they are all talking about the same thing when discussing a given requirement specification.
So, one more time: Measures fail reliability checks the lower their Cronbach’s Alpha (a) score is. And the lower their a’s (and RXXes[2]) the lower their correlations with other measures in the dataset. Low internal correlations mean that for any two or more variables (all considered pairwise), fewer data points or observations may move positively or negatively in tandem. False positives (Type 1 Error), or in this case, false negatives (Type II Error), are the result of degraded consistency and repeatability properties of the measurement, masking more reliable relationships between variables that are better correlated.
Go ask your CIO and your Data Curators: what do you know about all this? How reliable is our data? If you’re one of those who need reliable and valid data (all of us), you’ll know just what to say.
[1] Berman H.B., “Bias in Survey Sampling“, [online] Available at: https://www.stattrek.com/survey-research/survey-bias URL [Accessed Date: 8/3/2023].
[2] “Data Reliability and Analytic Validity for Non-Dummies”, Bill Luker Jr, originally published at https://www.predictiveanalyticsworld.com/machinelearningtimes/data-reliability-and-analytic-validity-for-non-dummies/9623/, and reprinted in an earlier blog posting here.

Leave a Reply