Psychology is a ‘soft’ science. Not because it’s pillowy or velvety to the touch, but because you can’t directly observe what’s going on in people’s minds like you can in ‘hard’ sciences like biology or chemistry or physics1. You can approximate how the mind influences the body (and vice-versa) through cortisol levels in saliva or brain waves. But you still can’t directly see or read thoughts, emotions, beliefs, or personality traits. If you can, let me know. That kind of skillset would come in mighty handy.
If you can’t see it, you have to ask people about it. And unfortunately, people are unreliable and biased on their best day. We can’t predict how we will act in the future because we think that how we feel today is how we’ll always feel, and we mostly remember people from our most recent interactions. We’re incoherent, inconsistent, and incompetent witnesses of ourselves and others.
The power of psychometrics
So how have scientists navigated these constraints for the last 150 years? They had to build their own tools, just like other scientific fields had to build the microscope and telescope and thermometer. Even if psychological states are much more prone to measurement error and inconsistencies than the ‘hard’ sciences, the field of psychometrics has allowed us to understand the human mind and develop therapies, interventions, and products that would have been impossible without it.
Through scientific and statistical rigor, researchers spend years writing, testing, and adjusting survey questions whose only job is to measure one thing. So if you want to measure a construct like burnout, personality, or positive and negative emotions, using published scales like the Positive and Negative Affect Schedule or the NEO Personality Inventory would be your best bet. You would be using an instrument that is highly reliable, valid, and consistent with the rest of the scientific literature.
But I want to observe people in the real world, not the lab
The reality of measuring people in the real world versus in the research lab is that these published scales aren’t always relevant to the context you’re trying to observe. You won’t need that much scientific rigor in your workplace, because you might change the questions a bit (or a lot) over time, your theory of performance or employee engagement or satisfaction might change, and your survey might have to change with it. You’ll have to decide on what is most important at that time, and balance rigor with relevance, usability, and speed.
If I’m trying to see how many peoples are engaged with their work after a senior leadership change or the rollout of new tools, I don’t have time to waste on developing the perfect psychometrically sound scale. It simply isn’t practical. I need the best questions now, that can get me the most accurate results for my goals.
I’m not suggesting that you cut corners and write whatever question comes to mind and call it a day. I’ve seen that happen and it can have underwhelming, if not disastrous effects when it comes time to interpreting results and making recommendations.
No. My goal is to walk you through some simple steps that will help you get the signal you’re looking for, report accurate findings to your leadership teams, and move the needle in the right direction, even if that needle isn’t perfect.
Let’s get cooking.
Be clear about what you want to measure
Before you write any question, decide what your survey will measure and what it won’t. It’s so easy to try and boil the ocean and just “add one more question” out of curiosity to “see what we get”. This ends up bloating your survey and exhausting your respondents. People have limited attention and you want to make sure it’s being used on getting the signal you actually need.
If you want your survey to measure performance but you also have questions about motivation and engagement and emotion, you’re also going to have a really hard time teasing apart what the results mean. If you do want to measure multiple constructs, make sure every one actually points to an answer you need. I suggest organizing them in specific sections, so that questions that are meant to measure the same thing are answered together and your respondents don’t get confused or lose interest halfway through.
Start big, cut down after
Start by writing an oversized question pool. I find it’s much better to start big and cut down later. As you discuss with your team and stakeholders, you’ll get a better idea on which questions matter and which ones don’t. I don’t usually show the full set to people unless they ask, but it gives me something to have in my back pocket in case discussions go in different directions, while still keeping the exercise on target (never lose sight of step number one).
You might be the one holding the pen, but writing questions is as much about incorporating what your leaders care about (including choice of words) as it is writing strong questions. Your job is to get consensus while still applying rigor. Cutting weaker questions later is straightforward, recovering coverage you never wrote is not.
Decide what the results will be used for
The intended use of your questions and survey will determine what evidence is required. Whether scores will be aggregated and/or used downstream in learning programs, promotions, rewards, or other algorithms will help you figure out how many questions to ask, in what format, and with how much statistical rigor.
Questions used to open a discussion or a retro are pretty low bar. A handful of decent questions are enough and their score doesn’t feed into anything other than deciding on topics up for discussion. But a survey used to track employee engagement over time needs to have questions and anchors that stay fixed so you don’t have to wonder if the changes in scores are due to real shifts in engagement or just differences in how you asked the questions each quarter.
I know this might not be the case for everyone reading this, but if you can work closely with a data scientist or people analytics team, they can gut check your assumptions and make sure that questions fed into higher stakes systems like promotion algorithms are rigorous enough. Come prepared, and leverage the right experts at the right time.
Make it reliable
A set of survey questions is reliable when they behave as though they are measuring one thing versus four or five different unrelated things, and this, every time you run the survey. You wouldn’t want your IKEA instructions to build a table one day, and a chair the next. Same goes with your surveys.
You can maximize reliability by the way you write your survey questions. You don’t have to be a scientist running experiments in a lab to get this right. All it takes is a few considerations.
One question, one thing. “My manager supports my growth and gives me feedback I can act on” is two questions in one: supporting growth, and giving actionable feedback. That’s what we call a double-barrelled question. One shot, two bullets. When one half is true and the other isn’t, people pick the middle answer and it becomes really hard to get a clear signal on what they actually think. Watch for ‘and’, ‘or’ in the middle of your question, that’s usually a good indicator it might be double-barrelled (not always though). And make sure you can answer yes or no, high or low for the entire question, not just one part of it.
Bound the recall window. “In the last quarter” gives people something specific to retrieve. Unbounded questions default to whatever happened most recently or most memorably, which means you might be measuring the last two weeks and calling it six months. If you’re measuring personality, you can anchor on a more general time-frame like ‘usually’ or ‘most of the time’. If you’re interested in a specific time window, anchor to that. You don’t have to repeat it for all questions though, that might be repetitive for no reason. If you have sections in your survey, you can write a short preamble asking people to reflect on the last six months, term, quarter, etc.
Drop vague qualifiers and absolutes. Words like ‘often’, ‘regularly’, ‘sometimes’, ‘always’, ‘never’, every time’ are best left out of question bodies, and included in the measurement scales themselves (you know, those friendly 1-5 or 1-7 ‘strongly agree’ to ‘strongly disagree’ that we’ve all answered ample times in our lives). I go into more detail on those below.
Write for a broad reading level and cut jargon. If a question contains an acronym, a framework name, a cultural saying or jargon, some portion of your respondents are more likely answering a question about vocabulary and language comprehension than they are their experience. Unless you’re certain a term is unanimously understood by everyone answering your survey, swap it out for something more universally understood.
Make it valid
Your scale or survey needs to be written and interpreted in the context it was meant for. The same instrument can be valid for one decision and invalid for another. This matters enormously in workplaces, where a scale built for one thing gets repurposed for another and the results simply don’t apply. The validity question should always be: “is this survey good for the decision I’m about to make with it?”. Never lose sight of that.
Construct under-representation. This happens when you’re only measuring part of a thing and calling it reality. If your performance survey asks entirely about speed, you’ll get a clean, reliable, highly consistent measure of speed, and you’ll report it as performance. Always make sure that you get domain experts to read and weigh in on survey questions to make sure you’re not missing anything important.
Asking people to infer rather than report. “My manager values my development” requires the respondent to run an inference about someone else’s internal state and report the conclusion. “In the last quarter, how many times did your manager and you discuss your development?” gets the respondent to observe behavior rather than make an attribution. It’s both easier for people to remember and to observe, and will get you a much more valid signal.
Be careful of reference groups. This isn’t a do or don’t, but be mindful that whenever you ask a question like “compared to your peers”, each respondent will picture someone different as their ‘peer’ unless you specifically define it for them. If you want the cleanest answer, it might be worth anchoring respondents on what you mean by ‘peer’ or ‘region’ or ‘team’ before having them answer comparative questions.
Don’t downplay the response scale
When you build a survey, you’re not just choosing which questions to write. You also need to put as much thought into how your respondents will answer. Are they agreeing or disagreeing? Reporting if they feel something ‘all of the time’ or ‘sometimes’? Are they choosing between 2, 3, 4, 5, or 7 options? All of these considerations will affect how people answer, the quality of signals you get back, and how easy or hard it is for you to interpret results.
Match the anchors to the stem. If you’re asking how often something happened, the answers should be frequencies all the way through (‘never’, ‘a few times’, ‘always’). I’ve seen people start a scale with agreement (‘strongly disagree’) and end with frequency (‘they do this all of the time’). If you do this, results that are low won’t mean the exact opposite of results that are high, like they should. They’ll mean completely different things. That’s messy (dare I say impossible) to interpret clearly, and definitely should never be fed into algorithms or used to make high stakes decisions.
Max five to seven points. You’re not actually differentiating anything extra by adding more numbers to your scales. Gains flatten after seven, because people can’t discriminate more finely than that. So you’d just be adding decimals to your reports that don’t really mean anything other than noise.
Label every point, not just the ends. Ok, I’m guilty of just labeling the ends to make scales more readable and faster to answer. If that’s your goal too, I’d suggest labeling the ends and the middle at least, so people know what it means. If you can label each number though, that’ll always get you clearer signal. Otherwise, people invent their own meaning for 2, 3, 4, etc. As a rule of thumb, leave the least to the imagination as you possibly can.
Decide about the midpoint on purpose. An odd number of scale options allows neutrality (‘I don’t know’ or even ‘I don’t care’), while an even number of scale options forces a choice and elicits clearer opinions, but it can also get people to manufacture an answer when they actually feel neutral. Neither option is better than the other, but you should make this an intentional choice rather than an after thought, because it will affect how you interpret your results.
Document and version it
Maintain an item bank with each item stored with its construct definition, it’s response scale, and the date. You can even store its psychometric performance if you want. If you can, make sure to work closely on this one with your people analytics/data science teams. They ideally should be the ones storing, modeling, and versioning this information.
It’s easy to have surveys drift when questions get edited between leadership changes that bring different aims and priorities. Having a historical record to go back to can help protect longitudinal data and reinforce the importance of certain historical questions that are ultimately there to serve the organization and its employees.
Save this post for later, it’s yours to keep
There’s a few things I haven’t covered in detail today, like the types of biases you might get from your respondents (social desirability, agreement, recency, etc), or how to deal with low response rates and over-inflated questions (everyone answered all 1s or all 5s). There’s so much more to survey design, interpreting results, and taking the right actions from those results, that wouldn’t fit into one essay without being overly-diluted. And if you know me, you know I don’t like to rush through important topics. So check back in soon for deep dives on those subjects.
For today though, I hope you walk away feeling more empowered in your ability to observe and understand people in all sorts of contexts. You don’t need to be a researcher to get this right. You just have to care about how you write questions, and what you plan to use them for. If you’re new to this, it might feel like a lot to remember. You can save this post and come back to it whenever you need. It’s yours to keep.
Until next time, happy writing!
Well, physics skirts both worlds when it comes to its theoretical aspects, but you get the idea.

