Intelligent Sensemaking
Measuring psychological safety
Measuring safety isn't like measuring temperature. Observing changes the thing we're measuring, people can't always explain why they stay silent, and metrics can hide more than they show. Here's what measurement can honestly do — and a tool to use with your people.
Two ways in. The self-check is for you alone, takes about five minutes and stores nothing. The team survey asks everyone anonymously and gives the whole team one shared picture — most people try the self-check first, then set up the team survey. Both are free, and anonymity is built in — explained below.
What a survey can and can't tell you
It surfaces issues, teaches people what psychological safety is, makes the subject discussable and — done well — encourages the very behaviours it asks about. The survey is an intervention, not a thermometer. That's also its limit: a score tells you where to look, not what is happening or what to do. This instrument pairs every low score with questions about mechanism, so the output starts a conversation rather than becoming a number to monitor.
Distributions, not averages
A team answering 3, 3, 3, 3 and a team answering 1, 2, 4, 5 share an average and need opposite conversations. The first has settled uniformity; the second contains someone having a very different experience from their colleagues — and a team can only be as psychologically safe as its least safe member. Every report shows the full spread, the lowest answer and the disagreement between team members — never a mean as the headline.
Why we don't benchmark
Your results are never compared to other teams — no averages, no percentiles, no "teams like yours". Comparison turns a conversation-starter into a target: competition between teams, complacency when the number is high, anxiety when it's low, gaming in between. Interpretation here is about your team's own next conversation.
Beyond the score: the calculus of voice
A note on how the themes fit together: speaking up is the behaviour psychological safety enables, so we measure it directly — it's the outcome theme. The others (belonging, cohesion, transparency, learning, innovation) are the conditions that make speaking cheap or expensive, and the voice-and-power questions read the mechanism underneath all of it. When someone stays silent, three different things may be happening: they can't predict how speaking would be received (ambiguity), they can predict it and it would cost them (consequence), or they can predict it and nothing would change (futility). Each needs a different response. So this instrument probes for mechanism when scores are low, and measures the power gradient directly — how much what's sayable depends on who's in the room — without asking anyone to reveal their own position. It treats "I'm not sure why I hold back" as a real answer, because the calculus of voice often runs below awareness.
Anonymity has to be architectural
Most tools promise anonymity. This one is built so the promise can't be broken.
No name, email, account or IP address. Answers are stored with a date, never a time, so nobody can correlate an answer with who was at their desk. No receipt or identifier ever goes back to a respondent.
Results are sealed until the survey closes and at least four people have answered — then released once, to everyone at the same time. Comments come back shuffled, detached from the answers their author gave.
When not to survey
Don't survey a short-lived team — a temporary project group hasn't had time together to answer honestly, and a conversation serves it better. Don't survey a team of fewer than about six, where anonymity is arithmetic rather than architecture; talk instead. And don't survey a team in acute conflict expecting the survey to fix it. Measurement is a starting point for teams stable enough to act on what they learn.
Validated research, or practical intervention? Two different jobs
Comparable, validated measurement and context-fitted intervention are different jobs needing different instruments. The question bank here is adapted from Amy Edmondson's work (The Fearless Organization, 2018), extended with sensemaking items of our own. If you're doing research, or you need scores comparable across teams and studies, use Dr Edmondson's original validated 7-item scale, unmodified — validation is precisely the property that adaptation destroys, and her scale is excellent. Adapt (or use ours) when your purpose is intervention in one team: then the right criterion isn't comparability but whether the questions fit your context and start the right conversation. That trade-off is real, and choosing it knowingly is the point.
Ready to look?
Three minutes for yourself, or set up an anonymous survey for the team.
