Methodology
Traditional pre-post surveys have a built-in flaw: participants don't rate themselves consistently before and after a program. Retrospective surveys fix this. Here's how, and why it matters for your evaluation data.
The standard approach to measuring program impact looks reasonable: survey participants before the program, survey them again after, compare the scores. But there's a problem researchers have documented since 1979.
Before a leadership program, a participant might rate themselves 4 out of 5 on "I give effective feedback." They believe they're pretty good at it. After the program, they now understand what effective feedback actually looks like. Using this new, more informed standard, they rate themselves 3 out of 5.
The data shows a decrease. The program looks like it made things worse. But the participant genuinely improved. What changed was their frame of reference, not their ability.
Identified by Howard et al. in 1979, response-shift bias occurs when a learning experience changes how participants interpret the rating scale itself. They recalibrate what "good" means, which contaminates the before-after comparison.
The result: traditional pre-post designs systematically underestimate program impact. Sometimes they show no effect, or even a negative effect, for programs that genuinely worked.
A retrospective survey (also called "post-then-pre" or "then-post") collects both ratings in a single sitting, after the program. For each question, participants rate where they are now and where they were before.
Because both ratings happen at the same time, participants use the same frame of reference for both. Their understanding of what "effective feedback" means is consistent across both scores. The shift you see in the data reflects actual change, not a recalibration of the scale.
Two data collection points. Attrition between rounds. Inconsistent frame of reference.
One data collection point. No attrition. Consistent frame of reference.
In ImpactCheck you pick which questions ask for a before and an after. Those show two columns, and participants rate where they were before the program and where they are now, side by side. Every other scale question stays a single rating, so you can measure shift on the behaviors the program targeted without turning satisfaction questions into something they are not.
"I am confident applying what I learned to my day-to-day work"
That +1.6 shift tells a client something satisfaction scores never can: participants believe their confidence measurably increased as a result of the program.
A five-question retrospective survey for a leadership program, filled in by one participant. Every question asks about something a colleague could watch you do. That is deliberate: a behavior can be rated against a fixed standard, where "how useful was the session" mostly measures the mood in the room that afternoon.
Thinking about how you work now, and how you worked before the program.
1.I give direct feedback to a colleague without avoiding the difficult part
+22.I adapt how I lead depending on what the situation and the person need
+23.I hold people to commitments they have made, without it becoming personal
+24.I make space in meetings for the quieter people to be heard
+25.I check how my team is experiencing my leadership rather than assuming
+16.What is one thing you now do differently?
6 questions · about 3 minutes
One participant's completed survey. In the real form, each scale point is its own row.
The participant fills in both rows after the program, using one yardstick for both. That is what removes the response-shift problem above: they are no longer comparing today's judgment against a version of themselves who did not yet know what good looked like.
Every question is answered twice, so five behaviors is really ten ratings. That is about as many as people will give you properly before they start clicking to reach the end. The last question moved by a single point, which is what real results look like. A set that all improves by the same amount usually says more about the questions than the program.
The scale here is Frequency, one of nine built in. There are others for agreement, confidence, effectiveness and extent, and you can rename the Before and Now headings to whatever suits the program.
Leadership development programs are especially vulnerable to response-shift bias. The whole point of a program is to change how people think about leadership. If it works, participants' standards change. Traditional pre-post surveys penalise programs for doing their job well.
Rohs (1999) found that traditional pre-post designs underestimated leadership program impact by 7-12% compared to retrospective measures. That's the difference between a program that looks mediocre and one that shows clear results.
The retrospective approach maps to Kirkpatrick's evaluation model at Level 2 (learning) and Level 3 (behaviour). It captures whether participants gained new skills and whether they're applying them, using a consistent frame of reference that traditional designs can't provide.
Retrospective mode isn't always the right choice. It depends on what you're measuring.
You want to know: did something change?
Kirkpatrick Level 2 (learning) and Level 3 (behaviour)
You want to know: how was the experience?
Kirkpatrick Level 1 (reaction)
ImpactCheck supports both. Before-and-after is a toggle on each scale question, so one survey can ask for a shift on the two or three behaviors the program targeted and a straight rating on everything else. Leave it off everywhere and you get an ordinary post-program or post-coaching feedback survey.
No method is perfect. The retrospective approach has known limitations, and being upfront about them makes your evaluation more credible.
These are the same limitations acknowledged in the research literature (Hill & Betz, 2005). The consensus is that retrospective surveys produce more accurate self-report data than traditional pre-post designs, while being simpler to administer. For most program evaluation contexts, the trade-off is worth it.
ImpactCheck has retrospective mode built in. Switch it on for the questions that measure change, rename the "before" and "after" columns if the program has its own language for them, and the survey collects both ratings in one sitting.
Start your free trial30-day free trial. No credit card required.