Did your additional training budget work? Here's how to test the impact of HR interventions
03.07.2026Suppose your organisation invests extra in training this year. One year later, the employee engagement score increases from 6.8 to 7.3 out of 10. A great result, right?
Maybe. But the most important question is not only whether something changed. The real question is:did that change actually result from the training?
Because in HR, rarely just one thing happens at the same time. Perhaps a new manager was appointed. Perhaps the workload was lower. Perhaps the team composition changed. Perhaps mainly highly motivated employees attended the training. Or perhaps engagement simply increased across the entire organisation, including among employees who did not attend any training.
This is the classic pre/post pitfall: we see a difference between “before” and “after” and quickly label it as impact. But change is not yet proof of impact.
De pre/post-valkuil
A simple pre/post analysis compares the situation before an intervention with the situation afterwards.
For example: Before the training, the average engagement score was 6.8. After the training, the average engagement score was 7.3. Therefore, the training led to an increase of 0.5 points.
Sounds logical, but that conclusion is premature.
A pre/post comparison only shows that a change occurred. It does not show why that change occurred. And that is precisely what makes impact analysis in HR so difficult. That is why you should always be careful with statements such as: “Engagement increased after the training, so the training worked.” A better way to phrase it is:
A better foundation: compare with a control group
A stronger analysis not only looks at the group that attended the training, but compares it with a group that did not receive the training.
At first glance, the training appears to have had an effect of +0.5. But the group without training also improved, by +0.2. So there may have been a general positive trend within the organisation for other reasons unrelated to the training, such as a faster response time to HR questions, a new feedback cycle with managers, or an organisation-wide salary increase.
In that case, the additional improvement attributable to the training is more likely: +0.5 – +0.2 = +0.3
In simple terms, we call this the additional increase associated with the intervention. It is not yet perfect proof, but it is already far stronger than simply comparing “before versus after”. The question therefore becomes: “Did the group that received training improve more than a comparable group that did not receive training?”
The word “comparable” is important. If the training group mainly consists of senior profiles in stable teams, while the control group mainly consists of junior employees in teams with high turnover, you may still be comparing apples and oranges. In that case, it is advisable to investigate further whether you find a similar impact across different job levels and within groups with and without turnover.
An analytical tool: G-computation
For those who want to take it a step further, there are statistical methods that help assess impact more accurately. One of these is G-computation.
G-computation is used to estimate so-called what-if scenarios. Instead of only comparing who happened to receive training and who did not, you build a model that predicts the outcome and simulate two scenarios:
- Scenario A: everyone receives the training
- Scenario B: no one receives the training
For each employee, the model then predicts the expected outcome under both scenarios. You can then compare the average predicted outcome.
This translates a statistical model into a much more understandable HR question: “What would the average engagement (or turnover risk, or performance level, ...) be if everyone received this intervention versus if no one received it?”
Suppose your model predicts that the average engagement score would be 7.27 if everyone received the training, and 6.93 if no one received the training. This results in an estimated difference of +0.34 points. The logic is similar to that of a control group, but more refined.
Why is G-computation useful?
G-computation is particularly interesting because it can account for differences in your data. With control groups, on the other hand, you ideally select a comparable group yourself and manually check differences between those groups to ensure the validity of your results.
In HR, interventions are rarely distributed randomly. People do not receive training “by chance”. Perhaps access is mainly given to employees at a certain job level. Or primarily to teams where the manager strongly promotes development. Or mainly to employees who are already performing well.
G-computation attempts to correct for this by taking the distributions in your data into account. For example, the model can consider differences between employees in terms of department, tenure, baseline engagement, job level or previous performance. This helps make the comparison fairer.
Instead of saying: “The people who received training score higher than the people who did not receive training,” you ask a better question: “What would happen if employees with the same characteristics did or did not receive training?”
This not only makes your analysis statistically stronger, but also easier to understand for HR and business stakeholders.
But G-computation is not magic
G-computation sounds powerful, but it is not a magical causality machine.
You can only interpret the results as more strongly causal if your data, timing and control variables are well structured. For example, the method assumes that you have measured and included all factors that influence both the intervention and the outcome. And that is often where things become challenging in HR. Many important factors are simply not captured neatly in a dataset. Consider:
- an employee’s motivation
- the atmosphere within a specific team
- a manager’s leadership style
- psychological safety
- informal support from colleagues
- personal events outside of work
- the quality of the training itself
If these factors influence both who attends the training and the outcome, but are not measured, bias may still remain.
For example:
Suppose that mainly motivated employees enrol in a training programme. Afterwards, they achieve higher engagement scores. Your model may correct for department, tenure and job level, but if motivation is not measured, part of the difference may still be incorrectly attributed to the training.
That is why interpretation remains important. G-computation helps you make better comparisons, but it does not replace critical reflection or HR context.
What should you best take away from this?
HR interventions deserve better analyses than just a quick before-and-after comparison. An increase in organisational outcomes is interesting, but the real question does not stop at “is there a change?” Also ask yourself: “was that change greater than in a comparable group?”, perhaps even: “what would have happened without the intervention?” and above all, from your role in HR: “Have we taken other explanations and business context into account?”
And that is precisely where the value of HR analytics lies: not in complex models themselves, but in asking better questions before drawing conclusions.
Want to know more?
Would you like help measuring the impact of your HR interventions? Or more inspiration on HR analytics? Be sure to subscribe to our bi-monthly newsletter for tips, trends, and insights, or contact us.