How to Build an OSFI E-21 Scenario Testing Program
For Canadian financial institutions, the question around OSFI E-21 is increasingly shifting from what do we need to do? to how do we actually do it?
As of September 1, 2026, federally regulated financial institutions should have identified and mapped their critical operations, established tolerances for disruption, developed a scenario testing methodology and begun testing. By September 1, 2027, scenario testing should be completed for all critical operations.
That gives resilience teams a year to turn methodology into evidence across their critical operations.
And OSFI has provided an important signal about what it expects from that testing. In a 2025 speech discussing supervisory findings, Deputy Superintendent Angie Radiskovic said reliance on tabletop testing was insufficient to ensure critical operations could continue through severe disruption.
So what should an E-21 scenario testing program actually look like?
1. Start with the critical operation, not the scenario
It can be tempting to start with an event:
Let’s run a cyber scenario.
Let’s test a cloud outage.
Let’s simulate a third-party failure.
Under E-21, the better starting point is the critical operation you are trying to keep running.
OSFI defines critical operations as services or products whose disruption could put the institution’s continued operations or safety and soundness at risk, or harm other institutions because of interconnectedness within the financial system.
Once identified, those operations should be mapped end to end across the dependencies required to deliver them, including people, technology, processes, information, facilities and third parties. Importantly, OSFI says this mapping should identify vulnerabilities that can inform scenario testing.
That gives teams a much stronger foundation for scenario design.
Instead of asking:
What happens if our cloud provider goes down?
Ask:
What happens to this critical operation when a technology dependency becomes unavailable - and how long can we continue operating within our established tolerance?
The disruption is the mechanism. The critical operation is what you are testing.
2. Define what you need to learn
A scenario needs a purpose beyond generating discussion.
Before designing the storyline, identify the capability or assumption you want to test.
For example:
- Can the operation continue if a critical technology service is unavailable for 6 hours?
- Can teams execute manual workarounds at the volume required during a peak period?
- Can decision-makers identify when a tolerance for disruption is approaching?
- Can the organization escalate quickly enough when multiple business units are affected?
- Does a third-party contingency plan actually work?
- Can the organization maintain service if several dependencies fail simultaneously?
These questions make scenario design considerably easier.
They also make the results more useful.
E-21 expects scenario testing to help institutions understand when tolerances for disruption would be breached. Results should ultimately inform senior management and board reporting, including identified deficiencies, an assessment of operational resilience and plans to address shortcomings.
A clearly defined testing objective gives you something concrete to assess.
3. Build scenarios from real dependencies
Once the testing objective is clear, scenario design can begin with the dependency map.
OSFI specifically identifies disruptions including power outages, large-scale technology failures, critical third-party disruptions, cyber incidents, natural disasters and pandemics. It also expects institutions to consider concurrent scenarios and disruptions of longer duration. (osfi-bsif.gc.ca)
The most valuable scenarios will often come from combining those events with vulnerabilities already identified in the critical operation.
Imagine a critical customer operation that depends on:
- a cloud-based application,
- an external identity provider,
- a specialist internal team,
- a third-party processing partner,
- and a manual fallback procedure.
A scenario could progressively remove or constrain those dependencies.
The application becomes unavailable.
The third party reports its own operational disruption.
Transaction volumes increase.
The manual workaround begins creating a backlog.
A key decision-maker becomes unavailable.
Customers begin contacting the institution.
At each stage, the question remains the same:
Can the critical operation continue within tolerance?
That is much closer to the end-to-end testing E-21 describes than evaluating each dependency independently.
4. Match the testing method to the question
Not every resilience question needs a full-scale simulation.
And not every resilience question can be answered by a tabletop.
E-21 explicitly anticipates a variety of testing methodologies based on criticality, including tabletop exercises, simulations and live-systems testing. Each can serve a different purpose.
- A tabletop can be effective for exploring governance, escalation and executive decision-making.
- A simulation can expose participants to evolving conditions, introduce consequences based on decisions and capture how individuals and teams respond under pressure.
- Live-system testing can establish whether technical capabilities, failover arrangements and recovery mechanisms work in practice.
The testing portfolio should reflect the questions the institution needs answered.
OSFI’s own supervisory observations reinforce this point. Its concern about overreliance on tabletop testing suggests institutions should think beyond whether exercises are being conducted and toward whether their testing methods provide sufficient evidence that critical operations can persist through severe disruption.
5. Make the tolerance part of the exercise
Tolerance for disruption shouldn’t live separately from the scenario. It should shape it.
Under E-21, a tolerance represents the maximum disruption an institution can withstand under severe but plausible conditions. OSFI notes that this might include measures such as outage time, diminished service, data loss or customer impact.
If a critical operation has a defined tolerance, the exercise should help determine what happens as the organization approaches it.
That could mean introducing elapsed time. Increasing transaction backlogs. Reducing available capacity. Expanding customer impact.Removing another dependency. Or extending the duration of the event.
This creates a more meaningful test than simply asking participants whether they believe they could recover.
It allows teams to explore where the operating model begins to break down and what happens before it does.
6. Test the connections between teams
Critical operations rarely sit neatly inside organizational structures.
A single operation may cross business units, technology teams, risk functions, operations, communications, third parties and executive leadership.
OSFI’s scenario testing expectations explicitly call for an end-to-end approach that considers the total impact across multiple business units and internal and external dependencies. It also says institutions should coordinate with critical third parties, where possible, to conduct broader exercises. (osfi-bsif.gc.ca)
This makes handoffs particularly valuable testing points.
Who notices the problem?
Who owns the decision?
When does escalation happen?
What information reaches the next team?
What assumptions does one function make about another?
What happens when a third party’s recovery timeline does not match yours?
Many resilience weaknesses become visible in the space between documented responsibilities. Scenario testing can make those gaps observable.
7. Capture evidence, not just observations
The exercise itself is only part of the testing process.
The other part is what you can learn from it.
Traditional exercise reporting often relies heavily on facilitator observations, participant feedback and after-action discussion. Those remain valuable, but simulations can produce another source of information: behavioral evidence.
For example:
How quickly was an issue recognized?
When was it escalated?
Which information influenced a decision?
Where did participants hesitate?
Which workarounds were selected?
Where did different teams respond differently to the same conditions?
Which dependencies repeatedly constrained the response?
Combined with operational measures such as outage duration, backlog, capacity and customer impact, this creates a richer picture of how the organization actually responds to disruption.
Over multiple exercises, patterns begin to emerge.
And that matters because E-21 treats scenario testing as an iterative process.
8. Let each test inform the next one
OSFI explicitly expects scenario testing to become more sophisticated over time and says previous test results should inform the design of future tests. Testing frequency and intensity should also reflect the criticality and risk of the operation, with additional testing when significant changes occur in the risk environment.
That creates an opportunity to move beyond an annual exercise calendar.
A test reveals that manual workarounds cannot handle expected transaction volume.
The next test examines that constraint more closely.
A simulation identifies inconsistent escalation decisions across teams.
A future scenario tests a revised escalation model.
A third-party dependency repeatedly creates problems.
A broader exercise includes that provider.
The testing program progressively builds evidence about where resilience is improving and where exposure remains.
From scenario calendar to testing program
With the September 2027 milestone approaching, many Canadian financial institutions will need to conduct a significant amount of testing across their critical operations.
The goal shouldn’t simply be to fit more exercises onto the calendar.
A mature E-21 testing program connects:
Critical operations → dependencies → tolerances → testing objectives → scenarios → evidence → improvement → retesting.
That creates something more valuable than a library of completed exercises.
It creates an accumulating body of evidence about how the organization performs under disruption.
At iluminr, we see scenario testing as an opportunity to turn practice and response into organizational insight. By combining simulations, structured evidence capture and repeatable testing, resilience teams can see where capability is developing, where vulnerabilities persist and where the next test should focus.
For institutions working toward the September 2027 milestone, that may ultimately be the most useful measure of progress: not simply how many critical operations have been tested, but what the organization knows about its ability to keep them running.





.png)