Skip to Content

What a Resilient Organization Actually Looks Like, and How One Is Built

Operational Excellence

At around eight in the evening on 17 March 2000, a lightning strike hit a power line in New Mexico, causing fluctuations across the state grid. A resulting surge started a fire in Fabricator No. 22 of a Philips semiconductor plant in Albuquerque. Staff smothered the flames within ten minutes, but sprinkler water and smoke particles had already contaminated effectively the entire stock of finished chips in the clean room. 

The plant supplied radio-frequency chips to two customers of consequence: Nokia and Ericsson, which together accounted for roughly 40 percent of its output of those components. What followed is instructive precisely because the triggering event was identical for both. 

Nokia's purchasing organization noticed within days that order numbers were not adding up, before Philips had formally notified anyone of a problem. When the notification came, with an assurance that the plant would be back to normal within a week, the manager receiving it was not especially alarmed, but he escalated it anyway. "We encourage bad news to travel fast," his supervisor later told the Wall Street Journal. The five affected components were placed under daily monitoring rather than the usual weekly review. 

That instinct proved decisive. When Philips revised its estimate on 31 March, Nokia calculated it was facing a shortfall of just under four million handsets, more than five percent of its annual sales. Within hours it had assembled a cross-continental team. Two alternative suppliers in Japan and the United States absorbed millions of additional units of power amplifier chips on five days' notice. Capacity at Philips plants in Eindhoven and Shanghai was rerouted under direct pressure from Nokia's chief executive, who joined a delegation to Philips' Amsterdam headquarters. Certain chip designs were reworked so they could be manufactured elsewhere. 

Ericsson received news of the fire at the same time, through what its head of investor relations later characterized as one technician talking to another. The initial one-week estimate was accepted. The executive responsible for the mobile division did not learn of the problem until early April. By then the available capacity in the market had been contracted, and Ericsson had no qualified alternative for several key parts, having deliberately removed backup suppliers in the mid-1990s to simplify its supply lines. Many of the executives in post at the time of the fire were new to their roles and unaware of the exposure this had created. "We did not have a Plan B," the company's marketing director for consumer goods conceded to the WSJ. 

Ericsson estimated it lost at least 400 million dollars in potential revenue as a direct result, some of which an insurance claim was expected to recover. Nokia met its production targets and took market share, most of it from Ericsson. 

The wider collapse of Ericsson's handset business should not be laid at the door of the fire alone, and it is worth being precise about this. When the company announced in January 2001 that it would outsource all handset manufacturing to Flextronics, it attributed the SEK 16.2 billion (1.68 billion dollar) loss in its mobile phone division to a combination of component shortages, a poor product mix and marketing failures. The fire was one contributing factor among several, in a business already under strategic pressure. It merged its handset operations into a joint venture with Sony later that year. 

What the episode does demonstrate cleanly is narrower, and more useful. Two companies of comparable sophistication faced an identical external shock, and one absorbed it while the other did not. The difference was not the quality of their risk registers. It was the speed at which a weak signal reached someone with the authority to act, and the existence of alternatives that had been created before anyone needed them. 


Resilience is a capability, not a condition

Much of the discussion around resilience remains stuck at the level of documentation. Organizations commission a business continuity plan, file it, and treat the matter as settled. This confuses evidence of intent with operational capability. 

A more useful working definition is behavioral. A resilient organization is one that can absorb a disruption without losing control of its critical services, adapt its operating model while the disruption persists, and recover to a defined level of service within a period it has determined in advance rather than discovered in the moment. 

Three verbs, three distinct capabilities, and most organizations are meaningfully strong in none of them. 

It is also worth separating resilience from robustness. Robustness is the capacity to withstand a shock without changing. It is expensive, it degrades with novelty, and it fails abruptly when the shock exceeds design assumptions. Resilience assumes the shock will land, and concerns itself with what happens next. In an operating environment characterized by compound and correlated risk (climate volatility, cyber, energy, geopolitical fragmentation of supply), the second proposition is the sounder investment.


Four observable characteristics

In practice, resilient organizations look different from their peers in four specific ways. 

They have decided what is genuinely critical. Not a comprehensive process inventory, but a short and contested list: the handful of services whose interruption creates material financial, regulatory, or societal consequence, each with a maximum tolerable period of disruption expressed in hours or days and validated by executive management. The exercise is uncomfortable because it forces trade-offs that are usually left implicit. It is also where most of the value sits, because organizations routinely discover that their assumed point of failure is not the real one. 

They understand their dependencies, including those they do not control. The 2011 Tōhoku earthquake severed Toyota's supply chains so comprehensively that it took six months to restore production outside Japan to normal levels. In the aftermath, the company estimated that procurement of more than 1,200 parts and materials might be affected, and reduced that to a list of 500 priority items requiring secured future supply. This is the first characteristic applied in practice: not an exhaustive inventory, but a deliberate act of prioritization. 

Semiconductors were on the list. Toyota concluded that chip lead times were structurally incompatible with a severe shock, and built a business continuity plan requiring suppliers to hold between two and six months of inventory, calibrated to order-to-delivery time. A decade later, when the global shortage forced Volkswagen, General Motors, Ford, Honda and Stellantis to slow or suspend production, Toyota raised its output forecast instead. 

Two qualifications, both material to how this should be read. 

The first is that Toyota paid for it. The stockpiling arrangement is funded by returning a portion of the annual cost reductions the company demands from its suppliers. Resilience was purchased on a recurring basis, out of margin that would otherwise have been extracted. It is also not attributable to inventory alone: Reuters' sources placed comparable weight on Toyota's long-standing refusal to accept supplier "black boxes", having designed and manufactured its own microcontrollers for three decades, which gave it the technical fluency to find substitutes quickly. 

The second is that the protection was partial. By August 2021, with pandemic disruption in Asia compounding the shortage, Toyota cut planned global output for the following month by 40 percent, some 360,000 vehicles across fourteen plants. Resilience does not confer immunity. It buys time, and time is the scarcest resource in any crisis. 

They have pre-positioned decisions, not only plans. Who has the authority to halt production, and at what threshold? On the basis of what information, escalated by whom, and within what timeframe? A continuity plan that does not answer these questions transfers the hardest problem to the least favorable moment. The Nokia and Ericsson divergence was, at its core, a difference in escalation culture rather than in planning documentation. 

They rehearse, and they learn from incidents. A single crisis exercise with the executive committee in the room typically surfaces more genuine weakness than six months of documentary audit, largely because it tests the two things documents cannot: decision-making under ambiguity, and the informal dependencies nobody wrote down.


A Belgian case, and the right lesson to draw from it 

On 13 January 2020, the Picanol Group in Ypres was hit by large-scale ransomware. Overnight, colleagues in China reported being unable to log into several systems; the Belgian headquarters and the Romanian site followed. Because the group's production process is computer-managed end to end, manufacturing stopped across all three locations. Around 1,500 workers in Belgium were placed on temporary unemployment, with roughly a hundred continuing in the departments where production is not computerized. Trading in the company's shares was suspended the following morning. 

The recovery is the part worth studying. Management declined to pay the ransom and involved the federal police. Technical teams worked through the weekend, running tests and adding security modules as systems came back. Production restarted step by step from Monday 20 January, beginning with the foundry and its 150 staff, and full capacity was rebuilt over the following days. 

On 31 January the group published a regulated disclosure setting out the financial consequences. Lost production days would be recovered over subsequent weeks and months, and the impact was therefore confined largely to the fees of the external specialists brought in to rebuild the IT estate. Picanol put those costs at under one million euros, with no material effect on operating result. Trading resumed on 3 February, with the share about 1.5 percent below its pre-attack level. 

That outcome deserves to be read carefully, because it cuts against the reflex assumption that a week of stopped production is catastrophic. It was not, because the group could make the volume up. Manufacturing capacity that runs below its ceiling is, in effect, an unrecognized resilience asset, and organizations operating at full utilization do not have it. 

The costs that could not be recovered sit elsewhere, and they are the ones that tend to be omitted from the business case. Fifteen hundred people were idle for a week. The company was unvaluable on a public market for three weeks. And the first authoritative figure to reach the press, in the confusion of day two, was an internal estimate of tens of millions of euros, an order of magnitude above what the company itself would report a fortnight later. Crisis communication under uncertainty is not a peripheral discipline. It is part of the exposure. 

The diagnostic for any executive committee is the one Picanol was able to answer and most cannot: how long would it take us to resume production without our core systems, how much of the lost output could we recover afterwards, and who has verified those two answers? 


How the capability is built 

There is no shortcut, but there is a defensible sequence. 

It begins with governance. An accountable sponsor at executive level, an explicit mandate, and a budget. Absent this, the work reverts to a compliance exercise owned by a function without the authority to change anything material. 

It continues with mapping. Critical services, the dependencies beneath them (suppliers, sites, systems, key personnel, and increasingly the second and third tiers of the supply chain), and a set of plausible disruption scenarios drawn on an all-hazards basis rather than a single-risk view. 

It becomes real when exposure is quantified. As long as resilience is expressed qualitatively, it competes poorly for capital. Expressed as a euro figure per day of interruption, tested against the cost of the mitigation, it becomes an investment decision that a CFO can evaluate on familiar terms. 

It is then built and tested. Continuity, crisis management, internal and external communication, recovery sequencing. Tested deliberately, before circumstances test it involuntarily. 

And it is sustained through the management cycle. A continuity plan three years old, referencing a supplier base and an IT estate that no longer exist, is a work of fiction with a governance signature on it.


Why the timing is not neutral 

Under Directive (EU) 2022/2557 on the resilience of critical entities, Member States were required to have identified their critical entities by July 2026. For organizations in scope (energy, transport, water, health, digital infrastructure, food, banking and financial market infrastructure), multi-hazard risk assessment and a documented resilience plan are no longer discretionary good practice. They are obligations, sitting alongside NIS2, DORA, and the resilience disclosure requirements of CSRD. 

Regulation, however, is only the proximate reason to act. Organizations that approach CER as a documentation exercise will produce a file that satisfies an authority. Organizations that use it as the occasion to ask questions they have been deferring will produce something considerably more valuable, which is the capacity to keep operating when their peers cannot. 

In March 2000, that capacity was worth several points of global market share. 

 

At ngage, we support organizations across that full arc, from governance design and dependency mapping through to crisis exercises and embedding resilience in the management cycle. If you want a clear view of where you currently stand, that conversation usually takes an hour.