micara
Embedded AI Subsurface Compliance Advisory Blog Contact
Blog · Compliance

Critical Infrastructure Is Not a Label. It Is an Operating Discipline.

1 September 2026 · 5 min read · -
-

When does a company cease to be merely important and become critical?

The obvious answer is: when the law says so. The useful answer is harder. A company becomes critical when the failure of a service it provides can propagate beyond its own balance sheet - into hospitals, homes, payment systems, transport networks, water supplies or public order. At that point, resilience is no longer only a private matter. It becomes part of the operating contract between business and society.

That is the central challenge of the German framework for critical infrastructure. Its purpose is not to award a badge. It is to establish, with defensible evidence, whether a critical service can continue, whether interruptions can be contained and whether recovery can be achieved under real conditions. The German BSI Act (BSIG) places cyber-risk management, incident reporting, management duties and evidence obligations in a single statutory architecture; the KRITIS-Dachgesetz adds the physical-resilience dimension.

This matters far beyond the audit team. It changes how boards allocate accountability, how operations describe dependencies, how procurement treats suppliers, how IT and OT teams preserve evidence, and how crisis managers rehearse the first 24 hours of an incident. The following ten insights translate these obligations into an executive operating agenda.

This article is not legal advice. Sector-specific law, implementing acts and current authority guidance must be checked for each organisation.

The ten insights at a glance

Classify precisely - the labels are not interchangeable. A NIS2-regulated entity is not automatically a KRITIS operator. A critical-facility operator is, however, a particularly important entity under the BSIG. Sector, size, facility category, service and threshold all matter.

Start with the critical service, not the server list. The protected object is continuity of supply to society. Technology matters because it enables the service, not because every asset is equally critical.

Cyber and physical resilience now form one operating problem. The BSIG and KRITIS-Dachgesetz address different dimensions, but management must connect cyber defence, physical protection, continuity and recovery.

Executive accountability is designed into the regime. Management must implement and oversee risk measures and maintain its own recurring competence development. This is a governance duty, not a technical delegation.

Scope is the first control - and the first point of failure. A defensible scope links critical services to processes, IT, OT, sites, interfaces and third parties. An incomplete scope makes every later assurance statement unreliable.

Risk has to be measured in service consequences. Commercial loss alone is not the test. The analysis must show how availability, integrity and other failures affect minimum supply, the public and interconnected sectors.

Designed, implemented and effective are three different claims. Policies may demonstrate intent. A test of one may show point-in-time implementation. Only evidence across a period supports an effectiveness conclusion.

Evidence is part of the control system. Logs, tickets, inventories, approvals, exercise records and remediation trails should be generated by normal operations. Reconstructing them before an audit is costly and fragile.

The incident clock is a rehearsed business process. The 24-hour early warning, 72-hour incident notification and one-month final report require prepared thresholds, roles, communications and evidence before a crisis.

The three-year audit cycle should become a continuous assurance cycle. Risk-based sampling, rotating coverage, recurring tests and disciplined defect closure turn compliance into a multi-year resilience programme.

1. Classification must be precise before compliance can be credible

The first trap is linguistic. NIS2-regulated companies, important entities, particularly important entities and operators of critical facilities are frequently discussed as though they were the same population. They are not. Under Section 28 BSIG, critical-facility operators are particularly important entities. But many particularly important or important entities qualify because of sector and enterprise size, even though no individual facility reaches a KRITIS threshold. The BSI's affectedness materials are therefore a useful starting point, not a substitute for a documented legal and operational analysis.

For most businesses, classification requires two different tests. The entity test asks what the organisation does, in which statutory sector, and whether the relevant employee, turnover and balance-sheet criteria are met. The critical-facility test asks a different set of questions: Which service is critical to the public? Which facility delivers it? Which facility category applies? Does the applicable supply threshold or other designation criterion make that facility significant? Group structures add another complication because partner and affiliated enterprises can influence the size calculation, subject to the statutory independence rules.

Affectedness is not a one-off memo. Supply values and organisational facts change. Plants are commissioned, networks are consolidated, services are outsourced, and thresholds can be crossed without a strategic announcement. A defensible organisation therefore assigns ownership for an annual threshold review and for event-driven reclassification after acquisitions, reorganisations, capacity changes or changes in service design. Registration follows the classification; it does not create it. Under the relevant statutes, registration is generally due no later than three months after the organisation or facility first qualifies.

Board question: Can we show, in one signed paper, which statutory category applies to every relevant legal entity and facility, who performed the analysis, what data were used and when the conclusion must be reviewed?

2. The unit of analysis is the critical service

Traditional security programmes often begin with assets: servers, firewalls, applications, buildings. The KRITIS logic begins one level higher. It begins with the service whose failure would cause serious supply shortages, threaten public safety or produce comparable societal consequences. This shift sounds semantic, but it changes the entire control architecture.

Once the service is defined, the organisation must state the quality and quantity that must be maintained in normal and abnormal conditions. A drinking-water operator, for example, cannot describe availability only as a percentage of server uptime. It must connect technical availability to potable-water quality, volume, pressure, coverage and the duration for which minimum supply can be sustained. A transport operator must connect systems to safe movement capacity. A data-centre service must be examined through the customer services and dependencies it enables. Protection objectives should be derived from the required level of supply.

This service-first logic prevents a common error: giving equal attention to every asset because each appears somewhere in an inventory. The relevant question is not whether an asset is valuable in general. It is whether, and how quickly, its compromise could affect the critical service. Business impact analysis, maximum tolerable outage, dependency mapping and restoration priorities should all point to the same service outcome.

Operational test: For every high-value IT or OT asset, complete the sentence: "If this fails, the critical service will be affected after ___ minutes or hours because ___. The minimum acceptable service level is ___." If the answer is unavailable, prioritisation is not yet defensible.

3. Cybersecurity and physical resilience are now one management problem

The BSIG and KRITIS-Dachgesetz address two closely connected dimensions of resilience. The first establishes the cyber and information-security regime for covered entities and critical facilities; the second establishes a broader physical-resilience regime. The two laws retain different authorities, instruments and sector-specific interactions, but an operator cannot manage them as unrelated compliance islands.

A malicious actor does not respect the boundary between cyber and physical controls. A compromised remote-maintenance account can disable a physical process. A fire can remove both production capacity and the redundant network path routed through the same building. A power failure can exhaust battery autonomy, stop pumps and disable communications. A personnel shortage can render a theoretically available manual fallback useless. The KRITIS-Dachgesetz expressly calls for prevention, physical protection, incident response, impact limitation and rapid restoration, supported by a resilience plan; operator risk analysis must also consider dependencies on other sectors and services.

The result should be one integrated view of failure, even if legal evidence is submitted through different channels. Cybersecurity, physical security, business continuity, emergency power, crisis communications, staffing, suppliers and restoration logistics must be connected through common scenarios. The organisational structure may remain federated. The risk story cannot.

Management implication: Avoid two disconnected registers labelled "cyber risks" and "physical risks". Use shared service-impact scenarios, common dependency data and a single executive view of residual risk.

4. Executive accountability cannot be delegated to the CISO

Section 38 BSIG gives management bodies a direct implementation and oversight duty for required cyber-risk measures and obliges them to update their relevant knowledge regularly. The KRITIS-Dachgesetz imposes corresponding implementation and oversight expectations for resilience measures. The point is not that directors must become technical auditors. It is that they must be able to recognise risk, interrogate the quality of risk-management practices and understand consequences for the services their organisation provides.

This changes the board agenda. A quarterly status report stating that the programme is "green" is insufficient if it does not reveal the underlying service risks, overdue high-risk findings, untested fallbacks, supplier concentration or the age of restoration evidence. Management needs a line of sight from statutory obligation to risk scenario, control owner, operating evidence, exception and remediation date.

Perfection is not a credible audit objective. Findings are expected in a living security system. The governance failure is not the mere existence of a finding; it is the inability to identify, classify, own, fund and close it. Mature management asks whether high-risk measures are operating, whether repeat findings persist and whether the organisation is learning faster than its risk changes.

Evidence for oversight: Retain agendas, decisions, challenge questions, risk acceptances, funding decisions, competence-development records and follow-up actions. Board attention should be demonstrable, not inferred.

5. Scope is the first control - and the first audit failure

Scope validation is foundational. If the scope omits a relevant process, system, component, site, network segment or provider, every later conclusion can be technically correct and still materially misleading. The audit becomes a moving target, and assurance is issued over the wrong object.

A robust KRITIS scope is not a rectangle drawn around the production site. It is a traceable chain: critical service to business and operational processes; processes to information, applications, databases and automation; systems to infrastructure, identity, monitoring and communications; and internal components to cloud services, managed providers, maintenance partners and upstream utilities. Shared corporate services require particular care. They may sit outside the plant organisation while remaining essential to plant operation.

The broad scope applying to important and particularly important entities must also be distinguished from the more focused facility-specific scope used for the enhanced KRITIS obligations and attack-detection requirements. Companies subject to multiple regimes should therefore use a layered scope model rather than forcing everything into one undifferentiated inventory. The scope document should be both textual and graphical, version-controlled and tied to named services and registered facilities. The current BSI evidence guidance and forms reinforce this expectation.

A practical scope chain: Service -> minimum service level -> critical processes -> IT/OT and physical assets -> interfaces -> supporting services -> third parties -> evidence owner. Every arrow needs a reason and an accountable owner.

6. A KRITIS risk analysis is incomplete until it explains public-service consequences

Many enterprise risk registers are financially literate and operationally thin. They score revenue loss, repair cost and reputational damage, but say little about the number of people affected, the duration of interrupted supply, degradation in service quality or cascading effects into other sectors. That is inadequate for a critical operator.

A credible assessment requires an all-hazards approach: natural events, technical failure, deficient services, hostile attacks, human error and organisational weakness. The current BSIG requires a cross-hazard approach for risk-management measures, while the KRITIS-Dachgesetz requires operator risk analysis at least every four years and expressly includes interdependencies within Germany, across neighbouring EU states and with third countries.

The hard part is not producing a longer threat catalogue. It is connecting threats to service outcomes and then selecting treatments that preserve delivery. Insurance may transfer financial loss, but it does not keep water flowing. Risk acceptance does not preserve a service unless credible compensating arrangements exist. A manual fallback does not reduce risk if the necessary staff, access, instructions, communications or spare parts will not be available during the assumed event. Every mitigation claim needs an operating mechanism.

The most useful risk statements are therefore causal. They explain the hazard, vulnerable dependency, loss of a protection objective, time to service impact, affected population or customer group, existing safeguards, residual exposure and restoration route. This makes proportionality visible: investment can be weighed against the plausible consequences of failure rather than against an abstract maturity score.

Risk-quality test: Reject treatments that only move money. Ask: What maintains the service? What limits the outage? What restores the service? Who is available to do it under degraded conditions?

7. Adequacy, implementation and effectiveness are three separate claims

A policy can be elegantly written and operationally irrelevant. A robust assurance process therefore distinguishes four phases: foundations, adequacy, effectiveness and result. First, it validates the audit basis, scope and risk management. Second, it tests whether the process design is suitable and aligned with the state of the art. Third, it tests whether measures actually operated over time. Only then can a defensible result be formed.

This distinction should shape internal assurance. A walkthrough and document review can establish whether appropriate activities exist. A point-in-time test can show whether one control instance was implemented on a given date. It cannot, by itself, establish sustained operation. Effectiveness requires a reliable population, representative samples, historical evidence, exception analysis and a reasoned inference from the sample to the population.

The BSI's RUN model makes the progression explicit: planning, controlled operation, establishment, measurement and improvement. It also resists the false comfort of partial maturity. A newly introduced management system may have sound processes on paper, but it cannot credibly claim high maturity until those processes have run repeatedly, produced measurements and driven improvements.

For management, this produces a clear rule: do not report that a risk is controlled merely because a policy has been approved or a tool has been purchased. Report separately on design, deployment coverage, operating effectiveness and improvement. Each statement should have a different evidentiary basis.

Three levels of proof: Design means the control could address the risk. Implementation means the control exists where required. Effectiveness means the control operated consistently over a relevant period and produced the intended risk-reduction outcome.

8. Evidence is not paperwork added after security; it is part of security

Audit evidence must make the work performed and the conclusions reached reproducible. Scope, objectives, time, place, team, tested objects, methods, population and sample logic all matter. Findings must be classified, linked to the critical service and accompanied by a concrete remediation plan with owner, date and status.

This is why evidence architecture should be designed into daily operations. Access approvals should exist because the workflow records them. Patch exceptions should have an owner and expiry because the exception process requires it. Backup restoration should generate test records because restoration is actually rehearsed. Supplier reviews, crisis exercises, competence records, alarm handling and corrective actions should leave reliable trails. Evidence is strongest when it is an unavoidable by-product of performing the control correctly.

Existing certifications can help, but they do not erase the KRITIS delta. ISO 27001 certification may support conclusions about management-system design and appropriateness, provided the scope and measures align. It may not, on its own, prove the operating effectiveness expected for every KRITIS-relevant measure. The BSI's Section 39 guidance similarly requires operators and auditors to reconcile scope, risk criteria, timing, effectiveness and missing areas.

The practical consequence is a controlled evidence repository, not a frantic pre-audit data room. Artifacts need owners, retention rules, traceability to controls and protection appropriate to their sensitivity. Evidence should also survive personnel changes and vendor transitions. If the organisation cannot reproduce how a conclusion was reached, the conclusion is weak even when the underlying technology is good.

Evidence principle: The best audit evidence is generated automatically or procedurally by normal control operation, retained in a protected source system and linked to the population from which testing samples are drawn.

9. Incident reporting is a clock that must be rehearsed before it starts

Under Section 32 BSIG, the reporting sequence for a significant security incident begins with an early warning within 24 hours of awareness, followed by a fuller notification within 72 hours and a final report generally within one month. For critical-facility operators, the notification must also describe the affected facility, the critical service and the actual or potential service impact. The BSI's current incident-reporting materials explain the same staged logic.

Those deadlines are not primarily a form-filling challenge. They are a decision-rights challenge. In the first hours, facts will be uncertain, counsel may be involved, operations may be degraded, vendors may control important logs and senior leaders may be unavailable. The organisation must still decide whether the threshold is met, preserve evidence, distinguish known facts from assumptions, coordinate cyber and physical information and communicate through an available channel.

A workable reporting process therefore needs pre-approved escalation criteria, a 24/7 contact model, deputies, authority to notify, secure communications, current facility and service data, and templates that separate the initial signal from later analysis. Tabletop exercises should include weekends, loss of normal email, incomplete telemetry and cross-border impact. The clock should be part of the incident drill, not explained for the first time during the incident.

First-hour checklist: Establish awareness time; protect life and service; activate decision-makers; preserve evidence; identify affected facility and service; assess reporting threshold; open the 24/72/one-month timeline; maintain a decision log.

10. Treat the three-year audit as a continuous assurance cycle

Section 39 BSIG requires operators of critical facilities to demonstrate the relevant measures through audits, examinations or certifications at the time set by the BSI - for newly qualifying operators no earlier than three years after first qualification and, after requalification, no later than three years - and then every three years. That does not make preparation once every three years a defensible operating model.

A complete examination of a large scope is rarely economical. The practical answer is risk-based sampling and, where defensible, a multi-year audit concept. High-risk areas should receive recurring attention; lower-risk measures can be rotated for effectiveness testing; critical processes must remain covered; and sample selection should vary across cycles. The population must be complete, drawn from a reliable source and tied to the review period. Sampling is not a shortcut around risk. It is a disciplined way to make a reasoned statement about a system too large to inspect item by item.

The strongest operators therefore build a rolling assurance calendar that combines internal audit, technical testing, continuity exercises, supplier reviews, facility inspections, detection testing and remediation validation. Previous findings are carried forward until closure is evidenced. Changes in services, topology, facilities and suppliers trigger scope and risk updates. The BSI's GAiN and evidence FAQ make clear that incomplete or implausible submissions can lead to document requests, revised scope or remediation plans and further supervisory attention.

KRITIS readiness is not the ability to survive an audit week. It is the ability to explain and demonstrate, on an ordinary day, why the public service will remain available on an extraordinary one.

The strategic shift: Move from "audit readiness" to "continuous service assurance". The first optimises a submission. The second improves the probability that the service survives.

A practical ninety-day agenda

The first management steps can be made concrete. The objective of the next ninety days should not be to create a perfect compliance programme. It should be to remove uncertainty from classification, service scope, accountability and evidence.

Period

Focus

Minimum output

Days 1-30

Classification and ownership

Confirm entity and facility status; identify sector-specific overlaps; appoint executive and operational owners; establish review triggers and the registration calendar.

Days 31-60

Service, scope and risk

Define minimum service levels; map critical processes, IT/OT, sites and third parties; reconcile inventories; rewrite top risks in service-impact terms.

Days 61-90

Proof and response

Identify evidence gaps; validate incident reporting; test one high-impact scenario; create a rolling assurance plan; fund the highest-risk remediation actions.

Conclusion: the audit is a mirror, not the destination

The clearest lessons belong in the boardroom. Classification must be demonstrable. The critical service must anchor scope and risk. Cyber and physical resilience must meet in common scenarios. Controls must be shown to work over time. Evidence must be generated by operations. Incidents must be reported through a rehearsed process. Defects must become owned improvements.

None of this guarantees that a determined attack, extreme event or compound failure will be harmless. That is not the standard. The standard is disciplined proportionality: understand what society depends on, identify how it can fail, implement measures appropriate to the consequences, test whether they work and improve them when evidence says they do not.

The serious operator does not ask, "Can we pass the KRITIS audit?"

It asks, "What would have to be true for our critical service to survive - and can we prove that those things are true today?"

Download PDF: schulungsfolien-kritis-auditor (19.0 MB)