At 2:13 a.m., the SOC receives an alert involving a cloud administrator.
At 2:16, analysts lose access to the SIEM because authentication depends on the identity provider now under investigation.
At 2:19, the incident collaboration channel disappears.
At 2:23, the on-call engineer discovers that the break-glass credentials are stored in the privileged-access platform—which also authenticates through the suspect identity provider.
At 2:27, the cloud console still shows green status indicators, but nobody can establish whether the logs are complete, delayed, or manipulated.
By 2:35, the organization has two incidents.
The first is the security event.
The second is the loss of its ability to respond to the security event.
That second incident is the one most response plans do not adequately address.
Incident Response Has a Hidden Assumption
Most incident-response plans assume the defensive machinery will survive the incident.
The identity provider will authenticate responders. The privileged-access system will issue administrative credentials. The cloud management plane will accept containment commands. The SIEM will provide reliable telemetry. The ticketing platform will maintain the timeline. Email, chat, conferencing, and document-sharing systems will allow the response team to coordinate.
Those systems are treated as infrastructure surrounding the incident rather than as potential components of the incident.
That is a dangerous assumption.
In a recent State of Security post, I argued that controls that look separate on an architecture diagram may actually be “branches of the same tree.” They may share an identity provider, administrative tenant, logging pipeline, automation layer, DNS infrastructure, certificate authority, or cloud control plane. When one of those shared foundations fails, several supposedly independent controls can fail with it. (State of Security)
The next question is harder:
What is the smallest defensive system the organization must be able to operate after that tree falls?
This is not conventional disaster recovery. It is not simply a matter of restoring the SIEM, activating a backup tenant, or retrieving an emergency password.
It is degraded-mode security operations.
The Control Plane Can Fail by Lying
Security teams usually test failures as availability problems.
The identity provider is down. The SIEM is unreachable. The ticketing platform will not load. The primary cloud region is unavailable.
Those are relatively clean failures. They are visible, bounded, and easy to describe during a tabletop exercise.
Adversarial failure is different.
A compromised identity provider may continue issuing tokens. A hostile administrator may alter conditional-access policies while leaving the service operational. A logging pipeline may continue displaying data while silently omitting selected events. A collaboration platform may preserve most messages while exposing the incident channel to the attacker. An automation platform may continue executing containment playbooks after its credentials or logic have been subverted.
The system is not down.
It is lying.
That distinction changes the recovery problem. An unavailable system can sometimes be restored. An untrustworthy system must first be excluded from the response path.
NIST’s cyber-resiliency work provides the right foundation for thinking about this. It frames resiliency around four outcomes: anticipating adverse conditions, withstanding them, recovering from them, and adapting afterward. It also applies that thinking to shared services, common infrastructure, and systems of systems—not merely individual applications. (NIST Computer Security Resource Center)
A resilient security program therefore cannot define success only as preventing compromise or restoring systems. It must be able to continue essential defensive operations while some of its own systems are unavailable, degraded, or hostile.
Start With Inversion
The normal planning question is:
How do we keep the security control plane available?
That is necessary, but incomplete.
Invert the problem:
Assume the primary security control plane is unavailable or cannot be trusted. How do we still defend the organization?
Assume that:
-
The primary identity provider cannot be trusted.
-
Existing administrative sessions may belong to the attacker.
-
The privileged-access platform is inaccessible.
-
Centralized telemetry is incomplete.
-
Normal collaboration and ticketing systems are exposed.
-
Managed responder workstations may be under hostile administrative control.
-
Cloud automation may be executing unauthorized changes.
-
DNS, certificate, key-management, or time services may be unreliable.
-
One or more key responders cannot be reached.
Now ask what responders must still be able to do.
Not which products must be restored.
Not which dashboards executives expect to see.
Not which recovery checklist should be opened first.
What capabilities must exist for the organization to remain defensible?
Define the Minimum Viable Defensive System
A Minimum Viable Defensive System is the smallest deliberately independent collection of people, authority, identities, devices, communications, telemetry, containment mechanisms, evidence storage, and recovery materials that allows the organization to continue defending itself when its normal security control plane is unavailable or untrusted.
It is not a second full-sized SOC.
It is not a duplicate of every production security tool.
It is not a collection of individually labeled “break-glass” features.
It is a composed operating system for degraded defense.
At minimum, it must provide six capabilities.
| Capability | Minimum viable state | Evidence that it is independent |
|---|---|---|
| Authenticate | Authorized responders can establish identity and emergency authority using protected credentials and trusted administrative devices. | The path does not require the primary identity provider, PAM platform, corporate network, normal endpoint-management plane, or primary email account. |
| Communicate and coordinate | Responders can establish a secure command channel, assign roles, reach critical internal and external parties, and maintain a decision and action log. | Accounts, provider, authentication, network path, and recovery information do not depend on the primary collaboration environment. |
| Observe | Responders can obtain trustworthy information from critical systems and identify gaps, delay, or tampering in normal telemetry. | Collection, storage, access, keys, and administrative authority do not all share the suspected failure domain. |
| Contain | Responders can revoke access, disable identities, isolate assets, stop dangerous automation, alter routes or policies, and protect critical systems. | Emergency containment does not require normal SSO, normal PAM approval, the primary SOAR platform, or access to the normal ticketing workflow. |
| Preserve evidence | Responders can store relevant logs, configuration data, images, exports, notes, and decision records with integrity and provenance. | Primary administrators cannot silently alter or delete the evidence, and responders retain access to keys, timestamps, capacity, and custody procedures. |
| Reconstitute | Responders can rebuild roots of trust and restore defensive services from known-good materials in a defined order. | Configurations, credentials, keys, backups, tooling, and administrative access do not depend entirely on the environment being rebuilt. |
These are capabilities, not products.
A security team may satisfy them through a mixture of cloud-native controls, offline materials, separate accounts, alternative providers, clean devices, manual procedures, direct system access, and preauthorized decision rights.
NIST makes a similar architectural point in its resiliency guidance: survivability comes from combinations of technology, architectural choices, engineering practices, operational procedures, and people—not from one product or control.
Break Glass Is Not an Account
Organizations often point to an emergency administrative account as proof that degraded-mode access has been addressed.
That is only one link in the chain.
Consider an emergency account that:
-
Is stored in the normal PAM platform.
-
Uses the normal identity provider for authentication.
-
Requires a managed laptop that cannot be unlocked without normal SSO.
-
Can be used only from the corporate network.
-
Depends on the normal DNS and certificate infrastructure.
-
Requires an approval recorded in the normal ticketing system.
-
Sends its alerts to the normal SOC mailbox.
The account exists.
The emergency capability does not.
Break glass is not an account. It is an end-to-end operating path.
Microsoft’s emergency-access guidance recommends multiple cloud-only accounts, authentication methods that differ from normal administrative authentication, secure storage, monitoring, and recurring validation. AWS similarly describes pre-created emergency identities and roles, dedicated emergency access arrangements, hardware authentication, and periodic testing for situations in which the centralized identity provider is unavailable or compromised. (Microsoft Learn)
Those are valuable design ingredients. The security program still has to compose them into a functioning response system.
The responder must be able to retrieve the credential, authenticate, use a trusted device, reach the management interface, execute an authorized action, record that action, observe its effect, and preserve the resulting evidence.
Testing only the login proves only the login.
Independence Is a Property of the Whole Path
Two tools can be different and still fail together.
A secondary messaging platform is not independent if it uses the same identity provider.
A backup SIEM is not independent if it receives data through the same collectors.
A separate cloud account is not independent if the compromised organization administrator can assume control over it.
An immutable log store is not useful during the incident if access to its decryption keys depends on the failed key-management plane.
A responder laptop is not clean if it must contact the suspect device-management service before allowing an administrator to sign in.
An alternate network path is not alternate if it converges on the same DNS, certificate, firewall-management, or telecommunications dependency.
This is why fallback design must be scenario-specific. Nothing is universally independent. It is independent only relative to a defined failure.
NIST’s guidance on diversity and redundancy warns that apparent alternatives can converge on the same underlying foundation. It specifically treats diversity of command, control, and communications paths—including out-of-band paths—as a resiliency technique, and notes that redundancy can be undermined when duplicated capabilities share common resources or homogeneous components.
CISA guidance has likewise emphasized out-of-band incident communications, separate management paths, and logging aggregated into protected out-of-band locations. These practices reduce the chance that the compromised environment can blind or isolate the response team. (CISA)
Again, those mechanisms are necessary.
They are not sufficient until they work together.
Restore Capabilities, Not Products
During a compound failure, organizations tend to restore whatever system has the most visible outage, the loudest executive sponsor, or the clearest recovery runbook.
That can produce the wrong order.
The goal is not to make the security stack look normal as quickly as possible. The goal is to restore trustworthy defensive agency.
A practical priority sequence looks like this.
1. Establish trusted command
The organization needs an incident commander, an emergency authority model, authenticated responders, a protected communications channel, and a functioning decision record.
Without trusted command, every subsequent action is debatable. Responders do not know who is authorized, which instructions are legitimate, or whether the attacker is participating in the response.
This layer should also provide access to an offline or separately protected contact roster covering executives, legal counsel, insurers, outside responders, critical vendors, law enforcement contacts, and communications personnel.
2. Establish a trustworthy view
Responders must determine what they can still observe and which sources remain credible.
This does not require immediately rebuilding the entire SIEM. It may involve direct access to native audit sources, an independently protected log archive, network telemetry, cloud snapshots, system exports, or read-only queries through emergency accounts.
The first question is not, “Are logs arriving?”
It is, “What evidence do we have that these logs are complete, current, and resistant to alteration by the suspected adversary?”
Evidence preservation begins here. Observation and preservation should not be separated into distant phases. The data available during the first hour may not remain available later.
3. Regain safe containment capability
Once responders have enough confidence to act, they need a limited but reliable way to constrain the event.
That may include:
-
Revoking sessions and tokens.
-
Disabling identities.
-
Isolating accounts, workloads, subscriptions, or network segments.
-
Blocking known infrastructure.
-
Suspending compromised automation.
-
Protecting backups and logging repositories.
-
Restricting administrative paths.
-
Moving critical services into a predefined defensive posture.
Containment paths should be narrow, preauthorized, observable, and reversible where practical.
They should not depend on the same orchestration layer whose trust is in question.
4. Reconstitute roots of trust
Only after responders have trusted command, sufficient visibility, and safe agency should they begin rebuilding the normal control plane.
Reconstitution must follow dependency order.
Identity may need to be restored before PAM. Key management may need to be restored before protected logging. Trusted administrative endpoints may need to be rebuilt before cloud policy is changed. Logging may need to be reestablished before production workloads are reconnected.
Restoring dependent tools before their foundations are trustworthy creates the appearance of recovery without its substance.
This is a priority model, not a waterfall. Observation, evidence preservation, and containment will often proceed in parallel. The point is to prevent teams from restoring familiar products while critical defensive capabilities remain absent.
Measure Time to Defensible State
Traditional recovery planning uses recovery-time objectives to express how long a system can remain in recovery before unacceptable harm occurs. That is useful for determining when an identity platform, logging service, or collaboration system must return. (NIST Computer Security Resource Center)
It does not answer what defenders can do while those systems remain unavailable.
A security program should add another measure:
Time to Defensible State
Time to Defensible State is the elapsed time between declaring the primary defensive control plane unavailable or untrustworthy and validating that the Minimum Viable Defensive System is operating through independent paths.
A defensible state might require proof that:
-
Emergency authority has been invoked.
-
Required responders have authenticated independently.
-
A secure command channel is operating.
-
A decision and action record is being maintained.
-
At least one trustworthy telemetry path is available.
-
Responders can execute and verify an emergency containment action.
-
Evidence can be deposited into an independently protected repository.
-
The team understands which normal services remain prohibited.
“Someone successfully logged in” is not a defensible state.
“We opened the backup chat room” is not a defensible state.
The state is reached only when the capabilities operate together.
Organizations should measure the component times as well:
-
Time to recognize and declare control-plane degradation.
-
Time to establish trusted responder identity.
-
Time to establish out-of-band command.
-
Time to obtain the first trustworthy telemetry.
-
Time to execute the first validated containment action.
-
Time to preserve the first evidentiary artifact.
-
Time to begin controlled reconstitution.
-
Maximum duration the degraded system can operate.
That last measure matters.
A fallback that works for 20 minutes but cannot support a twelve-hour investigation is not sufficient. Capacity, credential lifetime, battery life, communications access, evidence-storage volume, staffing, vendor support, and shift turnover are all part of the design.
Test Compound Failure, Not Components
Most organizations test their emergency mechanisms one at a time.
The emergency administrator logs in.
The backup conference bridge works.
The log archive accepts a test event.
The incident-response binder opens.
The cloud backup restores.
Every component passes.
Then the system fails during the exercise because nobody can retrieve the emergency credential without the PAM system, the clean laptop requires the unavailable identity provider, and the alternate communications channel does not include legal counsel or the cloud team.
Component success is not system success.
A meaningful exercise should begin with a compound condition such as:
The identity provider is suspected of compromise. Existing administrative sessions cannot be trusted. The centralized logging query plane is unavailable, although some collection continues. Normal email, chat, ticketing, and PAM systems are prohibited. The cloud management environment may contain unauthorized policy changes.
Then make the team perform the work.
Retrieve the emergency materials using the real custody process.
Start the designated administrative devices.
Use the actual alternate network path.
Authenticate to the necessary systems.
Generate a known event and prove it appears in the independent telemetry path.
Contain a sacrificial asset.
Preserve an artifact and validate its integrity.
Reach an external participant.
Maintain the incident log.
Operate long enough to perform a shift handoff.
Exit degraded mode, reconcile all emergency actions, rotate credentials, and restore monitoring.
Do not allow the exercise controller to hand the team imaginary access, pretend that a credential was retrieved, or declare a containment action successful without executing a safe equivalent.
Every simulated shortcut conceals a dependency.
NIST’s current incident-response guidance treats preparation, response, and recovery as integrated parts of cybersecurity risk management rather than as a separate binder activated after detection. Compound-failure testing is one way to turn that integration into operational evidence. (NIST Computer Security Resource Center)
The Fallback Creates Its Own Risk
An independent defensive system is powerful.
That also makes it dangerous.
Emergency identities may hold standing privilege. Alternative communications channels may escape normal monitoring. Offline credentials can become stale or be mishandled. Direct administrative paths can bypass approval systems. Separate evidence stores may contain highly sensitive information. Dormant devices may miss critical patches.
The fallback cannot simply be hidden and forgotten.
It requires its own controls:
-
Dual custody for the most powerful emergency materials.
-
Tamper-evident storage and access records.
-
Narrowly scoped emergency roles.
-
Independent monitoring of every activation.
-
Regular credential, key, and device validation.
-
Immediate review and rotation after use.
-
Defined activation and termination authority.
-
Reconciliation of emergency actions into normal records.
-
Periodic review for new shared dependencies.
This is a second-order tradeoff.
The organization creates exceptional access to survive the loss of normal access controls. It must then protect that exceptional access without reconnecting it to the same control plane it was designed to bypass.
There is no product setting that resolves that tension. It must be engineered and governed.
A Practical Design Standard
A security program should not claim degraded-mode readiness until it can produce six things.
1. A declared defensive minimum
The organization has explicitly defined what authentication, communication, observation, containment, evidence preservation, and reconstitution mean in its environment.
2. A dependency topology
Each capability has been traced through its identities, devices, networks, providers, administrators, keys, data sources, and human custodians.
Shared failure domains are visible.
3. Independence claims tied to scenarios
The organization does not say that a system is simply “independent.” It states which failures it is independent from and provides evidence supporting that claim.
4. Activation and exit criteria
Responders know who can declare the primary control plane untrusted, what changes when that declaration occurs, and what evidence is required before normal systems can be used again.
5. Measured recovery of defensive capability
The organization has established a target Time to Defensible State and measured it during an end-to-end exercise.
6. Evidence of sustained operation
The fallback has demonstrated that it can support a realistic incident duration, including staffing changes, evidence growth, credential use, containment actions, external coordination, and eventual reconciliation.
This turns degraded-mode response from a collection of reassuring statements into a testable design standard.
Small. Separate. Boring. Measured.
The Minimum Viable Defensive System should be small enough to understand.
It should be separate enough to survive the named failure.
It should be boring enough to operate under stress.
It should be measured often enough that leadership knows whether it is real.
Centralization is not inherently bad. Modern security programs need centralized identity, logging, administration, orchestration, and collaboration.
But centralized efficiency creates concentrated dependency.
A mature program does not pretend those dependencies will always survive. It plans for the moment when the normal tools are unavailable, the dashboards are questionable, and the attacker may be using the same administrative machinery as the defenders.
The question is no longer:
Do we have break-glass accounts?
The question is:
Can we prove that authorized responders can authenticate, communicate, observe, contain, preserve evidence, and begin reconstitution when the primary security control plane is unavailable or hostile—and can they do it within a measured period?
Until the answer is yes, the organization is resilient only while its assumptions hold.
Security programs must do more than defend the business.
They must remain capable of defending it while their own machinery is failing.
More Information and Assistance
MicroSolved, Inc. can help organizations:
-
Map security-control dependencies and shared failure domains.
-
Define a Minimum Viable Defensive System.
-
Design independent emergency identity, communications, administrative, logging, and evidence paths.
-
Establish degraded-mode activation and recovery procedures.
-
Run compound-failure tabletop and live validation exercises.
-
Define and measure Time to Defensible State.
Contact MicroSolved at info@microsolved.com or +1.614.351.1237. Relax. We’re on watch.
* AI tools were used as a research assistant for this content, but human moderation and writing are also included. The included images are AI-generated.
