Three Field Guides for the Security Problems Showing Up Right Now

Over the last several months, we have released three longer-form guides around problems we keep seeing in real environments:

  • The Antifragile SOC Book

  • The AI Agents Management Framework

  • The Security Leader’s Operating System

They cover different areas, but they share a common theme: how to operate effectively as complexity, automation, and pressure increase.

Books

Why We Created These Materials

These guides did not begin as content projects.

They came from working problems.

Over the years, we have accumulated a lot of models, processes, questions, and operating approaches that we use when helping organizations solve difficult security problems. Some came from client work, some from research, and plenty came from lessons learned the hard way.

Much of that knowledge lived in conversations, notes, presentations, and experience.

We decided it was time to write more of it down.

Not as theory, and not as another collection of cybersecurity predictions.

As practical field material people can use, challenge, adapt, and apply.

Why Release Them Now?

The timing matters.

Security teams are being asked to handle more data, more tooling, more organizational expectations, and more complexity.

At the same time, AI is moving rapidly from experimentation into real operational workflows.

And security leaders are increasingly being asked to manage all of this while making faster decisions with broader consequences.

The answer cannot simply be to do more.

More alerts. More automation. More dashboards. More meetings. More processes.

What we increasingly need are better operating models.

That is what these three guides are intended to provide.

The Antifragile SOC Book

Most SOCs are very good at generating activity.

The harder question is whether that activity is actually improving security.

The Antifragile SOC Book looks at how security operations can move beyond simply absorbing alerts and incidents toward continuously learning from them.

It is about building a SOC that gets better under pressure, focuses scarce analyst attention where it matters, and treats noise, friction, and failure as signals for improvement.

If you run, manage, or depend on a SOC, this one is for you.

Download The Antifragile SOC Book:

https://media.microsolved.com/The_Antifragile_SOC_Book.pdf

The AI Agents Management Framework

AI agents are quickly moving from interesting experiments into systems that can retrieve data, use tools, trigger workflows, and take action.

That creates a new management problem.

The AI Agents Management Framework is built around a simple idea: once an AI agent can act on behalf of the organization, it should be governed more like a digital worker than a traditional application.

The guide provides a practical way to think about ownership, authority, access, monitoring, risk, and accountability as agents become more capable.

If your organization is deploying agents beyond simple chat interfaces, now is the time to think about how they will be managed.

Get the AI Agents Management Framework:

https://signup.microsolved.com/ai-management-e-book/

The Security Leader’s Operating System

The third guide is the most personal.

Capable security leaders tend to attract work.

Eventually, incidents, exceptions, decisions, meetings, escalations, and unfinished problems can consume all of the time that was supposed to be used for leadership.

The Security Leader’s Operating System captures the approach I use to decide what should disappear, what should move to someone else, what should be simplified, what should be automated, and what genuinely deserves a leader’s attention.

It is not another productivity system for processing an infinite queue faster.

It is about redesigning the queue.

If you are a security leader who has become the default destination for every difficult problem, this one may be the most useful of the three.

Download The Security Leader’s Operating System:

https://media.microsolved.com/The_Security_Leaders_Operating_System.pdf

The Common Thread

I did not originally think of these as a series.

But they keep converging on the same questions.

Where should human attention go?

What should we automate?

What should remain under human judgment?

How do we distinguish useful outcomes from simple activity?

And how do we build systems that improve instead of merely accumulating more work?

Those questions are becoming more important as AI and automation make it easier to do more things, faster.

The challenge is no longer just whether something can be done.

It is deciding what is worth doing.

That is why we are releasing these materials now.

If one of these problems is showing up in your organization, grab the relevant guide, read it, challenge it, and put something from it into practice.

The Antifragile SOC Book
https://media.microsolved.com/The_Antifragile_SOC_Book.pdf

The AI Agents Management Framework
https://signup.microsolved.com/ai-management-e-book/

The Security Leader’s Operating System
https://media.microsolved.com/The_Security_Leaders_Operating_System.pdf

Use what works.

Discard what does not.

And, ideally, make the system better.

 

* AI tools were used as a research assistant for this content, but human moderation and writing are also included. The included images are AI-generated.

Introducing The Security Leader’s Operating System

There is one sentence I hear from information security leaders more than almost any other:

“I don’t have time to do what I need to do.”

The problem usually isn’t motivation, discipline, or effort. Security leaders already work hard. The deeper problem is that their ability to handle difficult work attracts more difficult work.

Questions, approvals, exceptions, vendor reviews, audit requests, incidents, meetings, unfinished decisions, and executive concerns all flow toward the person who has demonstrated that they can carry them. Eventually, being able to handle almost anything becomes the reason the leader has time to think about almost nothing.

I wrote The Security Leader’s Operating System to help change that.

Thinking

Security Leaders Don’t Need Another Productivity Trick

Most productivity systems concentrate on organizing work after it has already been accepted.

This field guide starts earlier.

Before deciding where a task belongs, it asks whether that work should exist in its current form at all. It applies the EDSAM mental model:

  • Eliminate work that does not need to happen.
  • Delegate outcomes that do not require the leader’s authority or judgment.
  • Simplify work until it produces the smallest useful result.
  • Automate the predictable parts of the work that remains.
  • Maintain only what must remain under the leader’s direct care.

The objective is not to help security leaders process an endlessly growing queue more quickly. It is to redesign the system so that less work requires the leader in the first place.

Protect the Work Only the Leader Can Do

Some work genuinely belongs with the security leader.

Setting risk appetite, challenging assumptions, translating technical exposure into business decisions, communicating during consequential events, developing people, and accepting residual risk cannot simply be optimized away.

But that work is often crowded out by management exhaust: status collection, repetitive reporting, preventable escalations, meeting preparation, evidence retrieval, routing questions, and processes that exist largely because they have always existed.

The guide introduces practical ways to distinguish leader-only work from work that can be removed, transferred, redesigned, or supported by better systems.

It includes:

  • A Money, Attention, and Time workload scan
  • The Eisenhower Matrix mapped into EDSAM
  • TaskGrid as an example of integrating urgency, importance, agency, commitments, planning horizons, and EDSAM dispositions
  • Information security workload playbooks
  • A recurring operating cadence
  • A 30-day reset
  • Capacity and decision tools
  • Field-ready worksheets and checklists
  • A NIST Cybersecurity Framework 2.0 scope check

AI Can Reduce Work—or Industrialize the Wrong Work

Modern LLMs, GPTs, skills, assistants, and agents make this conversation more urgent.

These tools can retrieve information, summarize evidence, compare documents, prepare decision packets, draft communications, and move bounded work through repeatable processes.

They can also accelerate bad processes, obscure missing context, accumulate excessive authority, and create new infrastructure that someone must maintain.

That is why the guide applies EDSAM before automation.

It explores how to divide work among humans, models, and deterministic systems; how to establish capability levels and autonomy budgets; and how to preserve accountable human judgment when AI-enabled workflows can create real consequences.

The important question is not simply, “Can AI perform this task?”

The better sequence is:

  1. Should this work exist?
  2. Who should own its outcome?
  3. What is the smallest useful result?
  4. Which parts are predictable enough to automate?
  5. What judgment, authority, or relationship must remain human?

A Field Guide, Not a Theory Book

This is intended to be used with the work already in front of you.

Choose one recurring meeting. One report. One approval queue. One source of preventable escalation. One hour of work that only you can perform.

Then apply the system.

A saved hour is useful. A redesigned workflow that prevents the same demand from returning every week is leverage. A clearer decision boundary makes the entire team stronger.

My broader goal is to leave behind a legacy of useful knowledge—methods that other people can test, challenge, improve, and carry forward. I hope this guide becomes part of that working library.

Download the Field Guide

You can download The Security Leader’s Operating System here:

Download The Security Leader’s Operating System

As you work through it, I would be interested to hear what you eliminate, delegate, simplify, automate, or deliberately maintain first.

Architecture Ages. Security Programs Should Plan for That

Every security architecture has an expiration date.

Most organizations simply do not know when it arrives.

An architecture may have been carefully designed, properly documented, and well defended when it was introduced. It reflected the business, technology, threat environment, and operational assumptions of its time.

Then the organization changed.

It adopted cloud services. It acquired another company. Employees became more distributed. Vendors received deeper access. Applications became API-driven. Automation expanded. AI entered business workflows. Identity became the connective tissue between systems that no longer shared a traditional perimeter.

The architecture did not suddenly fail.

It aged.

80sIBM

Continue reading

Account Recovery Is Becoming the New Identity Attack Surface

As passkeys and phishing-resistant authentication reduce password risk, attackers will move pressure to the recovery plane.

The industry is moving in the right direction.

Passkeys, FIDO2/WebAuthn, hardware security keys, conditional access, better MFA policies, and risk-based sign-in controls are all meaningful improvements. They reduce entire classes of credential theft. They make phishing harder. They remove reusable passwords from many authentication ceremonies. They shift more of the security burden from user judgment to protocol design.

That is good.

But it is not the finish line.

In my recent passkeys article, I called out a point that deserves its own treatment: passkeys do not solve weak account recovery, help desk social engineering, stolen session tokens, OAuth consent abuse, unmanaged vendor access, or excessive privilege. They are a major step forward, but they do not remove the rest of the identity attack surface.

That matters because attackers adapt.

Continue reading

Passkeys, Not Passcodes: A Practical Enterprise Guide to Moving Beyond Passwords

There is a small terminology problem in the identity world right now, and it matters more than it looks.

passcode or PIN is usually a local unlock secret. It unlocks a phone, a laptop, Windows Hello, an authenticator app, or a hardware security key. A passkey is different. A passkey is the standards-based replacement for passwords, built on FIDO2/WebAuthn. The user unlocks the passkey locally with a fingerprint, face scan, device PIN, pattern, or security key, but the website or application receives cryptographic proof — not a reusable password. FIDO defines passkeys as FIDO authentication credentials based on FIDO standards, tied to an account, and used with the same process the user already uses to unlock a device.

That distinction is not pedantry. It is the difference between a local unlock method and a replacement for one of the most abused controls in the history of computing.

Passwords have had a long run. They also have had a long list of failures: reuse, phishing, spraying, stuffing, database theft, weak reset workflows, help desk abuse, and user fatigue. We have spent decades trying to compensate for those failures with complexity rules, expiration schedules, password managers, SMS codes, mobile push prompts, training campaigns, and detective controls.

Some of those helped. Some just moved the pain around.

Passkeys change the model.

Continue reading

Rational Security in the AI Era: How Attackers Are Evolving and How We Must Respond

The weaponization of artificial intelligence by cybercriminals and nation-state actors has crossed a critical inflection point. We no longer live in a world where we can rely solely on traditional perimeters; the threat landscape has fundamentally shifted into what we might call “Extremistan,” where the speed and scale of attacks demand a completely new level of resilience.

SadKitty

At MicroSolved, our mission is to provide rational cybersecurity for an irrational world. To do that effectively, we must look unflinchingly at the data.

The Problem and the Metrics

The numbers tell a stark story of industrialization at machine speed. According to recent threat reports, AI-enabled adversaries increased their attack volume by 89% year-over-year. More concerning is the velocity: the average eCrime breakout time has collapsed to just 29 minutes, with the fastest recorded intrusion moving from initial access to lateral movement in a staggering 27 seconds.

The financial impact is equally severe. The FBI IC3 recorded over 22,000 AI-related complaints with adjusted losses exceeding $893 million in 2025 alone, including tens of millions lost to AI-enabled Business Email Compromise (BEC). AI is accelerating attack speeds by 4x, making human-speed incident response no longer viable.

Continue reading

The Evidence Supply Chain: How CISOs Build a Cyber Materiality Data Plane Before the Incident

A ransomware incident does not wait for the organization chart to catch up.

At 8:17 a.m., the SOC sees encryption activity on a file server. At 8:31, operations says the plant is still running. At 8:44, finance says revenue recognition may be affected if order processing stays down past noon. At 9:02, legal asks whether customer data was accessed. At 9:18, the forensic team says it is too early to tell. At 9:23, a vendor says the outage may have started in their environment. At 9:41, communications asks whether they should prepare a holding statement.

By hour two, everyone is working hard.

But they are not necessarily working from the same reality.

That is the problem.

Cyber materiality is often discussed as a decision problem. When does a cyber event become a board-level business event? When does it become reportable? When does it become material to investors, customers, regulators, lenders, or strategic partners?

Those are important questions. Public companies, for example, must disclose material cybersecurity incidents on Form 8-K within four business days after determining materiality, including the material aspects of the incident’s nature, scope, timing, and impact or reasonably likely impact.

But underneath that decision sits a deeper problem:

Continue reading

Cyber Materiality Engineering: How CISOs Pre-Decide When Risk Becomes a Board Event

A ransomware incident does not stay technical for very long.

For about the first fifteen minutes, it may look like a security operations problem. A strange alert. A locked server. A suspicious authentication chain. A vendor portal behaving badly. A handful of systems no longer responding the way they should.

Then the blast radius starts to widen.

Operations wants to know whether they can keep running. Finance wants to know whether revenue recognition, cash movement, reserves, or forecasts are exposed. Legal wants to know whether notification clocks have started. The CEO wants to know what can be said, to whom, and when. The board wants to know whether this is “material.” Investors may eventually ask the same thing, only with less patience and more lawyers.

This is where many organizations discover that their cyber incident response plan is not really an enterprise decision plan. It tells people who to call. It tells the SOC how to preserve evidence. It may even have a communications tree and a sample press statement.

But it often does not answer the question that matters most in the first few hours:

Continue reading

AI Agents Are Already Working for You. Who’s Managing Them?

AI Agents Are Not Applications. They Are Digital Workers.

Most organizations are adopting AI agents faster than they are learning how to govern them.

That is the problem.

A chatbot that answers questions is one thing. An AI agent that can access business data, use tools, trigger workflows, generate artifacts, make recommendations, or alter enterprise state is something else entirely.

At that point, the organization is no longer just deploying software.

It is introducing a new kind of operational actor.

That actor needs identity.

It needs boundaries.

It needs oversight.

It needs evidence.

It needs a human owner.

It needs a kill switch.

In other words, AI agents must be managed more like digital workers than ordinary applications.

AIAgentBanner

Continue reading

Rethinking Account Lockouts: Why 15 Minutes Isn’t a Strategy

There’s a moment in almost every security program where someone asks a deceptively simple question:

“Is 15 minutes a standard account lockout duration?”

The short answer? No.
The more honest answer? It’s common—but often wrong for the environment it’s deployed in.

And I’ve seen more than a few organizations learn that the hard way.

3Errors


The Myth of the “Standard” Lockout

If you go looking for authoritative guidance—from Center for Internet SecurityFFIEC, or CISA—you’ll notice something interesting:

They don’t tell you what number to use.

Instead, they consistently emphasize:

  • Risk-based decision making
  • Balancing usability and security
  • Detecting and responding to threats—not just blocking them

That’s not an accident. It’s an acknowledgment that static controls like lockouts are blunt instruments in a very dynamic threat landscape.


What We Actually See in the Real World

Across environments—financial services, healthcare, SaaS, manufacturing—the patterns are pretty consistent:

Setting Typical Range
Failed attempts before lockout 3–10
Lockout duration 5–30 minutes
Most common default 10–15 minutes

So yes, 15 minutes sits comfortably in the middle.

But “common” and “effective” are not the same thing.


Where 15 Minutes Breaks Down

1. It Punishes Users More Than Attackers

A 15-minute lockout sounds reasonable—until you multiply it.

  • A clinician locked out mid-shift
  • A call center agent missing SLAs
  • A trader unable to access systems during market hours

Now multiply that by repeated lockouts from cached credentials, mobile devices, or service accounts.

You don’t just have a security control—you have an operational problem.


2. It Doesn’t Stop Modern Attacks

Attackers have evolved. Most environments haven’t.

Today’s common attack patterns:

  • Password spraying (low-and-slow, avoids thresholds)
  • Credential stuffing (valid credentials, no lockout triggered)

A longer lockout duration doesn’t meaningfully impact either.

If anything, it gives a false sense of security while the real attack path goes untouched.


What Actually Works: A Layered Approach

This is where the conversation needs to shift—from “what’s the right number?” to “what’s the right strategy?”

1. Lockouts Are Supporting Controls—Not Primary Defenses

If you’re relying on lockouts as your main protection, you’re already behind.

At a minimum, you should be pairing with:

  • MFA everywhere it’s technically feasible
  • Conditional access (device, location, behavior)
  • Authentication throttling and smart detection

2. Tune for Risk, Not Defaults

A more balanced configuration tends to look like:

  • 5–10 failed attempts
  • 5–10 minute lockout
  • Reset counter after a defined cooldown window

This reduces user friction while still slowing down brute-force attempts.

More importantly—it acknowledges that lockouts are a speed bump, not a wall.


3. Progressive Delays Beat Hard Lockouts

One of the most underutilized strategies is progressive delay:

  • Attempts 1–2 → no delay
  • Attempts 3–5 → 30–60 second delay
  • Continued attempts → increasing delay

This approach:

  • Degrades attacker efficiency
  • Preserves user productivity
  • Avoids helpdesk spikes

It’s a far more surgical control than a blanket 15-minute lockout.


4. Detection Over Punishment

Modern security programs don’t just block—they observe.

You should be:

  • Logging all failed authentication attempts
  • Alerting on patterns (spraying, geographic anomalies)
  • Correlating identity signals across systems

Lockouts should be one signal among many—not the primary response.


Implementing This in Active Directory

Let’s get practical.

In on-prem Active Directory, you’re working primarily with Group Policy.

Recommended Baseline

In your domain or fine-grained password policy:

  • Account lockout threshold: 5–10 attempts
  • Account lockout duration: 5–10 minutes
  • Reset account lockout counter after: 10–15 minutes

Where to Configure

  • Group Policy Management Console (GPMC)
    • Computer Configuration → Policies → Windows Settings → Security Settings → Account Policies → Account Lockout Policy

Advanced Considerations

  • Use Fine-Grained Password Policies (FGPP) for high-risk accounts (admins, service accounts)
  • Monitor Event IDs:
    • 4625 (failed logon)
    • 4740 (account locked out)
  • Feed logs into your SIEM for correlation and alerting

Implementing This in Microsoft 365

In Microsoft 365, the model shifts significantly.

You don’t directly control “lockout duration” in the same way—because the platform is already applying smart lockout behavior.

Smart Lockout (Azure AD / Entra ID)

  • Automatically tracks failed attempts
  • Uses adaptive thresholds
  • Differentiates between familiar and unfamiliar locations

What You Should Do Instead

1. Enable and Enforce MFA

  • Conditional Access → Require MFA for all users (with staged rollout if needed)

2. Configure Conditional Access Policies

  • Block legacy authentication
  • Require compliant devices
  • Apply geographic restrictions where appropriate

3. Monitor Identity Signals

  • Azure AD Sign-in logs
  • Risky sign-ins and users
  • Integration with Defender for Identity / Sentinel

4. Tune Smart Lockout (if needed)

  • Default threshold is typically sufficient
  • Adjust only if you have a strong operational reason

The Bottom Line

A 15-minute lockout isn’t wrong.

It’s just incomplete.

  • ✔️ It’s common
  • ❌ It’s not a standard
  • ⚠️ It can create more operational pain than security value

The real shift is this:

Stop treating account lockouts as a primary control. Start treating them as part of a layered identity defense strategy.

Because in today’s environment, the goal isn’t just to block access.

It’s to understand it.

 

 

* AI tools were used as a research assistant for this content, but human moderation and writing are also included. The included images are AI-generated.