- Zen IT Technologies
- Security Readiness & Response
Your ability to respond is built before you need it.
Security response is not something we add after an incident. We build the capability into the environments we manage: security telemetry and the retention behind it, detection maintained against current attack patterns, evidence that can be collected without arranging physical access, and the inventory and reach to scope an incident across identity, endpoints, network and the SaaS estate.
When a new indicator of compromise is published, a dependency is found to be compromised, or an incident happens, the environment is already structured to answer the questions that decide the outcome. What is affected. What happened. What has to be contained. How you recover.
The problem
Most organizations discover their response capability during the incident.
The alert fires and the questions start immediately. Which machine. Which user. What did they touch. Was anything else using that credential. And underneath all of them, the one that actually decides the outcome: can we even find out?
The gaps are always the same, and always identified too late:
-
No forensic tooling in place.
Acquisition depends on physical access to a device that may be in another country.
-
Logs already gone.
Google Workspace, Microsoft 365, and source-control audit logs aged out on default retention while the incident was still being scoped.
-
No inventory to scope against.
Nobody can list which SaaS applications hold company data, or which service accounts exist.
-
No decision rights agreed.
The question of who can authorize taking production offline at 3 a.m. is being worked out during the incident.
-
Nothing written afterwards.
The incident resolves, and six months later a customer, auditor, or insurer asks what happened.
What incident response readiness covers
- Identity
- Endpoints
- Network
- SaaS and cloud
-
Response capability, built in
Response readiness is part of running an environment properly rather than a separate product. In practice it is a set of requirements placed on systems that already exist: acquisition tooling reaching every managed device through the endpoint management platform, so remote collection is normally available without first arranging physical access to a machine that may be in another country; audit-log retention set against realistic investigation timelines rather than platform defaults; detection logic maintained against current attack patterns; asset and identity inventory current enough to scope against.
We decide what has to be collectable and confirm that it is. The endpoint, identity and network platforms remain the control planes those requirements are implemented on, which is why readiness holds as the estate changes instead of being rebuilt for each incident.
-
Establish the baseline
We start by establishing what can actually be detected, investigated and recovered today: tooling coverage, log retention, inventory completeness, identity exposure, recovery capability and decision rights.
The distinction that matters is between a setting being enabled and the evidence being retrievable, so where it is practical we confirm the path rather than accept the configuration. A retention policy nobody has ever read back is a policy, not a capability. The gaps become part of the remediation roadmap. Most of them close with capability already owned and licensed rather than with a new platform, and the roadmap is maintained as the environment changes.
-
Monitoring and Observability
The security signal path has to be designed and kept working: which sources are collected, where they land, how long they are held, and what is allowed to raise an alert. The right shape is not the same at every company.
Most environments already produce more signal than anyone uses, so the first move is usually to get more out of the identity and endpoint platforms already in place. Authentication and risk events become a small number of notifications somebody will actually read; identity risk signals feed the conditional-access policy already enforced by the identity platform; an endpoint detection triggers the platform’s own response where the condition is clear enough to warrant it. For a smaller company that is often most of the practical coverage, out of licenses already paid for.
Centralized telemetry, longer correlation and SIEM pipelines get built where the estate justifies them. Either way the sources have to stay healthy, because one that quietly stopped reporting looks identical to one with nothing to report. Uptime and availability monitoring belongs to running the network and the endpoint fleet. What belongs here is whether the events that matter are collected, retained and noticed.
-
Response planning
Runbooks for the scenarios that actually occur: endpoint compromise, credential and session takeover, dependency compromise. Each with a containment sequence that preserves evidence rather than destroying it. Decision rights, escalation paths, and communication responsibilities agreed while everyone is calm.
A runbook is written against the tooling this environment has, which is why planning follows remediation rather than preceding it: a plan that assumes an acquisition path nobody has built fails at the first step.
-
Query and act across the estate
When an indicator of compromise is published, the questions are straightforward: which machines have it, do we hold the logs that would show it, can we query them, and can we deploy a check across the estate?
Answering takes four different systems and no one of them is sufficient. Endpoint telemetry and fleet-wide query come through the management platform; sign-in context, session state and access history come from identity; administrative and activity history comes from the SaaS and cloud audit logs; network telemetry comes from the edge.
We determine what has to be asked, read the answers together, and coordinate whatever action follows through the platform that owns it. The capability is not in any single one of those consoles. It is in being able to put their answers beside each other while the incident is still open.
-
Security Incident Response and Digital Forensics
When something is live the first job is triage: establishing what this actually is, how far it reaches, and what has to happen in the next hour. Scope is determined from evidence rather than assumed from the alert.
That means endpoint acquisition and analysis, audit-log acquisition across Google Workspace, Microsoft 365 and source-control platforms, and whatever the environment can be made to say about credential reuse and lateral movement. Containment is coordinated with whoever holds each control plane, so that what stops the attacker does not also destroy the record.
Afterwards there is a written report: timeline, entry vector, scope, containment actions, root cause and remediation status, in a form that survives being read by a customer, an auditor or an insurer. Remediation is tracked to closure, and what the incident revealed about detection goes back into the detection.
-
Recovery after a security incident
Response does not end at containment. Before anything is restored we determine what has to be preserved, because a rebuild performed too early destroys the evidence that would have explained the incident.
Restoration is then sequenced against containment rather than against pressure: systems come back once the access path is closed and the remediation is confirmed, not once the backup finishes. Backup, restore testing and continuity planning belong to running the wider environment. What belongs here is the security judgment applied to them: what to preserve, when restoring is safe, and how to establish that what came back is clean.
Technical notes
Reading a fleet-wide detection spike
The shape of the spike identifies it faster than the file analysis does.
How it runs
-
Assessment.
A documented review of your current response capability, with the gaps ranked by what they would cost you during an incident.
-
Remediation.
The gaps closed on the platforms that already run the estate: acquisition tooling reaching the fleet, retention extended where the default would expire mid-investigation, inventory made complete enough to scope against, detection written for the attacks that apply to you.
-
Planning.
Runbooks for the scenarios that apply, and decision rights settled in advance: who can authorize taking production offline, and who decides when that person cannot be reached.
-
Maintained.
Reviewed as the estate changes and as attack patterns change. A plan written once is a plan that will be wrong.
This is not an emergency service brought in only when something goes wrong. Readiness is maintained as part of the ongoing service, so when an incident occurs we are responding from an environment we already know and manage.
Proof
-
Forensic capability deployed before it was needed.
Technology company
Acquisition tooling rolled out fleet-wide through the existing management platform, so remote collection is the normal path rather than something arranged during the incident, including for staff and contractors working from another country. Where a device is reachable and enrolled, the investigation starts by collecting evidence rather than by deploying the means to collect it.
-
Audit log retention extended past platform defaults.
SaaS platform
Retention windows across identity, collaboration, and source-control platforms assessed against realistic investigation timelines, then extended where the default would have expired mid-investigation. Acquisition paths documented alongside them.
-
Detection rebuilt during a live campaign.
Technology company
Indicators from an active supply-chain campaign were turned into fleet-wide detection within the day, forensic collection was deployed across the estate, and the investigation produced a definitive assessment of exposure, delivered as a formal report.
-
A fleet-wide EDR malfunction contained and resolved with the vendor.
Technology company
An operating system update put the endpoint platform into a detection loop across the fleet, flagging components of the OS itself. The issue was triaged fleet-wide, an exclusion policy was engineered rather than applied per alert, and the fault was escalated to the vendor and tracked to a fixed release.
Client examples are anonymized by design. References are provided privately, on request, and with the client's consent.
Who this is for
Companies that would rather find out what they can investigate now than during an incident. Typically ahead of an audit, after a near miss, or when a customer security questionnaire asks a question nobody can answer.
Usually the systems are already there. Identity, endpoints, SaaS and the network each produce something useful, and the gap is that nothing joins them into an answer anybody can act on. That is a readiness problem rather than a budget one, and the work is sized to the environment: what a company of thirty people needs to be able to investigate is not what a company of three hundred needs, and neither of them gets there by buying the largest available security stack.
Frequently asked questions
-
Do you take emergency incident work from companies you do not already work with?
No. Response starts from an environment we already have access to, whose identity model we already know, and whose endpoints we already manage. Arriving cold during an incident removes all three, and a large part of the first day goes on getting them back. If you are in an active incident and not a client, you need a dedicated response firm that can mobilize today.
-
Do you provide 24/7 SOC or MDR monitoring?
Not as a staffed watch floor, and that is a narrower gap than it sounds. Telemetry, detections, notifications, automated workflows and response integrations keep running whether or not somebody is looking at a console. Configured identity risk events can raise an alert or tighten access, and endpoint detections can trigger containment where the platform and configuration support it, with both reaching the people who need to know through the channels they already use. For many smaller companies that is substantial practical coverage, built mostly out of platforms already licensed, with centralized telemetry or SIEM added where the estate justifies it.
What is genuinely separate is continuous human monitoring. An analyst queue watching around the clock is an MDR or SOC service, and where that is the requirement it can be added on top of this architecture rather than replacing it. It is not included here by implication.
-
What does a readiness assessment actually cover?
What you could investigate right now: forensic tooling coverage, audit log retention and acquisition paths, asset and identity inventory completeness, session revocation capability, and the decision-rights gaps that cost hours during a real incident. You get a written baseline and a prioritized remediation list. An assessment works as a defined piece of work on its own, whether or not we already run the environment. Responding to a live incident is the part that depends on already knowing it.
-
How long should we retain audit logs?
Longer than the default, which is shorter than most people assume and varies by platform and license tier. Investigations routinely need to reach back weeks to establish when access began, so a thirty-day window frequently fails at the moment it matters. Know the retention period for each platform and extend it where necessary.
-
Is EDR enough on its own?
Not on its own, and not because it is weak. A current EDR or XDR platform detects, collects evidence and responds at the device layer, including containing a machine that needs containing. What it does not hold is the rest of the estate: which SaaS data was reached, whether the same credentials were used elsewhere, what the sign-in and audit history shows. Scoping an incident across the estate needs those beside the endpoint picture, which is the work this service does.
-
Who writes the incident report?
We do, for environments we run. The report covers the timeline, entry vector, scope, data accessed, containment actions, root cause, and remediation status. It is written for leadership, customers, insurers, and auditors, with technical detail behind a plain-language summary.
Build response capability into your environment.
A 30-minute call with Jonny. We reply within one business day.