When something happens · 4.3

Being able to answer afterwards

Being able to answer afterwards. What it costs, who decides, and what usually goes wrong.

For an independent reference point, see CISA incident response resources.

The question

Can you answer what happened, three weeks later

Forensic readiness is not an incident response capability. It is the set of things that must already be true for anybody to reconstruct events afterwards, and almost all of it is decided long before anything happens.

The test is simple. If a machine were compromised on a Tuesday and you learned about it the following month, could you establish which files were touched, by which account, and where they went?

Logs

Four properties, and most environments miss two

They exist. Many machine controllers log nothing usable, which is why governance at the point of transfer matters more here than logging at the endpoint.

They are retained. Default retention is frequently days. An incident discovered a month later needs a month of history, and the standard has expectations about retention.

They are somewhere else. Logs held only on the affected system are the first thing an intruder alters and the first thing a reimage destroys.

The clocks agree. Time synchronisation across systems sounds like housekeeping and is the difference between a sequence of events and a pile of records that cannot be ordered. Reconstructing an incident across systems whose clocks differ by minutes is substantially harder and sometimes impossible.

Imaging

Being able to take a copy without destroying the original

Preservation is a contractual obligation and it is also the practical prerequisite for finding out anything. Somebody needs to know how to take an image of an affected system, and needs to know before the morning it is required.

For most manufacturers the honest answer is that this capability is bought rather than built, which is fine provided the arrangement exists in advance.

Who does the analysis

Not you

Digital forensics is a specialism and an incident is a poor time to learn it. What a supplier needs is a relationship arranged beforehand: a firm, a contact, an agreed rate and an understanding of what they would need from you.

Retainers exist and are not expensive relative to the cost of spending the first two days finding somebody. Your insurer may also require a firm from their panel, which is worth knowing before you call your own.

The floor

Where the evidence is thinnest

Machine controllers are the part of the environment least likely to record anything and most likely to be reimaged by a maintenance engineer trying to be helpful.

Two things help: a record kept at the point files are authorised and transferred, which exists independently of the machine, and an instruction to maintenance staff that a suspected compromise is not to be resolved by reimaging.

The tabletop

Two hours, once a year

A scenario, the people who would actually be involved, and a facilitator asking what happens next. No systems, no technology, a whiteboard.

What it finds is consistent: the contact list is out of date, nobody knows who decides, the certificate has expired, and two people believed somebody else was responsible for the same thing.

It also produces a record of an exercise, which the standard's incident response requirements expect to exist.

The paper copy

Because the network may be the problem

A printed plan with the contacts, the decision tree and the reporting requirements, held by the named person and their deputy.

This sounds old-fashioned until the day the file server is encrypted and the plan is on it.

Also

Elsewhere in when something happens