← UTIOM

From ATT&CK Evaluations to an Organisational Security Decision

Independent evidence is valuable. Context decides what it means.

MITRE ATT&CK Evaluations gives the industry something rare: independent, behaviour-based evidence of how security products and services perform against defined scenarios based on real-world adversary tradecraft. This is not a vendor demonstration, and the results are not vendor-controlled.

The 2026 methodology goes further than previous rounds. It introduces the Total Evaluation Score (TES), combining detection quality and protection quality into one comparable measure. It considers detection coverage, precision and speed, along with protection timing and false-positive performance. Techniques are weighted according to their criticality within the evaluated attack chain, and every weight needs a written justification. Investigation quality is scored separately, on whether the investigation established who, what, when, where and how. Precision penalises fragmented cases, not just false positives. And next to the score, MITRE makes something else visible: who or what produced the capability — platform automation, AI, human analysts, a managed service, or a combination.

This is good work. Several of those design choices — weight by criticality, reward one useful alert over ten noisy ones, refuse a score without a justification — are the same instincts UTIOM was built on. MITRE reached them independently.

UTIOM does not compete with it.

But it answers a different question from the one an organisation ultimately has to answer.

ATT&CK Evaluations asks: What did this product or service demonstrate against this scenario?

UTIOM asks: Does that capability matter to us, against our threats, on the attack paths that lead to what we actually need to protect?

Those two questions belong together. They should never be confused.

Figure 1 — A demonstrated capability is not an organisational capability Two boundaries. The evaluation proves what happened inside the first. Only the organisation can prove what happens inside the second. The evaluation range built by MITRE · one scenario · one adversary objective Adversary emulation from published CTI Product or service under test Telemetry the range provides Weighting by the adversary's objective (ACW) Scoring: coverage · precision · speed · timing Who acted: platform · AI · human · service Not in the range: your crown jewels · your configuration · your analysts your approval chain · your out-of-hours authority your business impact decisions · your rehearsed playbooks "The capability was demonstrated." evidence not a verdict Your organisation your estate · your threats · your clock Crown jewels and acceptable risk Threat profile and modelled attack paths 150–250 behaviours that matter, not 697 TID-CMM telemetry exists · detection engineered · validated against emulation TIR-CMM authority pre-granted · playbook rehearsed · margin against breakout Security Assurance evidence relevant · current · operational · traceable · independent Only measurable here: whether the behaviour is observable in your telemetry, and whether someone is permitted to act on it before the adversary moves "We can demonstrate that it works here." Independent evidence Organisational relevance Local capability Operational readiness Assurance Decision ATT&CK Evaluations UTIOM TID-CMM TIR-CMM Security Assurance UTIOM · CC BY-SA · Not affiliated with or endorsed by MITRE. MITRE ATT&CK® is a registered trademark of The MITRE Corporation.
Figure 1 — A demonstrated capability is not an organisational capability

A demonstrated capability is not an organisational capability

A product can perform extremely well in an independent evaluation and still be the wrong capability for a particular organisation.

That is not a weakness of the evaluation. It is the difference between evaluating a capability and designing a security operation.

MITRE's 2026 methodology recognises that not every adversary behaviour has equal significance. Attack Chain Weighting gives greater importance to techniques that are more critical within the evaluated attack path.

That is the right instinct.

But the weighting is anchored to the adversary's objective in the scenario being evaluated. It cannot be anchored to your crown jewels, because MITRE has never seen them.

The same is true of the environment.

The evaluation runs in a controlled environment built for the evaluation. Your architecture, telemetry, identities, configuration, analysts, business dependencies and approval chain are not part of that environment.

So a strong evaluation result proves something important: the technology or service demonstrated the capability to detect or stop the behaviour under the evaluated conditions.

It does not prove that the same behaviour is detectable in your estate, with your telemetry, your configuration and your operating model, or that your organisation can act on it in time.

The distinction is simple:

MITRE evidence: the capability was demonstrated. TID-CMM evidence: we can demonstrate that the capability works here.

Both are valuable. They are not interchangeable.

The evidence-to-decision chain

UTIOM places evaluation evidence inside a larger chain. MITRE's result enters from the side, as one input. It does not sit at the top, and it does not sit at the bottom.

Figure 2 — The evidence-to-decision chain Where independent evaluation evidence enters, and what it has to pass through before it becomes a decision. Business objective and acceptable risk Crown jewels and critical services Threat profile Relevant attack paths Relevant ATT&CK behaviours Independent capability evidence including ATT&CK Evaluations Local telemetry and detection capability TID-CMM Local response and containment readiness TIR-CMM Evidence and assurance Capability, engineering and procurement decision What matters to us anchored to the business, not to a scenario What works here measured in our estate, by our people, on our clock MITRE ATT&CK Evaluations 2026 TES = DQI + PQI IQI reported separately ACW chain weighting Source-of-action modifier "The capability was demonstrated." "We can demonstrate it works here." "We can act on it in time." "We can prove it." Rust marks evidence. The chain never rescores MITRE's result; it decides what the result means in the organisation that will depend on the capability. UTIOM · CC BY-SA · Not affiliated with or endorsed by MITRE. MITRE ATT&CK® is a registered trademark of The MITRE Corporation.
Figure 2 — The evidence-to-decision chain

The purpose is not to rescore MITRE's results. The purpose is to understand what those results mean in the organisation that will eventually depend on the capability.

Detection: ATT&CK Evaluations and TID-CMM

MITRE evaluates detection from the capability outward: coverage, precision, speed, context and correlation. It rewards useful signal, not more alerts. One alert that says who, what, when, where and how scores higher than fifteen that say "suspicious".

That matches UTIOM's own doctrine: alert fatigue is a system design failure, not a staffing problem.

TID-CMM approaches detection from the organisation inward.

It starts with your platforms, your relevant adversaries and your modelled attack paths, and from those derives the ATT&CK behaviours worth defending against.

That typically leaves a working set of around 150 to 250 techniques rather than treating all 697 Enterprise ATT&CK techniques and sub-techniques as a requirements list.

Then TID-CMM asks progressively harder questions:

Does the telemetry exist? Has the detection been engineered? Has it actually been proven to work?

If the telemetry required to observe the behaviour does not exist, no evaluation score can solve the problem locally.

That is not a gap in your rule set. It is a gap in physics.

An independent evaluation can show you that a technology is capable. TID-CMM determines whether that potential survived contact with your actual environment.

Response: ATT&CK Evaluations and TIR-CMM

The 2026 methodology also puts much more emphasis on protection. An earlier successful interruption is more valuable than stopping the adversary after critical actions have already occurred.

Timing matters.

TIR-CMM reaches the same conclusion from the organisational side: earlier containment buys the one thing response cannot manufacture — time.

But the evaluation boundary is different again.

MITRE can demonstrate that a product or service is technically able to stop adversary behaviour. TIR-CMM asks whether your organisation can actually use that ability when it matters.

A containment capability may exist technically and still fail operationally because the SOC is not authorised to execute it, because approval requires several management layers, because nobody accepted the business impact in advance, because there is no out-of-hours authority, or because the playbook has never been rehearsed.

A product can therefore be fast while the organisation around it is slow.

That is not a product problem. It is an operating-model problem.

MITRE's own methodology page names this layer — response, decisioning, resilience — and marks it as proposed for future versions. It is not scored yet.

That is the layer TIR-CMM is designed to expose, today.

Evaluation speed is not response speed

Product speed and organisational response speed are not the same thing.

MITRE runs two clocks. Detection Speed runs from adversary action to the first automated alert. Conclusion Speed runs from that alert to a defensible investigation conclusion. Both are the right clocks for measuring a product.

The organisation faces a longer clock:

Adversary action → Detection → Analysis → Decision → Authorisation → Containment

Figure 3 — Evaluation speed is not response speed MITRE's clocks stop where the product stops. The organisation's clock runs until containment lands — and it is racing the adversary. t = 0 time → Adversaryaction Firstalert Analysisconclusion Decision Authorisation Containmentlands MITRE ATT&CK Evaluations clocks product performance, measured in a controlled range Detection Speed action → first automated alert <15 min = 1.0 · 15–30 min = 0.75 · >30 min = 0.5 Conclusion Speed (IQI) first alert → defensible conclusion banded by scenario size not measured — "Layer 3: Outcome" is future scope in the 2026 methodology The organisation's clock — TIR-CMM operational capability, measured in your estate, against your management chain MTTD MTTDecide MTTC detect validate · analyse · decide · obtain authority execute containment invisible in every product evaluation, because your authority chain is not in the range Adversary breakout time breakout · average ≈ 29 min, fastest recorded 27 s Containment Margin = Breakout − (MTTD + MTTDecide + MTTC) Here the margin is negative: containment landed after breakout. The product was fast. The organisation was slow. A DS score of 0.75 ("acceptable delay", 15–30 min) can consume the entire average breakout window before anyone has decided anything. UTIOM · CC BY-SA · Not affiliated with or endorsed by MITRE. MITRE ATT&CK® is a registered trademark of The MITRE Corporation.
Figure 3 — Evaluation speed is not response speed

UTIOM treats that entire chain as operational capability.

TIR-CMM separates:

Mean Time to Detect Mean Time to Decide Mean Time to Contain

and compares their combined duration with adversary breakout time.

The result is the Containment Margin:

Containment Margin = Breakout Time − (MTTD + MTTDecide + MTTC)

If the margin is positive, containment lands before the adversary spreads. If it is negative, the organisation is already responding to a spreading intrusion, however good the original detection was.

A five-second technical containment capability is worth very little if permission to use it takes forty minutes.

That decision interval is one of the most fixable failures in incident response. And an external product evaluation cannot measure it, because your management and authority chain is not part of the evaluation environment.

How UTIOM uses an ATT&CK Evaluation

Say an organisation is choosing an endpoint, identity, XDR, SIEM, MDR or AI-SOC capability.

The usual approach starts with features, product categories or the highest score.

UTIOM starts somewhere else.

1. Establish what matters Identify critical services, crown jewels, trust boundaries and unacceptable consequences.

2. Model the relevant threat Determine which adversaries and which attack paths can create those consequences.

3. Derive the relevant behaviours Use ATT&CK to describe the behaviours required along those paths.

4. Read the independent evidence Use ATT&CK Evaluations to understand what candidate capabilities demonstrated against those behaviours, rather than treating every tested behaviour as equally relevant to your organisation.

5. Test local detection Use TID-CMM to determine whether the telemetry exists, whether detection has been engineered and whether it has actually been validated in your environment.

6. Test local response Use TIR-CMM to determine whether the organisation can investigate, decide, contain and recover inside the available Response Horizon.

7. Test the evidence Ask whether the evidence supporting each claim is relevant, current, operational, traceable and sufficiently independent for the decision being made.

8. Decide

The output is not: Which vendor scored highest?

It is: Which capability best reduces our relevant operational risk, and what must we change internally before we can depend on it?

TES is useful. It is not organisational fit.

TES makes detection and protection performance comparable inside MITRE's evaluation methodology.

MITRE publishes the detailed per-technique results alongside the score so that customers can see what was demonstrated and judge its relevance to their own environment. The score is the summary. The results are the evidence.

UTIOM takes that second step seriously.

A high evaluation score does not automatically mean high organisational value.

A capability can perform exceptionally against the evaluated scenario while contributing relatively little to the attack paths that dominate your actual risk.

Another capability may have a lower overall score while performing strongly against the smaller set of behaviours that matter most to your organisation.

The objective is not maximum generic coverage. It is: maximum defensible capability against relevant operational risk.

The roles

MITRE ATT&CK gives everyone a common language for adversary behaviour.

MITRE ATT&CK Evaluations provides independent evidence of what products and services demonstrated against controlled adversary scenarios.

UTIOM provides the operating model connecting business consequence, threats, engineering, detection, response and continuous improvement.

TID-CMM determines whether relevant adversary behaviour can actually be observed and detected in your environment.

TIR-CMM determines whether you can decide and act before the available time runs out.

Security Assurance determines whether the evidence behind those claims is strong enough for the decision you are about to make.

The chain becomes:

Independent evidence → organisational relevance → local capability → operational readiness → assurance → decision

That is where an evaluation becomes a security decision.

The principle

Independent evidence should inform a security decision. It should not make the decision for you.

ATT&CK Evaluations tells you what a product demonstrated. UTIOM tells you whether it matters to you. TID-CMM tells you whether you can see the behaviour. TIR-CMM tells you whether you can act on it in time. Security Assurance tells you whether you can prove any of it.


UTIOM is an independent project and is not affiliated with or endorsed by MITRE. MITRE ATT&CK® is a registered trademark of The MITRE Corporation. Figures are CC BY-SA.

Join the UTIOM community. Discuss, contribute evidence and share implementation experience. About the community →