Skip to content

Evaluation

Twelve questions to ask an AI SOC vendor

Feature comparisons in this category age badly, because the features converge within a quarter and the differences that survive are architectural and contractual. The twelve questions below are chosen on one criterion: every answer is demonstrable in a live session rather than assertable in a slide. They are grouped by the four functions of the NIST AI Risk Management Framework, which gives a buyer and a vendor a shared structure neither of them owns. Each question comes with the answer that should reassure you and the answer that should not, because a question without a calibration is just a prompt for a well-rehearsed reply.

Questions
12
Structure
NIST AI RMF
Criterion
Demonstrable
Red flag
A percentage
Written byMalthe Bang NorengaardCo-founder & CTOReviewed byRobin ÖsterdalFounder & CEO

Structured on the NIST AI RMF functionsLast reviewed 6 min read

Kort sagt

  • Feature comparisons converge within a quarter. Architecture and contract terms do not.
  • Every question here has an answer a vendor can show you rather than tell you.
  • Group by the NIST AI RMF functions so the structure belongs to neither party.
  • An accuracy percentage is a red flag, because it depends on an estate and a baseline that do not transfer.
  • Ask about failure paths specifically. Uncertainty, tool errors and an unreachable approver are three different cases.
  • For regulated buyers, where inference runs is a contract fact and belongs in the procurement record.

Why feature comparisons age badly here

A feature matrix in this category has a shelf life of roughly one quarter. Capabilities converge because they are largely downstream of the same model generation, and a differentiator described in a January briefing is table stakes by summer.

What does not converge is architecture and contract. Which actions run unattended, whether verdict reasoning is retained and exportable, where inference executes and under whose paper: those are decisions a vendor made and will not quietly change, and they are the ones you will still be living with at renewal.

There is also an asymmetry worth exploiting. A vendor can decline to answer a feature question by promising a roadmap. They cannot decline to show you a permission screen without telling you something.

Once this significant share is redacted, the top five targeted sectors in the EU include public administration (38.2%), transport (7.5%), digital infrastructure and services (4.8%), finance (4.5%) and manufacturing (2.9%).

What the European incident data frames the questions against

38.2 %
of EU incidents hit public administration, the most targeted sector

Källa: ENISA Threat Landscape 2025

7.5 %
hit transport, which emerged as a high-value sector in the period

Källa: ENISA Threat Landscape 2025

4.8 %
hit digital infrastructure and services

Källa: ENISA Threat Landscape 2025

4.5 %
hit finance, and 2.9 % manufacturing

Källa: ENISA Threat Landscape 2025

Govern: who decides what runs unattended

These three questions establish whether unattended action is designed or defaulted. They are first because every later answer is conditioned on them.

Ask which specific actions execute without approval, and ask to see that list in the product rather than in documentation. Ask who is permitted to change it, and whether the change appears in an audit trail. Then ask what happens when the model is uncertain, when a tool call fails, and when the approval channel is unreachable.

The last of those three separates most vendors. A system that acts when it cannot reach an approver has an approval feature rather than a boundary, and the difference only shows up at three in the morning.

An LLM-based system is often granted a degree of agency by its developer, the ability to call functions or interface with other systems via extensions.

Govern: three questions and their calibration
QuestionA good answerA concerning answer
Which actions run without approval?A visible list in the product, per action typeIt depends on your configuration
Who can change that, and is it logged?A named role, and the change is in the audit trailAny administrator, no separate record
What happens if the approver is unreachable?It holds and escalatesIt proceeds when confidence is high

Källa: OWASP LLM06:2025 Excessive Agency

Map: what is in scope and what is deliberately not

Scope questions surface the assumptions a product makes about your estate, and they tend to reveal more than feature questions do.

Ask what data sources the system reads and what it does when a source is missing or stale, since a verdict formed on partial context is a different object from one formed on complete context. Ask what parts of the estate it will not act on under any configuration. Ask how a new detection type is onboarded, and by whom.

The second question is the one worth pressing. A vendor with a considered answer will name specific exclusions such as domain controllers or production databases. A vendor without one will say everything is configurable, which means the boundary is yours to invent.

Limit the extensions that LLM agents are allowed to call to only the minimum necessary.

Map: what the answers tell you
Designed scopeConfigurable everything
Names specific systems it will not touch
Behaviour on a stale data source is defined
Verdicts record which sources were unavailable
Onboarding a detection type has an owner~
Boundary is the customer's to invent

Configurability is not a fault in itself. The question is whether the vendor has an opinion about the safe default, because that opinion is what you are buying alongside the software.

Källa: NIST AI Risk Management Framework

The vectors a scope answer should account for

60 %
of European cases began with phishing and its variants

Källa: ENISA Threat Landscape 2025

21.3 %
began with vulnerability exploitation

Källa: ENISA Threat Landscape 2025

9.9 %
came from botnets, and 8 % from malicious applications

Källa: ENISA Threat Landscape 2025

0.8 %
was insider access, the case correlation exists for

Källa: ENISA Threat Landscape 2025

Measure: how anyone knows the verdicts are right

This is where accuracy percentages appear, and where they should be declined. A number produced on someone else's estate with someone else's baseline does not describe what will happen on yours.

Ask instead whether the reasoning behind each verdict is retained and exportable, because a verdict you cannot re-examine cannot be reviewed and therefore cannot serve as evidence. Ask how disagreement is captured when an analyst overrules the system, and whether that feeds anything. Ask for two or three real cases where the system was wrong, and what happened next.

The last question is diagnostic in a way the others are not. A vendor who cannot produce an example of being wrong is either not looking or not telling, and both are worse than the errors would have been.

Excessive Agency enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM.

What the standards say to measure

4 functions
in the NIST AI RMF, with measure covering verification and evidence

Källa: NIST AI Risk Management Framework

10 risks
in the OWASP LLM Top 10 for 2025, of which misinformation applies directly to verdicts

Källa: OWASP Top 10 for LLM Applications 2025

24 hours
to an early warning under NIS2, which is what the verdict feeds into

Källa: NIS2 Article 23

Manage: what happens when it is wrong

The final three questions are about the day after, and they are the ones most likely to be answered honestly, because nobody rehearses them.

Ask what the rollback path is for each automated action class, and how long it takes. Ask what record exists of what the system did and, equally, of what it declined to do. Ask where inference runs, under whose contract, and whether that changes when a model provider is switched.

For an organisation inside NIS2 or DORA scope the last one is not preference but procurement. Article 21 makes supply chain security a risk management requirement in its own right, which means your detection vendor's own dependencies are inside your assessment whether or not you have looked at them.

Manage: three questions and their calibration
QuestionA good answerA concerning answer
What is the rollback path per action class?Named per class, with a time estimateEverything is reversible
What record exists of holds, not just actions?Both, in one exportable trailActions are logged
Where does inference run, under whose contract?A named region and a named providerIt varies by workload

Källa: NIS2 Directive Article 21, EUR-Lex

The regulatory questions that belong in the same conversation

For an organisation inside NIS2 scope, a detection vendor is a supplier inside a regulated estate, which means part of the evaluation is not about the product at all.

Article 21 makes supply chain security a risk management requirement in its own right, so the vendor's own dependencies are inside your assessment whether or not anyone has looked at them. That includes the model provider, the hosting arrangement, and anything the system calls out to during an investigation.

The sanction structure is worth having in mind while asking, not as a scare tactic but because it sets what proportionate diligence looks like. The ceiling is 2 percent of worldwide turnover for essential entities and 1.4 percent for important ones, which is the scale against which a procurement question about inference location is obviously reasonable.

Limit the functions that are implemented in LLM extensions to the minimum necessary.

The obligations the answers have to satisfy

24 hours
to an early warning of a significant incident, which the verdict feeds

Källa: NIS2 Article 23

72 hours
to an incident notification carrying an initial assessment

Källa: NIS2 Article 23

1 month
to a final report, which is where the audit trail is read

Källa: NIS2 Article 23

2 %
of worldwide turnover as the ceiling for essential entities

Källa: NIS2 Article 34

1.4 %
the corresponding ceiling for important entities

Källa: NIS2 Article 34

53.7 %
of recorded EU incidents involved essential entities

Källa: ENISA Threat Landscape 2025

The two questions that are not about the product

Ask what would make the vendor tell you this is not a fit. A vendor with a real answer will name an estate size, a sector, a maturity level or a data constraint where the product does not earn its cost. A vendor without one is describing a product that fits everyone, which no product does.

Then ask who else in your sector runs this at your size, and ask to speak to them. This is the only reference question that produces something checkable, because the answer is either a name or a reason there is no name, and both are informative.

Neither question is adversarial. Both are faster than a proof of concept at establishing whether the vendor has thought about where their product stops.

Utilise human-in-the-loop control to require a human to approve high-impact actions before they are taken.

How to run the conversation

Send the twelve questions in advance. The goal is not to catch anyone out but to get past the demo script to the people who know the answers, and a vendor with good answers will welcome the chance to prepare them.

Record the answers in the procurement file with the date. Products in this category change quickly, and an answer that was true in March is not evidence about the system you are running in November.

Keep one thing in view while you read the answers. The category's own risk register is unusually blunt about where the hazard sits, and a vendor whose answers do not engage with that has either not read it or would rather you did not.

An LLM-based system is often granted a degree of agency by its developer, the ability to call functions or interface with other systems via extensions.

Questions

Common questions

What should we ask an AI SOC vendor first?
Which specific actions execute without approval, shown in the product rather than described in documentation, and what happens when the approval channel is unreachable.
Why avoid accuracy percentages?
Because the number depends on an estate, a baseline and a definition of correct, none of which transfer to your environment. Ask for reproducible cases instead.
How do we compare two similar products?
On architecture and contract rather than features. Features converge within a quarter; unattended action policy, verdict exportability and inference location do not.
What is the single most diagnostic question?
Ask for two or three real cases where the system was wrong and what happened next. A vendor who has none is either not measuring or not telling.
Does the vendor's own supply chain matter?
Yes, if you are in NIS2 scope. Article 21 makes supply chain security a risk management requirement in its own right, so your detection vendor's dependencies sit inside your assessment.

Primärkällor

Källor

Varje regulatoriskt påstående på den här sidan går att spåra till en av källorna nedan. Ingen av dem är en konsultblogg.

Further reading

Working out what to ask a vendor in this category?

We will go through your evaluation criteria with you, including the questions that make us look worse. Half an hour, no preparation needed.