Structured on the NIST AI RMF functionsLast reviewed 6 min read
Kort sagt
- Feature comparisons converge within a quarter. Architecture and contract terms do not.
- Every question here has an answer a vendor can show you rather than tell you.
- Group by the NIST AI RMF functions so the structure belongs to neither party.
- An accuracy percentage is a red flag, because it depends on an estate and a baseline that do not transfer.
- Ask about failure paths specifically. Uncertainty, tool errors and an unreachable approver are three different cases.
- For regulated buyers, where inference runs is a contract fact and belongs in the procurement record.
Why feature comparisons age badly here
A feature matrix in this category has a shelf life of roughly one quarter. Capabilities converge because they are largely downstream of the same model generation, and a differentiator described in a January briefing is table stakes by summer.
What does not converge is architecture and contract. Which actions run unattended, whether verdict reasoning is retained and exportable, where inference executes and under whose paper: those are decisions a vendor made and will not quietly change, and they are the ones you will still be living with at renewal.
There is also an asymmetry worth exploiting. A vendor can decline to answer a feature question by promising a roadmap. They cannot decline to show you a permission screen without telling you something.
Once this significant share is redacted, the top five targeted sectors in the EU include public administration (38.2%), transport (7.5%), digital infrastructure and services (4.8%), finance (4.5%) and manufacturing (2.9%).
What the European incident data frames the questions against
- 38.2 %
- of EU incidents hit public administration, the most targeted sector
- 7.5 %
- hit transport, which emerged as a high-value sector in the period
Källa: ENISA Threat Landscape 2025
Källa: ENISA Threat Landscape 2025
Govern: who decides what runs unattended
These three questions establish whether unattended action is designed or defaulted. They are first because every later answer is conditioned on them.
Ask which specific actions execute without approval, and ask to see that list in the product rather than in documentation. Ask who is permitted to change it, and whether the change appears in an audit trail. Then ask what happens when the model is uncertain, when a tool call fails, and when the approval channel is unreachable.
The last of those three separates most vendors. A system that acts when it cannot reach an approver has an approval feature rather than a boundary, and the difference only shows up at three in the morning.
An LLM-based system is often granted a degree of agency by its developer, the ability to call functions or interface with other systems via extensions.
| Question | A good answer | A concerning answer |
|---|---|---|
| Which actions run without approval? | A visible list in the product, per action type | It depends on your configuration |
| Who can change that, and is it logged? | A named role, and the change is in the audit trail | Any administrator, no separate record |
| What happens if the approver is unreachable? | It holds and escalates | It proceeds when confidence is high |
Map: what is in scope and what is deliberately not
Scope questions surface the assumptions a product makes about your estate, and they tend to reveal more than feature questions do.
Ask what data sources the system reads and what it does when a source is missing or stale, since a verdict formed on partial context is a different object from one formed on complete context. Ask what parts of the estate it will not act on under any configuration. Ask how a new detection type is onboarded, and by whom.
The second question is the one worth pressing. A vendor with a considered answer will name specific exclusions such as domain controllers or production databases. A vendor without one will say everything is configurable, which means the boundary is yours to invent.
Limit the extensions that LLM agents are allowed to call to only the minimum necessary.
| Designed scope | Configurable everything | |
|---|---|---|
| Names specific systems it will not touch | ✓ | ✕ |
| Behaviour on a stale data source is defined | ✓ | ✕ |
| Verdicts record which sources were unavailable | ✓ | ✕ |
| Onboarding a detection type has an owner | ✓ | ~ |
| Boundary is the customer's to invent | ✕ | ✓ |
Configurability is not a fault in itself. The question is whether the vendor has an opinion about the safe default, because that opinion is what you are buying alongside the software.
The vectors a scope answer should account for
Measure: how anyone knows the verdicts are right
This is where accuracy percentages appear, and where they should be declined. A number produced on someone else's estate with someone else's baseline does not describe what will happen on yours.
Ask instead whether the reasoning behind each verdict is retained and exportable, because a verdict you cannot re-examine cannot be reviewed and therefore cannot serve as evidence. Ask how disagreement is captured when an analyst overrules the system, and whether that feeds anything. Ask for two or three real cases where the system was wrong, and what happened next.
The last question is diagnostic in a way the others are not. A vendor who cannot produce an example of being wrong is either not looking or not telling, and both are worse than the errors would have been.
Excessive Agency enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM.
What the standards say to measure
- 4 functions
- in the NIST AI RMF, with measure covering verification and evidence
- 10 risks
- in the OWASP LLM Top 10 for 2025, of which misinformation applies directly to verdicts
Manage: what happens when it is wrong
The final three questions are about the day after, and they are the ones most likely to be answered honestly, because nobody rehearses them.
Ask what the rollback path is for each automated action class, and how long it takes. Ask what record exists of what the system did and, equally, of what it declined to do. Ask where inference runs, under whose contract, and whether that changes when a model provider is switched.
For an organisation inside NIS2 or DORA scope the last one is not preference but procurement. Article 21 makes supply chain security a risk management requirement in its own right, which means your detection vendor's own dependencies are inside your assessment whether or not you have looked at them.
| Question | A good answer | A concerning answer |
|---|---|---|
| What is the rollback path per action class? | Named per class, with a time estimate | Everything is reversible |
| What record exists of holds, not just actions? | Both, in one exportable trail | Actions are logged |
| Where does inference run, under whose contract? | A named region and a named provider | It varies by workload |
The regulatory questions that belong in the same conversation
For an organisation inside NIS2 scope, a detection vendor is a supplier inside a regulated estate, which means part of the evaluation is not about the product at all.
Article 21 makes supply chain security a risk management requirement in its own right, so the vendor's own dependencies are inside your assessment whether or not anyone has looked at them. That includes the model provider, the hosting arrangement, and anything the system calls out to during an investigation.
The sanction structure is worth having in mind while asking, not as a scare tactic but because it sets what proportionate diligence looks like. The ceiling is 2 percent of worldwide turnover for essential entities and 1.4 percent for important ones, which is the scale against which a procurement question about inference location is obviously reasonable.
Limit the functions that are implemented in LLM extensions to the minimum necessary.
The obligations the answers have to satisfy
- 24 hours
- to an early warning of a significant incident, which the verdict feeds
Källa: NIS2 Article 23
The two questions that are not about the product
Ask what would make the vendor tell you this is not a fit. A vendor with a real answer will name an estate size, a sector, a maturity level or a data constraint where the product does not earn its cost. A vendor without one is describing a product that fits everyone, which no product does.
Then ask who else in your sector runs this at your size, and ask to speak to them. This is the only reference question that produces something checkable, because the answer is either a name or a reason there is no name, and both are informative.
Neither question is adversarial. Both are faster than a proof of concept at establishing whether the vendor has thought about where their product stops.
Utilise human-in-the-loop control to require a human to approve high-impact actions before they are taken.
How to run the conversation
Send the twelve questions in advance. The goal is not to catch anyone out but to get past the demo script to the people who know the answers, and a vendor with good answers will welcome the chance to prepare them.
Record the answers in the procurement file with the date. Products in this category change quickly, and an answer that was true in March is not evidence about the system you are running in November.
Keep one thing in view while you read the answers. The category's own risk register is unusually blunt about where the hazard sits, and a vendor whose answers do not engage with that has either not read it or would rather you did not.
An LLM-based system is often granted a degree of agency by its developer, the ability to call functions or interface with other systems via extensions.
Questions
Common questions
- What should we ask an AI SOC vendor first?
- Which specific actions execute without approval, shown in the product rather than described in documentation, and what happens when the approval channel is unreachable.
- Why avoid accuracy percentages?
- Because the number depends on an estate, a baseline and a definition of correct, none of which transfer to your environment. Ask for reproducible cases instead.
- How do we compare two similar products?
- On architecture and contract rather than features. Features converge within a quarter; unattended action policy, verdict exportability and inference location do not.
- What is the single most diagnostic question?
- Ask for two or three real cases where the system was wrong and what happened next. A vendor who has none is either not measuring or not telling.
- Does the vendor's own supply chain matter?
- Yes, if you are in NIS2 scope. Article 21 makes supply chain security a risk management requirement in its own right, so your detection vendor's dependencies sit inside your assessment.
Primärkällor
Källor
Varje regulatoriskt påstående på den här sidan går att spåra till en av källorna nedan. Ingen av dem är en konsultblogg.
- NIST AI Risk Management Framework— The four functions this question set is built on
- OWASP LLM06:2025 Excessive Agency— Why the govern questions come first
- OWASP Top 10 for LLM Applications 2025— Misinformation as the verdict-level risk
- NIS2 Directive Article 21, EUR-Lex— Supply chain security as its own requirement
- Artificial Intelligence Act (EU) 2024/1689— Why the inference location belongs in procurement
Further reading
