The AI governance conversation inside most enterprises has matured enough to produce principles documents, internal review boards, and responsible AI policies. What it has not always produced is a procurement process capable of testing whether a vendor's ethics claims are grounded in evidence or assembled for the pitch deck. That gap matters now more than it did two years ago, because the regulatory environment has moved. The EU AI Act, NIST AI RMF, and ISO/IEC 42001 are no longer aspirational frameworks. They carry roles, records, oversight obligations, and assessment requirements that transfer liability downstream to the enterprises that deploy non-compliant systems.
Companion piece to our broader work on AI accountability architecture. See Who Owns the AI Mistake? Building an Accountability Architecture Before Regulators Force Your Hand for a practical guide to role definitions, incident ownership models, and embedding accountability into the AI development lifecycle.
What the Frameworks Actually Require From Suppliers
The EU AI Act's risk-tiered structure places specific conformity obligations on providers of high-risk AI systems, including requirements for technical documentation, human oversight mechanisms, and logging of system behaviour. These are not voluntary disclosures. They are conditions that must be satisfied before a system can be placed on the market in the EU, and they flow directly into what you should be asking any vendor operating in scope.
NIST AI RMF organises governance into four functions: Govern, Map, Measure, and Manage. The Govern function in particular requires that AI risk tolerances are documented and that accountability is assigned to named roles, not distributed across an organisation in ways that make ownership untraceable. ISO/IEC 42001 adds a management system layer, requiring that an organisation's AI objectives are documented, auditable, and subject to continual improvement review.
The practical implication is that a vendor claiming compliance with any of these frameworks should be able to produce specific artefacts: a risk register, a documented oversight process, named accountability owners, and records of assessment against defined criteria. If they cannot, the compliance claim is a positioning statement, not a governance posture.
The Difference Between Bounded Claims and Aspirational Positioning
Chawla and Benanti draw a distinction that is directly useful in procurement contexts: the difference between evidence-bounded deployment, which limits claims to what has actually been evaluated, and measurement-bounded governance, which records constraints that favourable evidence cannot override (Chawla and Benanti, arXiv 2026). A vendor who tells you their system is fair has made an aspirational claim. A vendor who tells you their system was evaluated for demographic parity across defined subgroups using a specified dataset, and that those results are available for review, has made a bounded claim.
The distinction matters commercially because aspirational claims do not transfer accountability. If a vendor's system produces discriminatory outputs in your environment, a fairness principle on their website does not constitute a warranty. A documented evaluation against a defined criterion, scoped to a specific deployment context, at least establishes what was and was not tested.
When evaluating vendors, ask not only what claims they make but what evidence bounds those claims. The scope of the evaluation, the dataset used, the subgroups analysed, and the metrics applied are all part of the answer. Absence of that specificity is itself informative.
Where Institutional Failure Compounds the Risk
One of the more uncomfortable observations in recent responsible AI research is that AI systems do not arrive into neutral institutional environments. Chawla and Benanti argue that once deployed, AI becomes an intervention in existing institutional conditions, and that it can repair, compound, substitute for, or conceal the failures it encounters (Chawla and Benanti, arXiv 2026). For CTOs, this means that deploying a vendor's AI system into a broken process does not fix the process. It encodes the failure at scale and makes it harder to detect.
This has a direct bearing on procurement. Before evaluating a vendor's system, it is worth assessing the institutional process the system will enter. If that process lacks clear ownership, consistent data quality, or defined escalation paths, the AI layer will inherit those weaknesses. Vendor governance documentation cannot substitute for the organisational readiness assessment that should precede deployment.
The procurement question is therefore not only whether the vendor's system is compliant, but whether your organisation's deployment context is prepared to operate it in a way that preserves the governance properties the vendor has documented.
How to Structure the Vendor Evaluation
A governance-literate vendor evaluation should cover four areas.
Documentation and Auditability
Ask for the technical documentation required under the EU AI Act for any system in scope, the risk register maintained under NIST AI RMF, and the management system records required under ISO/IEC 42001. These are not bespoke requests. Any vendor claiming compliance should treat them as routine.
Accountability Assignment
Ask who holds named accountability for the system's outputs, what the escalation path is when the system produces an unexpected result, and how incidents are logged and reviewed. A vendor who cannot answer these questions with specific role titles and documented processes has not operationalised their governance framework.
Evaluation Scope and Limitations
Ask what the system was evaluated on, what it was not evaluated on, and what deployment contexts fall outside the scope of the vendor's documented testing. A vendor willing to state the limits of their evaluation is demonstrating the kind of epistemic honesty that correlates with genuine governance maturity.
Contractual Allocation of Liability
Ensure that the contract reflects the governance representations made during evaluation. If a vendor claims their system meets specific bias or safety criteria, those criteria should appear in the contract with defined remediation obligations if they are not met in production.
Reading the Signals in Vendor Responses
Vendors who respond to governance questions with polished one-pagers describing their ethical principles, without any reference to specific artefacts, named roles, or evaluation records, are demonstrating that their ethics programme exists at the communications layer rather than the operational layer. That is not necessarily bad faith. It may reflect an organisation that has invested in positioning before investing in process. But it is a signal that the governance claims will not hold up under regulatory scrutiny, and that the accountability gap will fall to you as the deploying enterprise.
Vendors who respond with specific documentation, who can name the person responsible for AI risk management, and who are willing to discuss the limitations of their evaluations alongside their results, are demonstrating a governance posture that has moved from principles to protocols. That is the standard the regulatory environment now expects, and it is the standard your procurement process should reflect.
Where Vector Labs Fits
We build AI systems with governance and certification requirements embedded from the design stage, not retrofitted at deployment. In our cardiovascular AI study, we structured validation from the outset to meet medical device software standards, incorporating subgroup analysis and comprehensive regulatory documentation, which resulted in Class 2A medical device certification delivered within the product launch timeline. If you are evaluating an AI vendor or preparing your own AI systems for regulatory scrutiny, contact us at vector-labs.ai/contacts.
FAQs
For any system classified as high-risk under the EU AI Act, you should request the technical documentation specified in Article 11, which covers system design, development methodology, training data governance, and performance metrics. You should also ask for the logs of system operation required under Article 12, and evidence of the human oversight mechanisms required under Article 14. If a vendor cannot produce these, their compliance claim is not operationally grounded.
NIST AI RMF is a voluntary framework that organises AI risk management into four functions: Govern, Map, Measure, and Manage. It does not carry regulatory force in most jurisdictions, but it provides a structured basis for evaluating whether a vendor has assigned accountability and documented risk tolerances. ISO/IEC 42001 is a certifiable management system standard, meaning a vendor can obtain third-party certification against it. Certification does not guarantee good outcomes, but it does mean an external auditor has reviewed whether the management system meets the standard's requirements, which is a stronger signal than self-attestation.
Any governance representation made during the evaluation process should be translated into a contractual obligation with defined performance criteria and remediation terms. If a vendor claims their system meets specific fairness or safety thresholds, those thresholds should be specified in the contract, along with the evaluation methodology used to establish them. The contract should also define what happens if the system fails to meet those criteria in production, including notification obligations, remediation timelines, and termination rights.
If an AI system is deployed into a process that already has inconsistent data quality, unclear ownership, or undocumented decision criteria, the system will operate on those conditions and produce outputs shaped by them. Because AI outputs can appear authoritative and are often produced at volume, the underlying process failures become harder to detect and challenge. The practical implication is that an organisational readiness assessment should precede any vendor deployment, covering data quality, process ownership, and escalation paths, not just the technical specification of the system itself.
The most reliable signal is specificity. A vendor with operational governance can tell you who holds named accountability for AI risk, what their system was evaluated on and what it was not, and what artefacts exist to support those claims. A vendor whose governance exists primarily at the communications layer will respond to these questions with principles statements, values documents, or high-level process descriptions that do not resolve to specific records or named owners. Asking for the limitations of their evaluation, not just the results, is a particularly useful test: a vendor willing to describe what their system has not been tested on is demonstrating the epistemic honesty that genuine governance maturity requires.

