The Direct Answer: Treat AI Suppliers as a Separate Risk Class

AI third-party risk management is the process of identifying, assessing, monitoring, and reducing risks created by external AI models, software agents, data services, and cloud infrastructure. The direct answer is that companies should govern these dependencies through a dedicated program rather than adding them as a few questions to an existing software-vendor review. A chatbot connected to customer records, an agent that executes transactions, and a conventional payroll application create different exposure despite all being purchased from third parties. Their failure modes, change rates, testing requirements, and contractual remedies also differ materially.

Also worth reading: What are the EU AI Act high-risk requirements for AI agents in 2026, and how do companies comply? · What is an enterprise AI governance control plane and how does it govern AI agents in production? · What is Amazon Bedrock AgentCore Gateway and how does it work as a security and policy control point for AI agents?

A workable program combines an inventory of AI dependencies, risk-tiered due diligence, contract controls, technical testing, and continuous reassessment after material updates. Reviews should occur before procurement, after deployment, and whenever a supplier changes its model, subprocessors, retention policy, or permitted uses. The objective is not to prohibit AI, because that can simply push experimentation into shadow IT and public tools. The objective is to make the risk visible and proportionate before the company connects a model to sensitive data or allows it to take consequential actions.

There is no universal certification that guarantees an AI vendor is safe. Companies should instead establish evidence thresholds, define decision rights, and record accepted residual risk. As a practical benchmark, a company might require enhanced review when a tool can access regulated data, represent 5% or more of a critical workflow, or act autonomously on financial or customer decisions. Those are governance examples rather than statutory thresholds, and organizations should calibrate them to their own size and exposure.

Why External AI Requires More Than Conventional Vendor Oversight

Traditional third-party risk programs often examine whether a supplier has security certifications, financial stability, insurance, and a credible incident-response process. Those checks remain useful, but they do not adequately describe how an AI system behaves. Model outputs can be wrong, biased, manipulated by prompt injection, or affected by changes made through a hosted application programming interface without any visible change to the supplier's contracts. IBM's guidance on governing third-party AI agents and research from RSM both emphasize that enterprises need controls specifically designed for AI-enabled services.

Agentic systems add another layer. A retrieval-augmented assistant may only recommend content, while an agent connected to enterprise tools may send email, modify records, or initiate purchases. The second system creates permission, delegation, and transaction risks that ordinary software reviews may overlook. Research associated with Bitsight has connected new forms of AI-enabled third-party risk to developments such as Claude Fable 5, illustrating why model changes deserve attention even when the vendor relationship itself remains unchanged.

The supply chain can also be deeper than the immediate contract suggests. A company may purchase an agent from one vendor, which then uses a foundation model from another, cloud services from a third, and data enrichment from a fourth. Contractual restrictions on the immediate supplier may not automatically reach every downstream provider. Company risk teams should ask which components are included, whether the supplier identifies material subprocessors, and what notice the customer receives before those providers change.

Regulatory expectations are making this operational issue harder to avoid. Under the EU AI Act, Article 26 places important duties on deployers of certain high-risk AI systems, including understanding intended use, monitoring operation, and keeping automatically generated logs. Most obligations under the Act became applicable on August 2, 2026, while obligations for certain embedded high-risk systems extend to August 2, 2027. The exact treatment depends on the system's role and context, so legal teams should not reduce compliance to a single date or checklist.

Build a Risk Taxonomy for Models, Agents, and AI Services

The first control is a clear taxonomy. A mature inventory should distinguish hosted foundation models, embedded AI features, custom-built systems using external models, autonomous agents, AI-enabled SaaS, and infrastructure or data providers. It should also capture whether the supplier retains prompts, outputs, embeddings, training records, or tool-call logs, and whether those records contain personal, regulated, or commercially sensitive information. A spreadsheet can be sufficient initially, but it needs a stable schema, accountable owner, and defined review frequency.

Classify systems using impact rather than vendor reputation. A low-impact internal drafting tool with public information can receive standard controls, while a claims-processing tool using health data or a pricing agent with authority to change customer contracts needs substantially deeper scrutiny. Useful criteria include autonomy, data sensitivity, decision scale, regulatory exposure, recoverability, and the difficulty of detecting errors. A system handling 10,000 customer interactions per day presents a larger monitoring burden even if each individual action appears routine.

Set evidence thresholds that correspond to each tier. Standard-tier providers might require security documentation, breach history, data-location details, and a subprocessor list. Higher tiers may also require model cards, system cards, evaluation results, red-team findings, prompt-injection testing, bias assessments, and an incident-notification deadline. Contracts for high-impact systems should include audit rights, change notification, cooperation with regulators, and prompt commitments for remediation within an agreed period.

Residual risk should be approved by a named business owner, not merely by procurement. Security, legal, privacy, data science, and operations may each hold veto rights in defined circumstances, but a separate committee should not become a ceremonial rubber stamp. Companies operating in multiple markets should maintain a common control baseline and add local legal requirements. A global standard is useful only if it prevents the weaker policy in one jurisdiction from becoming the de facto policy everywhere.

A Practical Control Cycle for AI Suppliers

Begin with discovery because most organizations underestimate how many AI services are already in use. Security teams can inventory SaaS connections, cloud API activity, browser extensions, code libraries, and employee accounts used with external assistants. For a mid-sized company, a sensible first target is to identify at least 90% of known AI services within 60 days, then investigate discrepancies rather than claiming complete coverage. Unapproved tools should be evaluated, restricted, or retired according to risk, with training aimed at explaining why an uncontrolled upload can expose confidential information.

Next, establish an intake process for proposed purchases and pilots. The requester should describe the intended purpose, data sources, user population, model provider, hosting location, logging behavior, autonomous permissions, and expected business value. Procurement can use a short questionnaire for low-risk tools and route higher-impact systems to a multidisciplinary review. Pilot contracts should avoid automatic production access and should contain a defined exit date or conversion decision, commonly 30 to 90 days, so experiments do not quietly become permanent dependencies.

Before production approval, run tests that match the intended use. This can include adversarial prompts, unauthorized tool-call attempts, sensitive-data leakage checks, role-based access tests, and reviews of whether the model follows documented escalation paths. For consequential decisions, compare error rates against a human or deterministic baseline and define what happens when results fall below an approved threshold. Example thresholds might include a false-positive rate above 3%, an unresolved critical vulnerability, or missing audit logs for more than 24 hours.

After launch, monitor both supplier events and actual system behavior. The program should track service availability, incident history, model changes, subprocessor changes, cost anomalies, and internal misuse. RSM notes that vendors can alter their risk profile between periodic reviews, which is why annual paperwork alone is inadequate. Continuous signals matter, but continuous model evaluation can be expensive, so organizations should increase testing frequency for high-impact systems and use sampling for lower-risk use cases.

Comparing Governance Approaches and Available Tool Types

Companies have four broad options: manual governance, conventional GRC platforms with AI extensions, AI-specific assurance platforms, and a combined operating model. None is automatically superior. Manual review can work for a small company with few dependencies, while a heavily automated program without accountable owners can produce a polished dashboard that does not change decisions. The best choice depends on portfolio size, technical maturity, and the consequences of failure.

FeatureConventional GRC platformAI-specific assurance platformManual program with cloud controls
Core strengthSupplier inventory, contracts, audits, and risk registersAI model evaluations, evidence collection, drift monitoring, and agent controlsFlexible judgment with low initial software cost
Setup timeOften weeks to a few months for a new programCommonly several months for a mature deploymentDays to weeks, but organized within the existing team
Ongoing effortModerate; questionnaires and document reviewsModerate to high where technical monitoring is enabledHigh per supplier and inconsistent at scale
Best fitRegulated enterprises with an existing GRC functionCompanies operating many models or agents in productionSmall teams with limited, low-risk AI usage
Main weaknessAI behavior may be reduced to generic vendor questionsCost and integration can exceed the value for simple toolsWeak consistency, limited alerts, and poor auditability
Illustrative annual cost$20,000-$150,000+$50,000-$300,000+$5,000-$30,000 in staff time and basic tooling
A combined model is usually the most credible starting point. Organizations can use their GRC system as the system of record, existing security tools to collect technical evidence, and specialist services for evaluations that internal teams cannot perform. The figures in the table are planning ranges rather than vendor quotes, and prices vary by users, modules, integrations, and implementation scope. Buyers should calculate the total cost of ownership, including data preparation, evaluation labor, model updates, and time spent by subject-matter experts.

No platform can replace contract interpretation, testing design, or risk acceptance. AI assurance tools can detect anomalous outputs or configuration changes, but a metric becomes meaningful only when an organization knows which errors matter. Similarly, ISO-oriented or sector-specific controls can provide structure, yet they do not automatically establish that a model performs acceptably in a particular workflow. Technology and governance should therefore reinforce rather than substitute for each other.

Common Mistakes That Produce False Assurance

The most frequent mistake is assuming that a security certification proves the AI system is safe. Certifications demonstrate that defined controls were assessed, not that every model output is accurate, unbiased, secure against prompt injection, or appropriate for a new use case. Another error is reviewing only the application vendor while ignoring the underlying model, cloud region, plugins, retrieval sources, and external tools. The visible interface may look familiar while several opaque dependencies operate behind it.

Companies also fail when they treat all model changes as equivalent. A minor display-language update does not warrant the same review as a new training method, expanded tool permissions, or modified data-retention policy. Instead of demanding approval for every release, suppliers and customers should agree on materiality thresholds, such as changes that affect safety evaluations, data use, legal basis, autonomy, or a regulator-relevant output. The contract should specify advance notice where commercially possible, even though no clause can guarantee real-time notice for an unannounced provider change.

A third mistake is measuring activity instead of outcomes. A security team might report 40 supplier reviews and 100 completed questionnaires, yet know nothing about failed evaluations, unresolved high-risk findings, or incidents caused by excessive permissions. Useful program measures include percentage of production AI services inventoried, critical findings closed within target time, percentage receiving monitoring, and median time from supplier change to internal reassessment. Metrics should be audited for gaming, because targets that reward speed can encourage premature closure.

Finally, organizations often delay governance until after a procurement dispute or public incident. Retroactive reviews are useful, but they cannot restore deleted evidence or reverse data already exposed. A limited 30-day stabilization effort can create an inventory, suspend unauthorized data uploads, and identify systems with autonomous permissions, but it should not be called mature risk management. Residual uncertainty should be recorded with an owner and an expiration date so temporary acceptance does not become permanent neglect.

When to Escalate, Restrict, or Exit a Supplier

Speed is not inherently better in third-party AI oversight, but delay can be risky. A new material model, evidence of a critical vulnerability, unauthorized retention of training data, or an unexplained change in output quality should trigger accelerated review. Regulated applications, decisions affecting employment, credit, insurance, healthcare, or essential services deserve senior attention because errors may affect people who cannot easily challenge an automated result. Companies should also escalate when a vendor refuses required evidence, lacks a viable incident-response process, or cannot support the required data location.

Restriction is often more effective than immediate termination. The company can disable external-model training, remove production write access, require human approval, restrict the tool to low-risk datasets, or move processing into an approved environment. For agents, a useful default is the principle of least privilege: read access for analysis, limited write access for a defined task, and no ability to transfer funds or make contractual commitments without human confirmation. Companies should test whether those restrictions survive retries, plugin changes, and account compromise rather than assuming they do.

Exit planning should begin before an emergency. Contracts should support data export, deletion certification, model-unlearning options where feasible, and transition assistance. Organizations should know whether switching providers requires rebuilding prompts, retrieval indexes, evaluation sets, and integrations, since the apparent cost of changing models can be much higher than the difference in token prices. A vendor that offers acceptable pricing but cannot provide logs or deletion evidence may create a concentration risk that is not visible in a benchmark chart.

Boards and executives should receive concise indicators rather than raw technical dashboards. Quarterly reporting might cover the number of critical AI suppliers, percentage with current due diligence, high-risk systems in production without approval, unresolved red-team findings, and material model or subprocessor changes. As of September 24, 2026, the EU AI Act's August 2 application milestone makes legal classification an immediate topic for companies serving the European market, although deployment details determine which rules apply.

Cost, Pricing, and Deciding What Not to Buy

There is no regulated market price for effective AI third-party risk management. Cost depends on the number of vendors, whether the company buys software or advisory services, and how much testing it performs internally. A small business might spend $5,000 to $30,000 annually on staff time and basic cloud security controls, while an enterprise GRC implementation can run from $20,000 to more than $150,000 annually. AI-specific assurance deployments may range from $50,000 to $300,000 or higher when they include integrations, continuous evaluations, and specialist workflows.

The relevant calculation is loss exposure and avoided rework, not merely license count. A tool used by 20 employees for public drafting may not justify a six-figure platform, even if it is innovative. A decision agent processing 500,000 claims or transactions annually may justify stronger monitoring because a defect can be repeated at scale. Many providers also charge separately for modules, evaluations, storage, API usage, and implementation, so short-term per-user prices can understate total cost.

Companies should avoid buying a large suite before specifying a control problem. Excessively broad platforms can add questionnaires without improving model testing, while narrowly selected tools may not integrate with existing GRC records. A pilot should have success criteria covering false-positive rates, evidence quality, integration time, and the time required to reassess a material vendor change. If the pilot cannot save meaningful staff hours or improve risk decisions within three to six months, the purchase should be reconsidered rather than expanded to satisfy a procurement deadline.

The Operating Model That Survives Continuous Change

The most durable program is a repeating cycle rather than a single project: discover, classify, contract, test, approve, monitor, and reassess. Ownership should be divided clearly. Procurement manages commercial records, security evaluates technical controls, legal translates requirements into clauses, data science or engineering conducts model tests, and the business owner accepts residual risk in the context of an actual workflow. For lower-risk systems, a lightweight path should remain available so controls do not unintentionally encourage unmanaged purchases.

Documentation should show not only what the system is but what changed. A review record ought to include the model version, applicable evaluation results, approved data types, tool permissions, unresolved findings, and next review date. Where suppliers offer version identifiers or change logs, retaining them makes post-incident analysis possible. Company evidence should also distinguish supplier assertions from independently verified results, because a report can be authentic without being complete.

AI third-party risk management should eventually become routine enterprise capability, but that does not mean every organization needs the same tooling. Some will use a conventional GRC platform, others will adopt AI-specific assurance services, and smaller firms may rely on managed security providers plus carefully restricted tools. The decisive test is whether leadership can answer five questions at any time: Which external AI systems matter most, what data can they access, what can they do, how will failures be detected, and who has accepted the remaining risk? If those answers are current and supported by evidence, the program is doing its job.