What Is an AI Third-Party Risk Review?
An AI third-party risk review is a structured assessment of an external model, API, software platform, data service, cloud infrastructure, or agent that can affect an enterprise’s operations, customers, or regulated decisions. It examines more than security: teams also test data handling, model behavior, contractual protections, regulatory exposure, operational resilience, and the vendor’s own dependencies. For example, a review of a coding assistant may cover prompt retention, training permissions, generated-code quality, license provenance, and whether developers can disable the service during an outage. A review of an agent with access to email or payment systems must additionally address tool permissions, action limits, approval gates, and audit logs.
Also worth reading: How Do Modern Enterprises Implement Effective Agentic AI Risk Management in 2026? · What Is AI Runtime Control Architecture and How Should Enterprises Adopt It in 2026? · How Can Enterprises Control AI Token Costs Without Slowing Agent Development in 2026?
The central answer is that enterprises should operate a continuous review process with scheduled re-certification and event-triggered reassessment, rather than treating procurement as a one-time questionnaire. AI vendors can change models, subcontractors, data practices, safety controls, or business ownership after approval. A vendor that passed a review in January may present a different risk profile six months later simply because a model update altered autonomous behavior or a new integration exposed customer data. As of September 25, 2026, that changeability makes AI third-party risk reviews a living governance activity, not static vendor documentation.
A defensible review produces four concrete outputs: a current inventory of AI dependencies, a documented risk rating, assigned control owners, and an expiration or reassessment date. It should also record which claims were independently tested and which rely only on vendor assurances. The goal is not to certify that a model is “safe.” That claim is usually impossible to prove. The goal is to define how the enterprise will use the system, what failure would look like, and what actions will be taken if the system exceeds its intended operating boundaries.
Why Traditional Vendor Assessments Are Insufficient
Conventional third-party assessments usually emphasize administrative safeguards, encryption, disaster-recovery plans, and whether a vendor has a security program. Those controls remain necessary, but they do not fully describe risks introduced by probabilistic and agentic AI. A model can comply with a security standard while producing discriminatory recommendations, fabricated citations, unsafe code, or unauthorized actions. It can also be technically available yet commercially unusable because a provider changes pricing, usage limits, or acceptable-use policies without notice.
The regulatory environment raises the stakes. Article 26 of the EU AI Act addresses deployers’ obligations in relation to high-risk AI systems, while EU guidance on classifying high-risk systems has continued to evolve. For general-purpose AI models, governance milestones have included August 2, 2025, for applicable provider and model obligations and August 2, 2026, for enforcement provisions, with additional high-risk-system requirements following later depending on the system category. Organizations using AI vendors in EU operations should therefore ask not only whether a provider has a compliance page, but whether its documentation identifies the relevant system role, intended purpose, and deployment conditions.
Agentic systems create another gap. A chatbot that drafts text has different exposure from an agent that can send emails, modify databases, or execute transactions. Human presence does not automatically make an agent low risk; it may simply mean an employee can approve a dangerous action without understanding it. Reviews should classify these systems by the authority and consequences of their actions. A useful internal threshold is to treat autonomous systems with financial, employment, healthcare, legal, or production-write access as high impact until testing demonstrates stronger controls. This is a risk-management recommendation, not a universal legal safe harbor.
What Should Each AI Risk Review Examine?
Each review should begin with a precise use-case description rather than a generic evaluation of the vendor. Record the model version, prompting approach, connected data, user population, autonomous capabilities, and business purpose. “Uses OpenAI for customer support” is inadequate. “Uses a model to draft replies, retrieves policy documents, and may issue refunds below $50 after a human approval” permits a materially better review. Version identifiers matter because behavior can change even when the product name and API endpoint remain the same.
Technical testing should then examine the system against realistic tasks. Teams can use a fixed test set, generally 100 to 500 representative cases for an early production review, and record task completion, false-positive rates, hallucination rates, refusal behavior, and failure severity. For coding systems, the set should include insecure patterns, dependency vulnerabilities, and unfamiliar languages. For agents, evaluators should attempt to trigger excessive permissions, prompt injection, data exfiltration, repeated tool calls, and approval bypass. These are internal engineering thresholds, not regulatory mandates, and they should be adjusted to the cost of error.
The review must also cover contracts and concentration risk. Contract terms should address data ownership, retention, model training, breach notification, subcontractors, audit evidence, service credits, regulatory cooperation, exit assistance, and deletion of customer data. Resilience analysis should identify whether production depends on one model provider, one cloud region, or one identity platform. A reasonable starting policy is to test recovery procedures at least annually and after any material architecture change, while maintaining a documented fallback for high-impact workloads. Vendor assurances should be mapped to named internal controls so that accountability does not disappear into a shared-responsibility statement.
How Often Should Reviews Happen?
Most enterprise AI applications need at least an annual baseline review, but calendar frequency alone is inadequate. High-impact deployments warrant quarterly control checks and event-driven reassessment, particularly when a vendor changes its model, ownership, training data policy, subprocessor list, incident history, or agent permissions. A trigger can also come from an internal discovery that the tool has expanded beyond its approved purpose. Waiting for the next annual questionnaire could leave newly introduced exposure unexamined for months.
A workable cadence assigns different depths by tier. Tier one can cover low-impact productivity tools with limited data and no external action; these may receive an annual questionnaire, automated configuration checks, and lightweight user reporting. Tier two includes tools that process confidential information or influence operational decisions, requiring annual testing plus semiannual control reviews. Tier three covers agents with production access, sensitive personal data, or legally significant decisions, making quarterly testing, continuous telemetry, and immediate reassessment after material events appropriate. Organizations should set their own tiers because a static tool-processing assessment does not fit every deployment.
Change management should have explicit notice periods where possible. A practical commercial target is 30 days’ advance notice of a material model or subprocessor change, with an emergency exception for security or legal reasons. Contracts may not always achieve that target, so monitoring remains necessary. Vendors may announce product changes through release notes, developer documentation, or customer portals rather than formal governance notices. A mature program samples those sources monthly and records whether the change altered a previously tested behavior.
Review frequency should also reflect evidence quality. If a vendor supplies reproducible audit artifacts, current independent assurance reports, and stable interfaces, an enterprise may rationally reduce manual sampling. If evidence is marketing-oriented, system behavior changes frequently, or the vendor resists logging requirements, the review burden should increase. Frequency therefore measures assurance needs, not simply organizational convenience.
Comparing Manual, Automated, and Consulted Reviews
Enterprises commonly combine three review methods rather than selecting only one. The best operating model usually uses questionnaires and contract analysis for legal and governance evidence, automated testing for repeatable technical checks, and specialist review for novel or high-impact systems. Buying a tool does not transfer accountability, while relying entirely on consultants can produce a polished report that is never connected to production controls.
| Feature | Internal Continuous Review | Automated Risk Platform | Specialist-Assisted Review |
|---|---|---|---|
| Best use | Routine inventory and control monitoring | Repeated configuration, policy, and behavior checks | Novel models, agents, and regulated use cases |
| Evidence | Logs, owner attestations, internal test results | Policy scanning, API checks, benchmark results, vendor monitoring | Independent testing, architecture analysis, expert interpretation |
| Typical cadence | Continuous monitoring; annual baseline | Weekly or monthly checks | At onboarding and material change |
| Cost structure | Staff time, internal engineering capacity | Subscription, implementation, integrations, usage | Project fees plus follow-up testing |
| Main limitation | Can be inconsistent or dependent on internal expertise | May create false confidence and miss context | Expensive and difficult to sustain if used for every review |
| Accountability | Enterprise retains ownership | Enterprise retains ownership | Advisory findings require internal control owners |
A Practical Review Process for 2026
The first step is to create an inventory that records the business owner, technical owner, vendor, model, data categories, permissions, affected populations, and review tier. Missing ownership is itself a governance finding. An unknown shadow AI tool should not remain outside oversight merely because it was adopted through a browser extension or employee account. The inventory should include services acquired through procurement, information security, legal, business units, developers, and acquisition targets.
The next step is a risk classification based on impact, autonomy, data sensitivity, reversibility, and regulatory relevance. A five-level scheme is workable, but thresholds should be explicit. A tool that only rewrites internal non-sensitive text might score low, while an agent authorized to change customer billing or access production credentials should score high even if the model provider is reputable. Scoring should produce obligations such as annual testing, human approval, restricted retention, or prohibition from sensitive data; a numerical score alone adds little value.
Evaluation and remediation follow. Test against representative tasks, document unacceptable results, assign owners, and set a deadline. A 30-day target is reasonable for disabling a clearly unauthorized integration or exposing credentials, while a 90-day target can be appropriate for engineering changes when service continuity must be preserved. Exceptions should have compensating controls, named approvers, and expiration dates. The process should end with a monitored approval rather than a PDF filed in a procurement system.
Common Mistakes and Weak Assumptions
A frequent mistake is accepting a vendor’s “enterprise-ready” label as an assessment. Enterprise readiness may describe availability, support, or administrative features, not suitability for a particular regulated decision. Another error is evaluating a vendor’s latest model while production still uses an older version. Reviews must identify the exact model and configuration, because model aliases can be updated and may not identify a stable underlying version.
Organizations also confuse an AI ethics statement with operational control. Public commitments about safety do not prove that a customer deployment has appropriate tests, permissions, monitoring, or incident procedures. Similarly, an impressive security score can obscure model-specific weaknesses such as prompt injection, harmful automation, or fabricated regulatory citations. These are distinct control domains and should be recorded separately.
The most damaging mistake is assuming that a human in the loop eliminates risk. Approval becomes weak when reviewers cannot see enough evidence, face excessive alert volume, or routinely accept the system’s recommendation. High-impact deployments should therefore measure review time, rejection rates, overrides, and whether approval changes outcomes. If 99% of agent actions are approved, that fact alone may indicate rubber-stamping or poor interface design rather than effective oversight.
When to Act and What to Do Next
Immediate action is warranted when an unapproved tool handles regulated data, an agent has production-write access, credentials are shared without scoping, or a vendor material change has not been assessed. These situations should move into an expedited review with interim restrictions where necessary. Temporary measures can include revoking unused permissions, disabling sensitive-data connections, requiring human approval, limiting actions to read-only mode, and increasing logging. The organization should not wait for a full procurement cycle before containing a known exposure.
For other deployments, the next step is a 90-day program establishing inventory, tiering, baseline testing, ownership, and reassessment triggers. Day one should identify the business units using AI and the systems capable of taking external action. By day 30, teams should have a preliminary inventory and ranked shortlist. By day 60, they can complete testing on the highest-impact systems and remediate obvious control failures. By day 90, the program should have documented approval criteria, exceptions, metrics, and an executive reporting cadence.
Leadership should measure more than the number of completed reviews. Useful indicators include percentage of AI assets with named owners, percentage of high-impact systems tested within the last 12 months, time to contain a critical finding, number of unapproved production connections, and proportion of material vendor changes reassessed. As of September 25, 2026, enterprises that treat these metrics as operating data will be better prepared than those that merely collect more questionnaires. The defensible standard is not perfect certainty; it is a repeatable process that notices material change, tests real behavior, and acts when the evidence no longer supports the approved use.