What AI Vendor Due Diligence Actually Requires

AI vendor due diligence is the structured process of deciding whether a third-party AI product, model, data source, or service provider is suitable for a defined business use. It is not a questionnaire marathon, a model benchmark review, or a single security questionnaire completed by sales. A defensible process connects the intended use to evidence about data rights, model behavior, security, operational resilience, human oversight, regulatory exposure, and the provider’s capacity to support the system over time.

Also worth reading: How Can Organizations Implement an Enterprise Agent Governance Blueprint to Control Autonomous AI Systems? · How Should Organizations Structure an Enterprise AI Architecture Roadmap for 2026 and Beyond? · What is the definitive agentic AI security posture for enterprise organizations in 2026?

As of September 24, 2026, the central issue is no longer whether AI deserves supplier oversight. Regulators already expect institutions to manage vendors as technology, data, and operational dependencies rather than as ordinary software purchases. The EU AI Act’s obligations concerning governance and general-purpose AI have applied since August 2, 2025, while most of its remaining provisions are scheduled to begin on August 2, 2026, subject to any late changes and sector-specific implementation. DORA has applied since January 17, 2025, making ICT third-party risk an operational concern for covered financial entities.

The correct diligence depth depends on what the AI will do. A tool that summarizes internal meeting notes presents a different risk profile from a system that rejects loan applications, recommends medical treatment, prices insurance, or makes hiring decisions. Low-impact productivity software may justify a focused review, while a high-impact system often needs model documentation, representative testing, data provenance checks, human-override rules, and an independently testable control environment. The output is a documented decision with conditions, not a universal claim that the vendor is “safe.”

Why Conventional Procurement Falls Short

Traditional supplier reviews tend to measure what is easy to count: certifications, uptime claims, contract terms, encryption features, and the number of supported languages. Those facts matter, but they do not establish whether a model performs acceptably in your organization. A vendor can hold ISO 27001 certification while still producing biased recommendations, training on data it lacks the right to use, or exposing confidential prompts to a human support process it never described in its sales deck.

AI changes the supply chain because performance cannot be separated from the customer’s data, operating procedures, and decisions about human review. The same model may behave differently after fine-tuning, retrieval settings, system prompts, or access permissions are changed. A model card written before deployment may describe the provider’s laboratory configuration rather than the production service being purchased. Due diligence must therefore examine the actual product, relevant documentation, contractual commitments, and intended deployment rather than treating the provider’s brand as a proxy for reliability.

There is also a hidden-provider problem. An apparent single supplier may depend on a cloud host, a foundation-model developer, an external data provider, a logging service, and subcontractors that process information elsewhere. Existing due diligence on unstructured documents can also miss liabilities that appear only in attached reports, spreadsheets, screenshots, or correspondence. As of 2026, many governance frameworks are not keeping pace with adoption, making internal evidence collection more important rather than less. Investors, audit committees, customers, and insurers may examine these dependencies even when the original contract names only one vendor.

A Risk-Based Diligence Method

Start with a written use-case statement that identifies the decision being supported, the people affected, the data involved, the expected business owner, and the consequences of error. Classify the system using consequences and regulatory status rather than the vendor’s marketing language. A useful internal threshold is to give enhanced review to any system that makes or materially supports decisions about credit, employment, housing, insurance, health care, education, payments, or public benefits. Systems that only draft content can still require review when they handle regulated, personal, or confidential information.

Next, obtain evidence that is specific to the purchased service. Request a system diagram, hosting locations, subprocessors, retention settings, deletion guarantees, incident-notification terms, encryption standards, disaster-recovery arrangements, and the identity of any foundation-model provider. For model-based products, ask for evaluation methods, known limitations, performance across relevant demographic or language groups, change-notification practices, and an explanation of whether customer prompts or outputs train shared models. A promise that customer data is “not used for training” is more useful when paired with contractual language, technical restrictions, and audit rights.

A practical evidence threshold is 80% coverage before pilot approval for a high-impact deployment, with every uncovered item assigned an owner and deadline. That is an internal management target, not a regulatory safe harbor. Less consequential pilots can use narrower thresholds, but they should still address data rights, security, incident response, and exitability. The objective is to create an evidence trail showing why the vendor was selected, which risks remain, who accepts them, and what would cause the deployment to stop.

Evaluation AreaEnterprise-Grade AI VendorLower-Risk Pilot or Standalone Tool
Intended useMay support regulated or material decisionsDrafting, search, summarization, or personal productivity
Evidence expectedRepresentative testing, data lineage, model documentation, human controls, audit rightsSecurity review, data-use terms, owner approval, limited test data
MonitoringOngoing performance, drift, incidents, and vendor-change reviewPeriodic spot checks and access review
Contract postureStrong audit, notification, IP, indemnity, and exit termsStandard terms may suffice if exposure is contained
Typical review cycleQuarterly for material changes; at least annuallyBefore renewal or scope expansion
## Testing What the Vendor Does Not Volunteer

Desk research should be followed by tests using data that resemble the real workload without exposing unnecessary personal or confidential information. For retrieval systems, include realistic documents and questions so reviewers can observe whether citations are present, accurate, and relevant. For classification or prediction tools, define error costs separately. A 2% false-negative rate may be tolerable in an internal search feature but unacceptable in a system used to flag suspected financial crime.

Security testing should cover identity and access management, tenant separation, API authentication, secrets handling, logging, vulnerability management, and deletion. Ask whether administrators can turn off logging when troubleshooting and whether support staff can access prompts, files, or evaluation results. Define a maximum support-access window where the vendor cannot provide evidence, such as zero for privileged access without a documented business need. Verify encryption in transit and at rest, but do not stop there: a correctly encrypted database can still contain improperly collected data or weak authorization rules.

Performance testing should include edge cases and groups that may perform differently from the vendor’s headline benchmarks. For general-purpose AI systems, NIST’s AI Risk Management Framework 1.0, published in January 2023, and its Generative AI Profile, released as NIST AI 600-1 in July 2024, provide useful structure even for buyers outside the federal government. Documentation based on a generic model card remains a starting point. Buyers need production-relevant acceptance criteria, monitoring thresholds, a process for reporting material model changes, and a right to suspend use when those thresholds are breached.

Comparing Vendors, Build, and Open Models

AI vendor due diligence is sometimes framed as a choice between buying a product and building a system. That comparison is too narrow because most products combine third-party models, cloud services, proprietary data, and internal automation. A buy decision may still require substantial integration, governance, and evaluation work, while building a system can preserve control without eliminating dependencies on open-source maintainers, hosting platforms, or data licensors.

Decision FactorBuy From a VendorBuild or Fine-Tune InternallyUse an Open or Self-Hosted Model
Time to initial testOften fastest, but proof of concept may take 4–12 weeksUsually slower because staffing and controls are neededFast software start, slower operational readiness
Model controlDepends on provider contracts and settingsHighest control over training and deployment choicesHigh control if the team can operate the stack
Ongoing costSubscription plus usage, integration, and potential overage feesTalent, compute, security, evaluation, and maintenanceInfrastructure and specialist labor; licensing must be checked
Regulatory evidenceCan be strong if supplied and contractually enforceableOrganization creates the evidence directlyVaries widely by project and license
Best fitOrganizations needing speed and vendor-managed operationsRegulated teams with mature engineering and model-risk capabilitySensitive workloads with strong internal capacity
Open weights do not mean open-source governance, and self-hosting does not remove model risk. An internal system still needs testing, access controls, monitoring, patching, and accountable ownership. Conversely, a commercial vendor may offer stronger security operations and faster updates than a small internal team can reproduce. The most credible comparison uses total cost over 24–36 months, not merely license price, and includes integration, evaluation, governance staff, infrastructure, and the expected cost of retraining or replacing the service.

Common Due Diligence Mistakes

The most common mistake is treating a completed questionnaire as approval. Questionnaires are snapshots assembled by sales or security staff, and they rarely demonstrate performance in the buyer’s environment. Another error is asking only for overall accuracy. Accuracy averaged across a benchmark can conceal poor performance for smaller populations, uncommon languages, edge cases, or high-cost errors. A vendor’s 95% benchmark result also says little about your workload if 5% of misclassifications correspond to the most consequential cases.

Organizations also overstate the value of certifications. ISO 27001 addresses information-security management, not whether a model is fair, explainable, or appropriate for a regulated decision. SOC 2 reports provide control observations over a defined period, not assurance that every AI output is correct. Teams similarly make the mistake of accepting a model card without checking the deployed configuration, ignoring that retrieval, prompts, tool access, and fine-tuning can change behavior after the document was written.

The final major error is failing to plan the exit. If prompts, embeddings, evaluation sets, and workflow logic cannot be exported, the buyer may be locked into a vendor whose prices rise or whose service changes. Require a data-export format, a deletion certificate, transition assistance, and a defined period for continued access during migration. Annual review is too slow when a provider changes its model, acquires a competitor, introduces a new subprocessor, or materially changes data handling. Material changes should trigger notice and reassessment rather than waiting for the next procurement cycle.

Contracts, Costs, and Decision Timing

The commercial review should price the full deployment, not just the advertised per-seat or per-token rate. Small pilots may cost a few thousand dollars, while enterprise implementations can reach tens or hundreds of thousands of dollars in the first year once integration, security review, evaluation, and governance work are included. General-purpose API prices may look inexpensive, but retrieval pipelines, long prompts, repeated evaluations, and tool calls can increase usage charges. A vendor may also charge separately for storage, premium models, private networking, audit exports, or support.

For higher-risk deployments, negotiate specific rights: restrictions on training on customer data, defined subprocessor notice, incident notification within a stated period, audit evidence, intellectual-property allocation, regulatory cooperation, and transition assistance. Reasonable contractual targets include advance notice of a new subprocessor and notification of a confirmed security incident within 24 to 72 hours, depending on the sensitivity of the data. These are negotiating benchmarks, not universal legal requirements. Contracts should connect notice to information that lets the customer investigate rather than requiring the vendor to disclose irrelevant forensic details.

Act before a production pilot, not after procurement complains about the invoice. A focused review can occur in 2–4 weeks for a contained tool, while a high-impact system may need 6–12 weeks or longer for representative testing and legal work. Renewals and material scope changes deserve a new review. The first deployment should use limited data, a named owner, defined acceptance thresholds, and a date on which the organization will decide whether to expand. That discipline keeps pilot enthusiasm from becoming permanent operational exposure.

How Different Organizations Should Scale the Review

A small company should not reproduce the process of a global bank, but it can apply the same logic at a suitable level. A 30-person company testing an internal writing assistant may need a data classification check, account controls, a confidentiality review, and a short acceptance test. A larger financial institution evaluating a credit-decision model may need legal review, model validation, fairness testing, independent security assurance, resilience exercises, and board-level reporting. The difference is driven by data sensitivity, decision impact, regulatory scope, and the organization’s ability to absorb failure.

Vendor due diligence should also include internal readiness. Assign one accountable business owner, one risk or compliance contact, and one technical evaluator. Train users to report harmful outputs, keep an inventory of AI use, and define prohibited uses such as entering regulated or export-controlled information into unapproved tools. For material systems, monitor usage, error rates, overrides, incidents, and changes in model behavior. A dashboard showing token consumption without quality or risk measures creates activity data, not governance evidence.

The defensible standard is traceability: someone should be able to explain which system was approved, for what purpose, on what evidence, under which controls, and with what unresolved risks. If a vendor cannot supply that evidence, the organization must either reduce the use, redesign the deployment, or decline it. This approach is more demanding than collecting badges, but it is proportionate and clearer than assuming either that AI vendors are inherently trustworthy or that every new model presents unacceptable risk.