What AI Vendor Due Diligence Actually Requires
AI vendor due diligence is the structured process of deciding whether a third-party AI product, model, data source, or service provider is suitable for a defined business use. It is not a questionnaire marathon, a model benchmark review, or a single security questionnaire completed by sales. A defensible process connects the intended use to evidence about data rights, model behavior, security, operational resilience, human oversight, regulatory exposure, and the provider’s capacity to support the system over time.
Also worth reading: How Can Organizations Implement an Enterprise Agent Governance Blueprint to Control Autonomous AI Systems? · How Should Organizations Structure an Enterprise AI Architecture Roadmap for 2026 and Beyond? · What is the definitive agentic AI security posture for enterprise organizations in 2026?
As of September 24, 2026, the central issue is no longer whether AI deserves supplier oversight. Regulators already expect institutions to manage vendors as technology, data, and operational dependencies rather than as ordinary software purchases. The EU AI Act’s obligations concerning governance and general-purpose AI have applied since August 2, 2025, while most of its remaining provisions are scheduled to begin on August 2, 2026, subject to any late changes and sector-specific implementation. DORA has applied since January 17, 2025, making ICT third-party risk an operational concern for covered financial entities.
The correct diligence depth depends on what the AI will do. A tool that summarizes internal meeting notes presents a different risk profile from a system that rejects loan applications, recommends medical treatment, prices insurance, or makes hiring decisions. Low-impact productivity software may justify a focused review, while a high-impact system often needs model documentation, representative testing, data provenance checks, human-override rules, and an independently testable control environment. The output is a documented decision with conditions, not a universal claim that the vendor is “safe.”
Why Conventional Procurement Falls Short
Traditional supplier reviews tend to measure what is easy to count: certifications, uptime claims, contract terms, encryption features, and the number of supported languages. Those facts matter, but they do not establish whether a model performs acceptably in your organization. A vendor can hold ISO 27001 certification while still producing biased recommendations, training on data it lacks the right to use, or exposing confidential prompts to a human support process it never described in its sales deck.
AI changes the supply chain because performance cannot be separated from the customer’s data, operating procedures, and decisions about human review. The same model may behave differently after fine-tuning, retrieval settings, system prompts, or access permissions are changed. A model card written before deployment may describe the provider’s laboratory configuration rather than the production service being purchased. Due diligence must therefore examine the actual product, relevant documentation, contractual commitments, and intended deployment rather than treating the provider’s brand as a proxy for reliability.
There is also a hidden-provider problem. An apparent single supplier may depend on a cloud host, a foundation-model developer, an external data provider, a logging service, and subcontractors that process information elsewhere. Existing due diligence on unstructured documents can also miss liabilities that appear only in attached reports, spreadsheets, screenshots, or correspondence. As of 2026, many governance frameworks are not keeping pace with adoption, making internal evidence collection more important rather than less. Investors, audit committees, customers, and insurers may examine these dependencies even when the original contract names only one vendor.
A Risk-Based Diligence Method
Start with a written use-case statement that identifies the decision being supported, the people affected, the data involved, the expected business owner, and the consequences of error. Classify the system using consequences and regulatory status rather than the vendor’s marketing language. A useful internal threshold is to give enhanced review to any system that makes or materially supports decisions about credit, employment, housing, insurance, health care, education, payments, or public benefits. Systems that only draft content can still require review when they handle regulated, personal, or confidential information.
Next, obtain evidence that is specific to the purchased service. Request a system diagram, hosting locations, subprocessors, retention settings, deletion guarantees, incident-notification terms, encryption standards, disaster-recovery arrangements, and the identity of any foundation-model provider. For model-based products, ask for evaluation methods, known limitations, performance across relevant demographic or language groups, change-notification practices, and an explanation of whether customer prompts or outputs train shared models. A promise that customer data is “not used for training” is more useful when paired with contractual language, technical restrictions, and audit rights.
A practical evidence threshold is 80% coverage before pilot approval for a high-impact deployment, with every uncovered item assigned an owner and deadline. That is an internal management target, not a regulatory safe harbor. Less consequential pilots can use narrower thresholds, but they should still address data rights, security, incident response, and exitability. The objective is to create an evidence trail showing why the vendor was selected, which risks remain, who accepts them, and what would cause the deployment to stop.
| Evaluation Area | Enterprise-Grade AI Vendor | Lower-Risk Pilot or Standalone Tool |
|---|---|---|
| Intended use | May support regulated or material decisions | Drafting, search, summarization, or personal productivity |
| Evidence expected | Representative testing, data lineage, model documentation, human controls, audit rights | Security review, data-use terms, owner approval, limited test data |
| Monitoring | Ongoing performance, drift, incidents, and vendor-change review | Periodic spot checks and access review |
| Contract posture | Strong audit, notification, IP, indemnity, and exit terms | Standard terms may suffice if exposure is contained |
| Typical review cycle | Quarterly for material changes; at least annually | Before renewal or scope expansion |
Desk research should be followed by tests using data that resemble the real workload without exposing unnecessary personal or confidential information. For retrieval systems, include realistic documents and questions so reviewers can observe whether citations are present, accurate, and relevant. For classification or prediction tools, define error costs separately. A 2% false-negative rate may be tolerable in an internal search feature but unacceptable in a system used to flag suspected financial crime.
Security testing should cover identity and access management, tenant separation, API authentication, secrets handling, logging, vulnerability management, and deletion. Ask whether administrators can turn off logging when troubleshooting and whether support staff can access prompts, files, or evaluation results. Define a maximum support-access window where the vendor cannot provide evidence, such as zero for privileged access without a documented business need. Verify encryption in transit and at rest, but do not stop there: a correctly encrypted database can still contain improperly collected data or weak authorization rules.
Performance testing should include edge cases and groups that may perform differently from the vendor’s headline benchmarks. For general-purpose AI systems, NIST’s AI Risk Management Framework 1.0, published in January 2023, and its Generative AI Profile, released as NIST AI 600-1 in July 2024, provide useful structure even for buyers outside the federal government. Documentation based on a generic model card remains a starting point. Buyers need production-relevant acceptance criteria, monitoring thresholds, a process for reporting material model changes, and a right to suspend use when those thresholds are breached.
Comparing Vendors, Build, and Open Models
AI vendor due diligence is sometimes framed as a choice between buying a product and building a system. That comparison is too narrow because most products combine third-party models, cloud services, proprietary data, and internal automation. A buy decision may still require substantial integration, governance, and evaluation work, while building a system can preserve control without eliminating dependencies on open-source maintainers, hosting platforms, or data licensors.
| Decision Factor | Buy From a Vendor | Build or Fine-Tune Internally | Use an Open or Self-Hosted Model |
|---|---|---|---|
| Time to initial test | Often fastest, but proof of concept may take 4–12 weeks | Usually slower because staffing and controls are needed | Fast software start, slower operational readiness |
| Model control | Depends on provider contracts and settings | Highest control over training and deployment choices | High control if the team can operate the stack |
| Ongoing cost | Subscription plus usage, integration, and potential overage fees | Talent, compute, security, evaluation, and maintenance | Infrastructure and specialist labor; licensing must be checked |
| Regulatory evidence | Can be strong if supplied and contractually enforceable | Organization creates the evidence directly | Varies widely by project and license |
| Best fit | Organizations needing speed and vendor-managed operations | Regulated teams with mature engineering and model-risk capability | Sensitive workloads with strong internal capacity |
Common Due Diligence Mistakes
The most common mistake is treating a completed questionnaire as approval. Questionnaires are snapshots assembled by sales or security staff, and they rarely demonstrate performance in the buyer’s environment. Another error is asking only for overall accuracy. Accuracy averaged across a benchmark can conceal poor performance for smaller populations, uncommon languages, edge cases, or high-cost errors. A vendor’s 95% benchmark result also says little about your workload if 5% of misclassifications correspond to the most consequential cases.
Organizations also overstate the value of certifications. ISO 27001 addresses information-security management, not whether a model is fair, explainable, or appropriate for a regulated decision. SOC 2 reports provide control observations over a defined period, not assurance that every AI output is correct. Teams similarly make the mistake of accepting a model card without checking the deployed configuration, ignoring that retrieval, prompts, tool access, and fine-tuning can change behavior after the document was written.
The final major error is failing to plan the exit. If prompts, embeddings, evaluation sets, and workflow logic cannot be exported, the buyer may be locked into a vendor whose prices rise or whose service changes. Require a data-export format, a deletion certificate, transition assistance, and a defined period for continued access during migration. Annual review is too slow when a provider changes its model, acquires a competitor, introduces a new subprocessor, or materially changes data handling. Material changes should trigger notice and reassessment rather than waiting for the next procurement cycle.
Contracts, Costs, and Decision Timing
The commercial review should price the full deployment, not just the advertised per-seat or per-token rate. Small pilots may cost a few thousand dollars, while enterprise implementations can reach tens or hundreds of thousands of dollars in the first year once integration, security review, evaluation, and governance work are included. General-purpose API prices may look inexpensive, but retrieval pipelines, long prompts, repeated evaluations, and tool calls can increase usage charges. A vendor may also charge separately for storage, premium models, private networking, audit exports, or support.
For higher-risk deployments, negotiate specific rights: restrictions on training on customer data, defined subprocessor notice, incident notification within a stated period, audit evidence, intellectual-property allocation, regulatory cooperation, and transition assistance. Reasonable contractual targets include advance notice of a new subprocessor and notification of a confirmed security incident within 24 to 72 hours, depending on the sensitivity of the data. These are negotiating benchmarks, not universal legal requirements. Contracts should connect notice to information that lets the customer investigate rather than requiring the vendor to disclose irrelevant forensic details.
Act before a production pilot, not after procurement complains about the invoice. A focused review can occur in 2–4 weeks for a contained tool, while a high-impact system may need 6–12 weeks or longer for representative testing and legal work. Renewals and material scope changes deserve a new review. The first deployment should use limited data, a named owner, defined acceptance thresholds, and a date on which the organization will decide whether to expand. That discipline keeps pilot enthusiasm from becoming permanent operational exposure.
How Different Organizations Should Scale the Review
A small company should not reproduce the process of a global bank, but it can apply the same logic at a suitable level. A 30-person company testing an internal writing assistant may need a data classification check, account controls, a confidentiality review, and a short acceptance test. A larger financial institution evaluating a credit-decision model may need legal review, model validation, fairness testing, independent security assurance, resilience exercises, and board-level reporting. The difference is driven by data sensitivity, decision impact, regulatory scope, and the organization’s ability to absorb failure.
Vendor due diligence should also include internal readiness. Assign one accountable business owner, one risk or compliance contact, and one technical evaluator. Train users to report harmful outputs, keep an inventory of AI use, and define prohibited uses such as entering regulated or export-controlled information into unapproved tools. For material systems, monitor usage, error rates, overrides, incidents, and changes in model behavior. A dashboard showing token consumption without quality or risk measures creates activity data, not governance evidence.
The defensible standard is traceability: someone should be able to explain which system was approved, for what purpose, on what evidence, under which controls, and with what unresolved risks. If a vendor cannot supply that evidence, the organization must either reduce the use, redesign the deployment, or decline it. This approach is more demanding than collecting badges, but it is proportionate and clearer than assuming either that AI vendors are inherently trustworthy or that every new model presents unacceptable risk.