What AI Vendor Due Diligence Actually Means

AI vendor due diligence is the structured process of determining whether a third-party AI product, service, or underlying data supplier is suitable for a particular business use. It is not a paper exercise, a short security questionnaire, or a sales demonstration scored against a generic feature list. The process connects vendor claims to testable evidence: how the system was trained or configured, what data it retains, which subprocessors participate, how human oversight works, and what happens when the model produces an incorrect, discriminatory, or prohibited result. For an organization evaluating customer support automation, for example, “accurate” has little value unless the buyer defines the acceptable error rate for the intended population and workflow.

Also worth reading: How Can Organizations Implement an Enterprise Agent Governance Blueprint to Control Autonomous AI Systems? · How Should Organizations Structure an Enterprise AI Architecture Roadmap for 2026 and Beyond? · What is the definitive agentic AI security posture for enterprise organizations in 2026?

The reason AI deserves separate treatment is that conventional vendor review often examines only the vendor, the contract, and the service uptime. An AI system adds a model layer, training or retrieval data, prompts, external tools, deployment infrastructure, and potentially a changing provider underneath the brand presented during the sale. That chain can introduce risks even when the immediate supplier has a respected security program. AI vendors may also acquire datasets, plugins, or infrastructure providers after a contract is signed, so due diligence must include software bills of materials, model documentation, and change-management commitments rather than relying only on an annual questionnaire.

A practical definition is therefore “evidence-based suitability review.” Suitability depends on the use case, the data classification, the people affected, and the organization’s tolerance for disruption. This approach is neither a guarantee of safety nor a reason to reject all third-party AI. It is a way to identify where the organization can accept residual risk, where contractual and technical controls are needed, and where the proposed deployment is not appropriate at all.

Why Traditional Procurement Is No Longer Enough

Traditional procurement usually starts with functional fit, price, implementation effort, and vendor financial health. Those factors still matter, but they do not answer the central AI questions. A legal agreement may restrict how a customer uses outputs, yet it cannot by itself prevent a model from reproducing personal data, making a materially biased recommendation, or generating a plausible but false operational answer. Similarly, an ISO 27001 certificate can show that a security management system exists, but certification does not prove that a particular model was evaluated for robustness, privacy leakage, or performance under real operating conditions.

The regulatory environment has increased the cost of treating AI as an ordinary software purchase. The EU AI Act entered into force on August 1, 2024, with obligations for general-purpose AI models applying from August 2, 2025, and most remaining provisions scheduled to apply from August 2, 2026. Certain higher-risk uses face later application dates. The exact duties depend on the provider’s role, the system’s classification, and the context in which it is placed on the market. US financial institutions also face sector-specific attention; the National Credit Union Administration has published supervisory material on artificial intelligence, while the NYDFS has required covered institutions to address third-party risk and cybersecurity concerns in their programs.

Regulation still leaves substantial room for interpretation. A bank may be accountable for a vendor-generated decision even when the bank did not train the model, and regulators have shown increasing interest in whether oversight is real rather than ceremonial. Organizations should therefore test whether named individuals can stop a deployment, investigate a failure, correct records, and notify affected parties. If nobody can answer those questions, a completed security questionnaire is not persuasive evidence of readiness.

A Due Diligence Process Built Around Evidence

The first stage is to define the use case precisely. Record the business owner, intended users, affected populations, data categories, required decisions, and unacceptable outcomes. For a bank, the proposal might involve summarizing loan applications for human review, while for a hospital it might involve prioritizing clinical messages; the required evidence is not interchangeable. A useful threshold is to require documented performance and safety testing before production, define an escalation path for uncertain outputs, and establish a rollback plan that can be exercised without waiting for a quarterly governance meeting.

The next stage is to map the complete supply chain. Ask for the immediate provider, foundation-model creator where known, hosting providers, retrieval sources, plug-in developers, monitoring services, and all material subprocessors. Require notice before adding a subprocessor that changes data location, model training practice, or the system’s core control environment. Review software bills of materials, container or dependency inventories, API logs, and access-control records where proportionate. The objective is not to know every incidental network connection; it is to identify dependencies that can materially alter confidentiality, integrity, availability, or regulatory compliance.

Evidence should be current and specific to the offered product. A vendor policy describing “responsible AI” should be paired with model cards, system cards, evaluation reports, data lineage, red-team results, and incident records. For high-impact uses, the buyer should request performance broken down by relevant demographic or operational groups, not only an aggregate accuracy figure. Many vendors have legitimate commercial limits on disclosing training data, so the contract should require lawful commitments, audit rights, and clear explanations where raw training records cannot be shared. “The model is enterprise ready” is not evidence; reproducible test conditions and a named owner for remediation are evidence.

Review areaMinimum evidence to requestRed flag that should delay approval
DataData inventory, retention schedule, deletion process, training-use restrictionsVendor cannot say whether prompts or outputs train shared models
PerformanceProduct-specific evaluation with error rates, slices, and known limitationsOnly marketing claims or aggregate accuracy are provided
Human oversightNamed reviewer, escalation path, override authority, monitoring rulesDeployment is mandatory and users cannot challenge outputs
Supply chainSubprocessor list, model dependencies, software bill of materialsMaterial providers are undisclosed or can change without notice
SecurityIndependent reports, penetration-test summary, access controls, incident historyCertifications are offered as a substitute for product evidence
Exit and remediesExport format, transition assistance, deletion confirmation, incident notificationVendor controls data but offers no usable transition path
## Comparing In-House, Open-Source, and Managed AI

There is no single “safe” AI sourcing model. A managed commercial service may offer stronger investment in safety research and faster access to capable models, but the customer has less control over model changes and may depend on external infrastructure. A self-hosted open-source model can provide greater configuration control and may reduce ongoing per-token costs, but it transfers more responsibility for hosting, patching, evaluation, and monitoring to the buyer. A smaller specialist vendor may understand a narrow workflow better than a general provider, yet it may have a thinner security program and a less diversified operational base.

The correct comparison is risk-adjusted, not feature-driven. Calculate the total cost of ownership for a period such as 24 to 36 months, including licenses, inference, integration, data preparation, evaluation, security review, monitoring, retraining, legal work, and the cost of responding to an incident. A low subscription price can be misleading if the vendor requires expensive data segregation, manual review, or a costly exit. Conversely, an open-source model can look inexpensive while the internal team spends months building a reliable control environment around it.

Decision factorManaged vendorOpen-source modelHybrid approach
SpeedUsually fastest; platform infrastructure is providedSlower because deployment and controls are built internallyModerate; a managed API supports selected workflows
ControlLower control over model updates and infrastructureHighest control over configuration and hostingMore control for sensitive or regulated tasks
TransparencyProvider documentation is often easier to request, but training details may be limitedTeams can inspect code and architecture, but training data may remain opaqueDepends on which model and data each component uses
Cost profileRecurring API, seat, or platform fees plus usageInfrastructure and engineering labor, sometimes with lower variable costSelected managed services plus internal monitoring and evaluation
Best fitFast deployment with a mature provider and acceptable contractual termsOrganizations with strong MLOps, security, and legal capacityEnterprises separating lower-risk automation from sensitive work
The best option is frequently hybrid. A company can use a managed model for low-risk drafting while keeping sensitive records, approvals, or regulated decisions in a controlled environment. That design is not automatically safer: a hybrid architecture can create more integration points and unclear accountability. It should be adopted only when the data flow, human checkpoints, and ownership of every output are explicit.

Costs, Timelines, and Decision Thresholds

AI vendor due diligence has no universal market price. A lightweight review of a low-risk internal assistant might take one to two weeks, while a review of a bank’s credit decisioning, medical prioritization, or employment system can take three to nine months. External assessments commonly range from roughly $10,000 to $50,000 for a focused review, while broader validation, red-team testing, or regulatory analysis can reach $100,000 or more. Internal work can be cheaper in cash but expensive in staff time if it requires legal, privacy, cybersecurity, data science, and business specialists to coordinate.

Use thresholds that reflect consequence, not novelty. Require a full review before production when the system processes regulated, confidential, biometric, financial, or health data; makes decisions affecting eligibility, employment, credit, safety, or essential services; uses a child’s information; or creates automated decisions without meaningful human review. For low-risk summarization with public data, a shorter review may be reasonable if outputs are clearly labeled, users can verify them, and no consequential action follows automatically. Even then, a baseline record should be retained so that problems can be investigated.

A practical governance trigger is to pause the purchase when the vendor cannot identify the model version, cannot explain data retention, or refuses contractual incident-notification terms. A 72-hour notification commitment may be a useful starting negotiation point for security events, while material model or subprocessor changes may need at least 30 days’ notice. Those are negotiating positions, not universal legal requirements. The organization should calibrate them to the ability to migrate or shut down, the sensitivity of the data, and the contractual rights available under applicable law.

Budget for monitoring after approval. An initial review is only a snapshot; model updates, prompt changes, new integrations, and drift can alter the risk. Establish quarterly reviews for high-impact systems, annual reassessments for stable low-risk systems, and immediate review after a serious incident, acquisition, material model release, or change in data use. Contractual commitments should state who pays for revalidation and remediation when a new release changes the system’s behavior.

Common Due Diligence Mistakes

The most common mistake is treating a questionnaire as a technical assessment. Responses such as “we use encryption,” “we conduct bias testing,” or “we follow industry standards” are too broad to support a decision unless they identify the scope, date, test method, exceptions, and accountable owner. Another frequent error is allowing a sales demonstration to define the risk profile. A polished interface can hide a brittle backend, an undocumented data source, or an operational process that depends on manual corrections.

Organizations also make the mistake of asking for assurance without defining a remedy. If a vendor admits that its model has known limitations but offers no correction period, credit for failed deployment, audit cooperation, or termination assistance, the buyer carries much of the risk. A related error is confusing accuracy with suitability. A model can score well on broad benchmarks while performing poorly for a particular language, document type, regional population, or edge case that matters to the buyer.

The final mistake is assuming a human in the loop solves the problem. A reviewer who sees 200 recommendations per hour, lacks time to investigate, and faces pressure to approve them is not meaningful oversight. Human review should be tested with representative workloads, explicit uncertainty signals, authority to reject an output, and a process for correcting the underlying record. If the organization cannot measure override quality and downstream outcomes, it should not claim that human oversight controls the residual risk.

When to Act, Reassess, or Walk Away

Act promptly when AI is moving into a consequential workflow, even if the initial pilot is informal. Shadow deployments, browser extensions, spreadsheet macros, and employee subscriptions can introduce unapproved data processing. Keep an inventory of tools, owners, purposes, data sources, and business units, and require business owners to distinguish experimentation from production use. A useful governance question is whether the organization would be comfortable explaining the deployment to a customer, regulator, or journalist without relying on undocumented assumptions.

Reassess when the model version changes materially, the vendor changes its subprocessor, the tool begins processing a new data class, or its role shifts from drafting to decision-making. Also reassess after a data breach, repeated quality failures, employee complaints, an audit finding, or evidence that performance has declined for a particular group. The review interval should be shorter for dynamic or high-impact systems; annual review alone is often too slow when a provider updates model behavior frequently.

Walk away when the vendor refuses basic transparency, cannot support a lawful use of the data, will not accept reasonable security obligations, or makes claims that cannot be tested. It may also be necessary to reject a specific use even with a capable vendor, such as automated eligibility decisions where reliable validation and meaningful appeal rights cannot be delivered. Walking away is not a failure of innovation. It is recognition that some systems require capabilities the organization, vendor, or legal framework cannot currently support.

The Best Next Step for Buyers

Start with a one-page use-case and risk statement before selecting a vendor. Identify the data, affected people, decision authority, failure consequences, and the exact evidence needed for production. Then request a standard evidence package from each candidate: architecture and data-flow diagram, model and dependency information, security materials, evaluation results, subprocessor schedule, human-oversight design, incident history, and proposed contract terms. Compare responses by evidence quality, not by how quickly a sales team answers.

The strongest due diligence process ends with a documented decision, not merely a contract signature. Record the accepted risk, unresolved questions, remediation deadlines, monitoring metrics, approval authority, and conditions that would trigger suspension. Revisit that record after deployment. The organizations that handle AI responsibly are not those that claim every model is safe; they are those that can show what they tested, what they do not know, who can intervene, and what happens when the system is wrong.