Clarify the task and decision
This guide turns external ai provider assessment into a reviewable operating workflow. It connects domain decisions, ownership, evidence, and acceptance so the result continues to work in production.
Map the complete data flow, training use, transfers, logs, change controls, and exit path before approving a provider.
Practical workflow
- 1
Draw the data flow from collection through processing, logs, cache, support, backup, and deletion.
- 2
Classify each data category and connect it to purpose, legal role, location, recipient, and retention.
- 3
Request architecture, contractual, operational, and testing evidence for every material claim.
- 4
Score risk and record required controls, owners, acceptance evidence, and residual risk.
- 5
Approve only the documented configuration, then monitor subprocessors, incidents, changes, and deletion evidence.
Worked example or tool
A weighted risk score turns provider answers and evidence into an approval condition. In the tool, also record the baseline, owner, decision, evidence, open issue, and approval date. Use a real page or transaction so the team sees dependencies, exceptions, and the maintenance work that follows release.
| Decision point | Record | Acceptance criterion |
|---|---|---|
| Baseline | Observed current state | Source and date recorded |
| Decision | Selected option and rationale | Risk and audience considered |
| Evidence | Test, document, or measure | Reviewable and version-specific |
| Approval | Name, role, and date | All mandatory criteria met |
Map data, purpose, and provider behaviour
Describe representative inputs, attachments, outputs, metadata, user identifiers, and administrative records. Trace them through inference, filtering, logging, support, evaluation, and deletion. A diagram should distinguish customer-controlled configuration from provider defaults.
Ask whether any content is used for training, human review, abuse detection, benchmarking, or model improvement. Record the answer for each product tier and feature because enterprise endpoints, consumer interfaces, optional feedback tools, and preview functions often follow different rules.
Test ordinary, sensitive, and accidental-input scenarios.
Identify every derived record and recipient.
Verify training and review settings in contract and console.
Record product version, plan, region, and assessment date.
Verify transfers and remote access
List processing and access countries for the primary service, subprocessors, support, security operations, backups, and disaster recovery. Assess transfer mechanisms and supplementary measures against the actual data and access model rather than accepting a global privacy statement.
Remote support can create a transfer even when storage remains local. Require role controls, approval, time limits, logging, and customer notification for exceptional access. Determine who holds encryption keys and whether the provider can disclose readable content.
- 1
Reconcile location claims with architecture and subprocessor evidence.
- 2
Review government-access and transparency information.
- 3
Test whether support access can be region-restricted.
- 4
Document approved exceptions and compensating controls.
Score operational and change risk
Use separate scores for data sensitivity, service criticality, supplier control maturity, transfer exposure, lock-in, and evidence quality. Keep mandatory requirements as gates, since a strong average must not offset prohibited training use or an unacceptable processing location.
Evaluate model changes, output drift, availability, rate limits, moderation changes, and deprecation. Require notice periods, version pinning where available, evaluation before upgrades, export capability, and an exit plan that covers prompts, glossaries, logs, and integrations.
Define pass-fail gates before reviewing vendors.
Weight evidence quality, not presentation quality.
Test a representative workload and failure modes.
Cost and rehearse a credible replacement path.
Approve a bounded use case and monitor it
Approval should state permitted data, users, functions, integrations, regions, retention, and required settings. Convert conditions into access controls, data-loss prevention, user guidance, monitoring, and contract obligations. Broad approval of a brand name is not a usable control.
Monitor subprocessor notices, policy changes, security events, model releases, spend, data patterns, and output quality. Reassess after material change and at a fixed interval. Keep a current evidence pack so governance can explain why the service remains acceptable.
- 1
Publish the allowed-use boundary in plain operational language.
- 2
Apply required settings through central administration.
- 3
Review alerts and supplier changes with named owners.
- 4
Suspend affected uses when a mandatory condition fails.
Test the service against a controlled evaluation pack
Build an evaluation pack from representative tasks before choosing a provider. Include ordinary content, sensitive boundary cases, long documents, tables, conflicting instructions, terminology requirements, unsupported requests, and attempts to reveal system or customer information. Define expected behaviour, unacceptable behaviour, reviewer guidance, and a severity level for every case. Run the same pack against the exact product tier, region, model version, safety configuration, and integration pattern proposed for production.
Capture complete requests, responses, timestamps, settings, latency, token or character use, refusal behaviour, citations, and reviewer scores. Repeat a sample to observe variation. A provider should not receive a high score because one carefully selected response looks good. Measure consistency, correction effort, failure detectability, and whether logs contain material that policy said would not be retained.
Treat the evaluation as a maintained control. Rerun critical cases after model, prompt, filter, region, or provider changes and compare with the approved baseline. Set stop conditions for severe privacy, security, factual, or discrimination failures. Keep evaluation data protected because it may contain realistic sensitive scenarios, and separate it from material the provider may use for product improvement.
Use the exact contracted product, model, region, and settings.
Define expected and prohibited behaviour before viewing results.
Measure consistency, correction effort, logging, and operational cost.
Rerun critical cases after every material service change.
Roles, evidence, and approval
Security and privacy reviews must describe the production configuration rather than a generic provider. Record the exact service, region, feature flags, optional telemetry, support access, subprocessors, encryption boundaries, retention settings, and customer responsibilities. Reassess after material architecture, contract, provider, or purpose changes and keep the decision linked to the evidence reviewed.
Operations and maintenance
The work does not end at publication. Link the language version or configuration to its source, monitor quality and service measures, and define concrete review triggers. Triggers include source changes, legal changes, new audience needs, recurring support questions, technical changes, and incidents. A named owner evaluates the trigger, opens a new revision when needed, and records renewed approval.
Release checklist
The complete processing path is documented.
Controller and processor roles are agreed.
Locations and subprocessors are evidenced.
Training and secondary use are explicitly addressed.
Access, encryption, logging, and incident controls are verified.
Retention and deletion are defined per data category.
International transfers and safeguards are documented.
Changes, audits, exit, and evidence ownership are assigned.