In July 2025, an AI coding agent built by Replit deleted a live production database mid-project, despite an explicit code freeze instruction from the client. When questioned, the agent admitted to running unauthorised commands, fabricated test results to cover the gap, and initially told the client a rollback was impossible. The client hadn’t touched his own code.
vendor’s AI had changed itself, and nobody had told them.
The 30-second take
Most AI risk frameworks are built to assess a vendor once, at onboarding, and very often it is just a risk and compliance step.
But AI vendors don’t stand still: they retrain models, swap data sources, and ship silent updates on their own release schedule. When those changes land without notice, the organisation using the tool inherits the risk without ever seeing it coming.
Vendor change management has to be a standing control, not a one-off assessment.
A live production database, gone in nine days
Jason Lemkin was testing Replit’s “vibe coding” agent over a 12-day sprint. On day nine, the assistant ran destructive commands against a live database holding records on more than 1,200 executives and nearly 1,200 companies, then wiped it. Reporting from Fortune and The Register confirmed the agent had been told explicitly not to make changes without approval during a code freeze and did so anyway, then fabricated data and misrepresented its own recovery options when challenged.
Replit’s CEO publicly acknowledged the failure and rolled out new safeguards, including hard separation between development and production environments.
The lesson for everyone isn’t about Replit specifically. It’s that the vendor’s system behaved in a way no contract review or onboarding questionnaire would have caught, because the risk didn’t exist until the vendor’s own product changed mid-engagement.
What the frameworks already say, and why it’s not being applied
NIST’s AI Risk Management Framework is explicit on this point.
Its GOVERN 6.1 function calls for policies addressing third-party AI risk, and GOVERN 6.2 requires contingency processes for failures or incidents in third-party data or AI systems specifically. That is a standing obligation to monitor, not a one-time sign-off.
ISACA’s 2025 review of the year’s AI incidents reached a similar conclusion from the opposite direction: looking at real failures, including a hiring platform left exposed by a test account with default credentials, it found that procurement is treated as the finish line for AI risk management rather than the starting point.
Its recommendation is blunt: ask vendors where models run, what data they retain, how they monitor and test it ongoing, how incidents are handled, and who is accountable in their organisation, and keep asking as the relationship continues, not just at signing.
Ask your organisation these questions
Do we have contractual notice rights when a vendor materially changes a model, data source, or service architecture we depend on?
Who in our organisation would actually find out if a vendor’s AI system changed behaviour overnight, and how long would that take?
Is vendor AI risk reviewed only at onboarding, or is there a standing agenda item that revisits it through the life of the contract?
Do we maintain a current inventory of every third-party AI component in use, including who owns the relationship and when it was last reassessed?
If a vendor’s AI system took an unauthorised or destructive action tomorrow, do we know what our contractual and operational recovery options actually are?
Have we tested a vendor change or incident scenario, rather than just reviewed the policy on paper?
Operational resilience means monitoring the relationship, not just the contract
The organisations most exposed here are the ones that treated vendor due diligence as a procurement milestone rather than an important and ongoing discipline.
A signed contract and an initial “process driven” risk assessment tell you almost nothing about what a vendor’s AI system will do today or a year into the relationship.
Too often I have heard CEOs and senior executives tell me, “why do we make this so hard“, “they are a large organisation using this already“, and “we are so far behind, even the Board is saying we need to get on with it and we cannot even get this through“.
Embedding change controls, standing initial and ongoing assurance/testing review points, and real contingency testing into vendor oversight is what separates organisations that catch a vendor-side failure early from ones that find out from their own customers, or from a headline.
Want to see how your organisation’s vendor oversight measures up?
Try our readiness snapshot and practical tools to strengthen your AI governance approach.
Free 3–5 minute AI diagnostic
Know where your AI governance stands in five minutes.
Use a short diagnostic to test practical AI governance, oversight and risk controls. Get an immediate visual result and suggested next focus areas.
Practical tools for boards, executives, auditors and risk professionals.
Privacy note: your individual results are not stored by Innovation of Risk. Results stay in your browser; we only track aggregate usage such as page views and average score once you leave our page.
AI Readiness Snapshot
Quick Snapshot
Artificial Intelligence Risk Readiness Snapshot
A compact readiness check to help leaders see where AI governance, oversight and risk controls may need attention before moving into the full toolkit.
Privacy note: this Quick Snapshot runs in the browser only. It does not send answers to this site, does not call ChatGPT and does not generate a server-side workbook. Use it as a light indicator, not a complete assessment.
View
Please complete all areas below:
0%
Not startedSelect a group on the left to answer the 10 questions.
Response map
Capable but informal
Responsible AI maturity
Uncontrolled experimentation
Policy theatre risk
Responsible-use behaviour ↑
Formal governance / controls →
Average
Snapshot positionAnswer the groups to move this marker.
Suggested next focus
Complete the snapshot to identify the lowest-scoring areas.
Domain signals
Domain movement guide
Each coloured line on the visual relates to a domain below. Domains already near advanced may show little or no movement line.
Full AI Maturity Assessment capabilities
Extend the snapshot into a supported AI governance review with:
Role-based survey support and detailed governance assessment
Target State Planner, Scenario Lab and Dependency Mapping
Service Provider AI Maturity and Action Plan Map
AI Risk Assessment module for individual AI use cases
Use Commence Snapshot or the blue area buttons to begin with Strategy & Governance and continue through each question group.
Note: we do not hold your individual answers or any identifying details from this Quick Snapshot. We only retain the anonymous average outcome of each completed or updated snapshot response to show the overall average for all users.
Strategy & Governance
AI use-case ownership, accountability and board or executive visibility.
Strategy & ownership · Q1
AI use cases are identified, documented and owned by the business.
Strategy & ownership · Q8
Accountability is clear across business, risk, compliance, technology and executive teams.
Human oversight · Q10
Board or executive reporting includes AI risk, maturity and responsible-use progress.
EmergingAd hoc or not yet consistent
DevelopingSome practices exist but are uneven
ManagedDefined and mostly embedded
AdvancedMature, monitored and improving
Risk, data & third parties
Risk assessment, escalation, data/privacy/security review and third-party AI oversight.
Assessment & escalation · Q2
AI risks are assessed before pilots, procurement, deployment or material change.
Assessment & escalation · Q3
High-risk AI use cases are escalated for senior approval before they go live.
Data, privacy & security · Q4
Data, privacy, cyber and information-security risks are reviewed before AI tools are used.
Third-party AI · Q6
Third-party AI tools, vendors and embedded AI features are assessed before use.
EmergingAd hoc or not yet consistent
DevelopingSome practices exist but are uneven
ManagedDefined and mostly embedded
AdvancedMature, monitored and improving
Oversight, monitoring & controls
Human oversight, control monitoring and learning from incidents or unintended outcomes.
Human oversight · Q5
Human oversight is defined for AI-supported decisions or outputs that matter to customers, staff or operations.
Monitoring & controls · Q7
AI controls are monitored after implementation, not only checked at launch.
Monitoring & controls · Q9
AI incidents, errors, complaints or unintended outcomes are captured and reviewed.
EmergingAd hoc or not yet consistent
DevelopingSome practices exist but are uneven
ManagedDefined and mostly embedded
AdvancedMature, monitored and improving
How useful was this snapshot?
Your answers are not stored. This short survey only records usefulness and optional feedback.
15
4/5
Email snapshot results
Enter the recipient address and the plugin will send the results through the site email service.
The plugin sends this email through WordPress mail. It does not store the individual snapshot answers.
Sample-data DemoView the controlled demo without opening the paywalled full toolkit.
Sample-data demo
Explore the AI Maturity & Risk Assessment Toolkit
A controlled demonstration using sample data so users can see the toolkit outputs without entering organisational information.
Controlled demo: This demo shows representative maturity outputs, AI risk model classification, action planning and browser-local workbook messaging. Export, email and participant submission paths are disabled in demo mode.
View
Sample organisation snapshot
This view uses realistic sample data to show the type of conversation the full toolkit supports.
62%
Managed, with clear gaps
Governance and monitoring are forming, but third-party AI and data/privacy review need stronger consistency.
Capable but informal
Responsible AI maturity
Uncontrolled experimentation
Policy theatre risk
Responsible-use behaviour ↑
Formal governance / controls →
Domain signals
Strategy & ownership63%
Assessment & escalation55%
Data, privacy & security48%
Human oversight58%
Monitoring & controls72%
Third-party AI38%
Maturity outputs with sample data
The full toolkit combines role-based behaviour signals, detailed maturity scoring, evidence prompts, human-focus indicators and target-state planning.
Demo mode is view-only. Real assessment entry, encrypted save, report email and workbook export remain available only in the full toolkit.
Example management insight
“AI usage is increasing faster than formal control ownership. The next uplift should focus on procurement gates, data/privacy review and post-implementation monitoring.”
AI risk model builder preview
This sample use case shows how the full toolkit helps classify a specific AI initiative and prepare a browser-local workbook.
Use case
Customer-service generative AI assistant using internal knowledge articles.
Initial path
Enhanced review recommended due to customer interaction and data/privacy considerations.
Human risk
Medium-high: customer impact and quality of advice need oversight.
Data/security risk
Medium: internal content, access controls and logging need validation.
In demo mode the workbook download is disabled. In the full toolkit, workbook generation is browser-local.
Action plan map preview
Scenario Lab and target-state actions can seed a practical action map for management discussion.
AI governance and decision rights4 / 5 • 2 plans
Risk assessment, testing and assurance3 / 5 • 1 plan
Data privacy and security controls3 / 5 • 2 plans
Human oversight and responsible decisioning4 / 5 • 2 plans
Monitoring, incidents and control review3 / 5 • 1 plan
2 plans
Program Group 1
Governance foundations and decision rights
Action Plan 1
Confirm named AI decision-rights owner and escalation pathway.
Governance foundation
Action Plan 2
Introduce a lightweight AI approval gate for high-impact use cases.
Governance foundation
2 plans
Program Group 2
Assurance, oversight and control lift
Action Plan 3
Define human-in-the-loop review for customer-facing AI outputs.
Control lift
Action Plan 4
Create post-implementation control indicators and review cadence.
Control lift
Demo privacy and control posture
The demo is intentionally controlled. It uses sample data only and does not ask users to enter real organisational assessment content.
Disabled
Email reports, participant submissions, full workbook export and real assessment save paths.
Shown
Representative visuals, sample scoring, action map examples and privacy messaging.
Purpose
Help users understand the value of the full toolkit before requesting access.
Next step
Use the full toolkit for real assessment work, private session mode, encrypted browser-local save and browser-local workbook generation.
Microsoft Azure AI Foundry and Amazon Bedrock show how model retirement can shorten notice periods, stop requests and require code changes. The EU's DORA framework shows why notification, objection and exit rights must connect to a tested operational response.
The Digital Transformation Agency’s Microsoft 365 Copilot trial shows how a bounded experiment can produce evidence about benefits and limitations. OECD adoption research and UK Government assurance guidance point to an operating model that helps organisations test, scale or stop AI responsibly.
The Australian Cyber Security Centre's procurement and AI supply-chain guidance shows why vendor assurance must be refreshed when services change. OAIC guidance adds a clear requirement for organisations to conduct privacy due diligence on commercially available AI products.
The FCA's 2026 AI Live Testing cohort and Supercharged Sandbox show how Barclays, Experian, Lloyds Banking Group, UBS and other firms are building evidence through controlled testing. NIST's tailorable AI RMF Playbook provides a practical basis for proportionate triage.