Context
The organisation had the same relationship with generative AI that most services businesses had in 2025: real enthusiasm, a handful of people using it well in private, no standards, and a leadership team that wanted a strategy rather than a collection of anecdotes. I was made the AI adoption lead with a broad mandate and no existing playbook.
The problem
The instinct in this situation is to run training. I'd argue the constraint is almost never capability or knowledge — it's that people can't tell which of their tasks is a good fit, don't trust the output enough to stake their work on it, and have no safe place to fail while they find out. Training that doesn't address those three things produces a spike in usage and a return to baseline six weeks later.
So the real problem was: how do you make adoption an evidence problem rather than a persuasion problem?
Constraints
- A services business, where billable hours are the unit of account and time spent learning is time not billed.
- Client confidentiality, which ruled out a large class of tools and workflows outright and made the governance question load-bearing rather than theoretical.
- A wide skill spread — from engineers who wanted to build agents to delivery staff who wanted a better way to write a status update.
- No budget assumption. Anything I proposed had to justify itself before it could ask for spend.
What I did
The first wave deliberately targeted internal, reversible, high-frequency tasks: story writing, backlog grooming, documentation, status synthesis. These weren't the highest-value workflows in the business. They were the ones where a wrong output costs a minute and everyone can see whether it worked.
The Business Analyst agent became the template that mattered more than the tool. Built with Claude Code and the Atlassian MCP, it drafts Jira stories, grooms backlogs, and generates acceptance criteria from a brief. The design decision that counted was making it configurable, documented, and forkable, so other teams could adopt and adapt it themselves rather than wait on me for a feature. It ended up used well beyond the team it was built for.
Production automation ran behind human-in-the-loop guardrails: n8n orchestrating LLM APIs, cutting more than 60% of the manual effort on targeted workflows. Every automation answers three questions before it ships: what a wrong output costs, who catches it, and how long that takes. Low cost and reversible, it runs unattended. Everything else keeps a permanent review step built into the design.
The responsible-AI standard went in before any incident forced it: internal guidance on evaluation, guardrails, and human-in-the-loop design, covering what a use case has to demonstrate before it goes near client work, what data can't cross which boundary, and what "good enough" means in writing. Leadership cited it, peer teams adopted it, and that mattered more than its content — it gave AI proposals a standard to be assessed against.
Enablement ran as cohorts, each one producing something real: a working automation, an agent config, a prompt library for their own practice. Fifty-plus staff trained, and the output of training was infrastructure, not attendance.
Outcome
- More than 60% of targeted manual effort removed on the workflows we instrumented, with guardrails that survived contact with client work.
- A reusable agent pattern adopted across multiple delivery teams — the thing I actually optimised for.
- A written responsible-AI standard cited by leadership and adopted by peer teams.
- 50+ people trained, each leaving with a working artefact rather than a certificate.
What I'd do differently
I should have instrumented the baseline harder before starting. We have credible before-and-after numbers on the workflows we automated deliberately, but we're weaker on the diffuse gains — the hours people saved on their own once they knew what was possible. That's the number leadership asks for, and capturing it retroactively is nearly impossible. Next time, a two-week measurement window before any enablement runs.