Most businesses do not need another AI tool demonstration. They need one useful workflow, built with context, guardrails and human review.
Most AI conversations reach the same moment.
The business is interested. People can see the potential. Then someone says the work needs mapping first, and the energy in the room drops. That step feels like delay. Like overhead. Like something to sort out once the tool is chosen.
So the conversation turns to the tool. ChatGPT, Claude, Copilot, Gemini. Which one should we use?
It is a fair question, asked in the wrong order.
In the calls we have had lately, two things come up together. People do not know where to start. And when training comes up as the answer, it stalls, because no one can find the time to pull the team into a programme, and they are not convinced they need one yet.
Both resolve faster than they look. Not by choosing a tool, and not by booking a course. By taking one piece of work and making it reliable.
That is the gap between AI activity and AI capability. The business is chatting with AI. It has not yet made AI part of how the work reliably gets done.
That gap is wider, and more common, than most leaders assume.
| 69% | ~1.5 hrs | ~90% |
|---|---|---|
| Of firms are already using AI | Average executive use, per week | Report no measurable productivity or employment impact over three years |
A February 2026 working paper from the National Bureau of Economic Research surveyed close to 6,000 executives across the US, UK, Germany and Australia. Around seven in ten firms are using AI. Yet the average executive spends only about 90 minutes a week with it, a quarter are not using it at all, and close to nine in ten report no measurable change in productivity or employment over the past three years. (National Bureau of Economic Research, February 2026)
The tools are in the building. What they have changed about how the work actually gets done is, so far, very little. That gap is the thing worth solving.
AI use almost always starts informally. One person tries a tool on a task that has been annoying them for months, gets a result that saves an hour, and keeps going. Nobody planned it. Nobody approved it. It just worked, so it stuck.
That is a fine way to start. It is not a reliable way to run a business.
The trouble shows up at the edges. Some people get good results and others get mediocre ones, and the business cannot easily see why. It cannot see what information is going into these tools, what is coming back out, or which outputs can be trusted without a second look. The early wins stay locked inside the person who found them. When that person is on leave, or moves on, the capability walks out the door with them.
We see this across the service businesses we work with. A professional services firm where one associate produces sharp first-draft client updates in minutes, while everyone else still starts from a blank page. A club where the operations manager has a clever way of turning committee notes into a clean summary, and no one else knows how she does it. The activity is real enough. What is missing is anything that turns it into something the business can repeat without her in the room.
So the question is not whether AI is being used. It almost certainly is. The question is whether the business can rely on the way it is being used. Right now, for most, the honest answer is not yet.
The first demo nearly always impresses. You give a tool a rough brief, and a few seconds later something usable appears on the screen. A draft proposal. A tidy meeting summary. A board paper that would have taken an hour. The room goes quiet, then someone says the obvious thing. This changes everything.
It feels like progress. Often it is the moment a pilot quietly begins to stall.
Speed is not the same as readiness. The demo worked because the tool is fast and the example was forgiving. The next attempt is messier. A different person runs it with a different brief and gets something slightly off in tone. Someone uses it on a sensitive document without thinking about where that document just went. A result comes back confident and wrong, and nobody catches it until a client does.
When we look at the pilots that fade, the reasons are consistent. The work was never mapped, so the tool was guessing at a process no one had written down. The business context was missing, so the output read like it came from any firm rather than this one. Quality was never defined, so nobody could say whether a draft was good or just fast. Human review was vague. Data boundaries were unclear. Or the pilot tried to fix five things at once and proved none of them.
None of those are tool problems. They are clarity problems. And this is the part most leaders underestimate: AI does not fix unclear work. It usually makes unclear work faster.
The demo trap.
A fast, impressive first result is the easiest thing to produce and the worst thing to plan around. It proves the tool can do the task once. It says nothing about whether the business can do it reliably, next week, with a different person at the keyboard.
Claude, ChatGPT, Copilot, Gemini. Any of them can be genuinely useful. None of them is your operating model.
The same tool, in two hands, produces very different work. Give it a one-line instruction and you get something generic. Give it your services, your audience, an example of what good looks like, the tone you use with clients and the line you will not cross on what data goes in, and the output changes completely. The model did not get smarter between those two attempts. The instructions did.
This is why the chase for a better tool so often disappoints. A business switches from one assistant to another, expecting the new one to fix the inconsistency, and finds the same problem waiting on the other side. The tool was never the limiting factor. The limiting factor was that the business had never explained the work clearly enough for any tool to do it well.
A useful AI workflow needs more than a good model. It needs business context, real examples, clear review rules, defined data boundaries and a named owner. Supply those, and an ordinary tool produces reliable work. Skip them, and the best tool on the market produces confident noise.
There is a lot of attention on prompting at the moment, and some of it is useful. A well-constructed prompt can lift one person’s output noticeably. But a prompt improves a draft. It does not build a process.
Here is the difference in plain terms. A prompt helps one person produce one good summary, one good report, one good email. A workflow makes that result repeatable across the team. It carries the quality standard, decides who owns the output, builds in the review step, sets out how work is handed over, and lets you measure whether it actually helped.
The clearest sign of a prompt masquerading as a process is fragility. The person who wrote the clever prompt keeps it in their head, or in a note on their desktop. When they are away, the workflow stops. When a new staff member joins, they start from nothing. The organisation never moves past its most capable individual, because the capability never left that individual.
Prompting asks a personal question. Can AI help me with this? Implementation asks a business one. Can we rely on this process again next week, when someone else is running it?
This is the shift in mindset most businesses are missing. It is also the one that separates AI activity from AI capability.
| What businesses often start with | What they should start with |
|---|---|
| What tool should we use? | What work are we trying to improve? |
| What can we automate? | Which part of the work repeats often enough to prove? |
| How fast can AI produce this? | What does a good output actually look like? |
| Can we connect it to our systems? | What data is safe to use? |
| Can we roll this out to the whole team? | Can we prove one workflow first? |
AI works far better when it understands the business it is working for. That sounds obvious, and it is routinely skipped.
The useful context is specific. What services do you offer, and to whom. What tone carries authority with your clients or members. What a strong piece of output looks like, with a real example attached. The internal rules that shape the work. The review standard a draft has to meet before it goes anywhere. And the data boundaries that say what can and cannot be put into a tool.
This is not an instruction to pour everything you have into an AI tool and hope for the best. Dumping the entire shared drive into a chatbot is not context. It is risk. The skill is supplying the right context, in the right order, with the right boundaries, so the tool understands the work without being handed things it has no business holding.
We learned this on our own work before we ever recommended it to a client. We got the order wrong at first, reaching for the tool before we had properly described the work, and the early results were uneven because of it. When we built an AI-assisted proposal process inside Dovetail, the gains did not come from the tool. They came from the hours spent first, defining what a good proposal needed to achieve, what a strong draft looked like, and where a human had to stay in control. The structure did the heavy lifting. The tool just moved faster once the structure existed.
It is also why most of what we do starts with a diagnostic rather than a proposal. You cannot build a reliable workflow on top of work the business has never actually described.
Connecting AI to your email, calendar, files, CRM or finance system can be powerful. It can also be the point where a sensible pilot turns into a quiet exposure. The order matters. The permission rules come first, then the connection. Not the other way around.
A simple way to think about it is three tiers, in plain language.
Green: information that is safe to use freely. General queries, public material, internal drafts with nothing sensitive in them.
Amber: information that needs care. Internal documents, draft commentary, anything you would not want quoted out of context.
Red: information that should not go into an unsanctioned tool at all.
Member and client data, financial detail, board materials, anything that carries privacy obligations or professional liability.
Most businesses have never drawn that line, which is exactly why staff cross it without meaning to. Someone pastes a member list into a free tool to clean it up. Someone drops a confidential board paper in to get a summary. They are trying to be useful, and they have never been told where the boundary sits.
The second rule is just as important, and just as often missed. There is a real difference between asking AI to draft, summarise or prepare something, and letting it send, publish or decide. The first is assistance. The second is authority. A human should hold the second, every time, until the business has very good reason to think otherwise.
And before anything gets scheduled or automated, test it by hand. Run the workflow manually, with a person watching the output, often enough to trust it. Automation is what you do once a workflow is proven, not the thing you reach for to prove it.
For the data boundary question specifically, the Office of the Australian Information Commissioner has published practical guidance on privacy and the use of commercially available AI products, including free tools and chatbots that handle personal information. It is a useful reference point for any business holding member, client or employee information.
Our blog We Didn’t Start With the Task goes into this in more detail. “The limiting factor was that the business had never explained the work clearly enough for any tool to do it well.”
Interested. Read our blog Human at the Helm on how to know the rules before you connect.
There is a temptation to point AI at the biggest, most painful problem in the business first. It is usually the wrong choice. The biggest problem is big precisely because it is tangled, high-stakes and hard to define. That is the last place you want to learn whether AI-assisted work holds up.
The better first pilot is the repeatable one. The task that runs every week or every day, where the risk is manageable and the result is easy to compare before and after. It does not have to be glamorous. It has to be the kind of thing you can prove.
Good candidates share a few traits. They repeat often, so you get many chances to test them. They have a clear owner. The data involved is low or manageable risk. You can put the old way and the new way side by side and see the difference. And the output is genuinely useful even after a human has reviewed it, rather than something that only looks finished.
In a service business, that usually points to one of a handful of workflows. Meeting follow-ups and action summaries. An internal briefing that gets written the same way every week. A first draft of a board report. A first draft of a proposal. Email triage. Turning one piece of content into several. A short brief that helps a manager prioritise the week. None of these will transform the business on their own. All of them are safe enough to learn on.
That is the point of the first workflow. It is not there to transform the business. It is there to prove the model, so the second and third decisions are made on evidence rather than hope.
That is also why the first implementation step should be narrow enough to govern properly: one workflow, one owner and one repeatable result before the business tries to scale AI more broadly.
If your team is using AI but you are not sure what you are actually building, here is a practical way to begin. It does not need a budget or a project plan. It needs one workflow and an honest hour.
One workflow. One owner. One honest look at where the judgement lives. That is enough to move from chatting to working.
This is the practical shift we spend most of our time on. Helping service businesses move from scattered, individual AI use to one governed workflow that actually holds up under real conditions.
That does not mean buying more tools, and it does not mean booking another generic training session first. Training still matters. It is what makes capability compound across a team. But it lands better once a first workflow has shown people what good actually looks like, so we start with the workflow and let it become the place the learning happens. It means taking one repeated task and building it properly. On real business context. Tested by hand. Reviewed after it has been used for real. Then improved or extended only once the first one has earned it.
We are preparing a practical workflow activation pathway for businesses ready to turn one repeated task into a governed, AI-assisted workflow. Until that is live, the best starting point is the AI Impact Report.
It gives you a structured read on where AI is already showing up in your business, where the gaps are, and whether your current use is building capability or just creating more scattered activity. That is a clearer foundation than any tool decision you could make today.
The AI Impact Report gives you a practical first read on where AI is helping, where it is inconsistent, and which next step is likely to matter most: training, governance, a workflow pilot or a structured activation pathway.
Most AI pilots stall because they begin with a tool or a prompt before the business has mapped the workflow, defined what good output looks like, clarified data boundaries or assigned ownership. The first result often looks impressive, but without that structure it is hard to repeat, review or scale, so it quietly fades after the initial demo.
AI activity means individuals are using AI tools and getting occasional useful results. AI capability means the business has repeatable workflows, shared standards, clear data boundaries and review processes that make AI-assisted work consistent across the organisation, regardless of who is doing it.
Before connecting AI to email, files, CRM, finance or other systems, the business should classify its data, decide what AI can and cannot access, confirm what AI is permitted to do versus what requires a human, and test the workflow manually before any automation is added.
A good first AI workflow is repeated often, has a clear owner, uses low or manageable-risk data, can be tested by hand, and produces a result you can compare before and after. It is chosen to prove the model reliably, not to transform the whole business at once.
No. For most service businesses, governance starts with practical controls: approved tools, clear data boundaries, review checkpoints and named ownership. Governance should match the risk and maturity of the business, then grow as AI becomes more embedded and touches more sensitive work.
Still feeling stuck? You’re not alone but you don’t have to figure it out solo.
Our DMA helps you cut through the noise and focus on what matters most.