Commercial teams have been shown a lot of AI this year. Some of it is automation with a better deck — genuinely useful, often overdue. Some of it changes what a twelve-person brand team can cover. The feature list rarely separates the two.
The test that does
Language models are dependable on work that is high volume, well specified and independently checkable. Reconciling a feed against the prior period. Drafting a first-pass territory summary. Flagging a variance against a threshold. Assembling a briefing pack from governed data. In each case the person receiving it can see within seconds whether it is right, and one team we work with has taken eighty percent out of the time spent researching answers to stakeholder questions this way.
The picture changes where being wrong is expensive and the error is invisible on inspection. Interpreting an ambiguous coefficient. Choosing what goes to a customer. Judging whether an evidence gap is material enough to delay a submission. Those outputs look the same whether they are right or not.
Looks right
- Every output carries the query, source and timestamp behind it
- A named approver signs anything that leaves the building
- It writes to data, never to the structure of your systems
- Every run is logged and can be replayed
Worth questioning
- Outputs that cannot be traced to a source
- An approval step treated as a formality
- Recommendations with no stated reason
- Anything you could not switch off on a Tuesday
Deploying it so the team keeps using it
Reconciliation and drafting. If it returns real hours there, the team will extend it themselves. If it does not, you have learned that at low cost.
A named person approves, with the capacity to actually read it. That capacity is the control; the signature is not.
Log what got rejected and why. Three months in, the rejection rate is the most informative number you have.
Scope grows when the rejection rate falls, rather than when the quarter turns.
The gain is not fewer people. It is that the people spend their hours on judgement, which was always the part in short supply.
Worth being clear about the limit: none of this tells you whether the model reasoned well, only whether the output was acceptable. That distinction matters most where you are least able to check it, which is why the boundary between automated and approved is worth drawing generously in favour of the person.