The risk with AI in professional work is not that it is unreliable. It is that it is reliable enough to stop being checked.

A tool that fails obviously trains you to verify. A tool that is right nineteen times out of twenty trains you to trust it — and the twentieth answer goes out on your letterhead, under your professional indemnity, with your name on it. The responsibility does not transfer.

The three things that actually go wrong

1. Figures that came from nowhere

The most common failure is a number that appears in a summary but exists in no source document. It usually looks plausible — a total that is close to the real one, a subtotal that was inferred rather than read. It is the hardest error to spot precisely because nothing about it looks wrong.

Check: take every figure that will appear in the advice and find it in the source. Not a similar figure — that figure. If the tool cannot show you where it came from, treat it as unverified.

2. Rates and thresholds from the wrong year

Tax rates change, and models are trained on years of text in which the old rate was correct. An answer using last year's dividend allowance or a superseded NIC threshold reads exactly like an answer using this year's.

Check: any rate, allowance, threshold or band gets checked against the current HMRC published figure. This takes seconds and catches the error class that causes the most rework.

3. Confident answers to underspecified questions

Ask about a client's position without stating their residence, their year end, or whether a company is close, and you will still get an answer. It will be a good answer to a question you did not ask.

Check: before accepting the conclusion, read back what the answer assumed. If the assumptions are not stated, they were still made.

A review method that survives January

Elaborate checking procedures get abandoned the moment a practice is busy, which is exactly when they are needed. This one is short enough to survive:

  1. Trace every figure that reaches the client back to a source document or a computation you can see.
  2. Verify every rate against the current published figure.
  3. State the assumptions in the advice itself. If you cannot write them down, you do not yet understand the answer.
  4. Read it as the client will. An answer that is technically correct and practically misleading is still a problem.

Steps 1 and 2 are mechanical and take a few minutes. Steps 3 and 4 are the professional judgement that no tool replaces.

What to demand from the tool

Most of step 1 should not be your job. A tool used for professional work should show, for every figure it produces, which document and which part of it the figure came from — so checking is reading a citation rather than re-deriving the number.

That is a deliberate design choice, not a feature that appears by accident. TaxStats AI traces figures back to their source document and flags any it could not verify, which turns the review from a re-computation into a read-through. Where a figure cannot be traced, it says so rather than presenting it with the same confidence as one that can.

It is also worth asking what the tool does when it does not know. An answer of "the records do not contain this" is far more useful than a confident guess, and a tool that never says it is a tool that is guessing somewhere.

The bottom line

Use AI for the reading — the bank statements, the management accounts, the pile of receipts. That is where the hours are, and it is where the work is most mechanical.

Keep the judgement, the assumptions and the sign-off. The checking method above costs a few minutes per piece of advice, and it is the difference between a tool that makes a practice faster and one that eventually makes it liable.