AI for accountants: what actually works, from someone selling nothing
Which accounting tasks AI genuinely handles, which need dedicated software, and which you should not hand over at all. No vendor pitch attached.
Search for AI in accounting and every result is a company selling accounting software. That is not a conspiracy, it is just who has a reason to write about it. It does mean the genuinely useful answer, which is that a general-purpose assistant you already pay for handles a surprising amount of this and dedicated software is only worth it for a narrower set of tasks, does not get written down very often.
I have nothing to sell you here. Here is the split.
The three-way split
Every task in an accounting practice falls into one of three buckets, and knowing which one you are looking at saves most of the wasted money in this category.
| Bucket | What it means | Examples |
|---|---|---|
| A general assistant already does this | ChatGPT or Claude, no new subscription | Drafting client emails, explaining a treatment in plain English, summarising a standard, first-pass review of a long contract |
| Needs dedicated software | Because it has to touch your ledger, or has to be auditable | Bank feed categorisation, reconciliation, invoice capture at volume, anomaly detection across a full year |
| Do not hand this to AI | The failure is expensive and silent | Filing positions, anything you sign, final numbers, judgement calls a regulator would ask you to justify |
Most of the disappointment in this area comes from buying bucket-two software for a bucket-one problem. You do not need an AI bookkeeping platform to write a better email to a client who has not sent their receipts.
What a general assistant genuinely does well
Explaining things to clients. This is the highest-value, lowest-risk use and almost nobody talks about it because there is no software to sell. Take the treatment you have already decided on and ask for it in language a non-accountant will understand. You are not asking the model for the answer, you have the answer. You are asking it to translate, which is what these things are actually best at.
First-pass document review. A forty-page lease, a new client's prior-year file, a supplier contract. Ask what the unusual terms are, what is missing compared to a standard version, what you should look at first. It does not replace reading it. It tells you which ten pages to read carefully, and it is reliably good at that.
Drafting the boring correspondence. Chasing documents, explaining a fee, the eighth version of the same onboarding email. Keep a handful of your own good examples and hand them over as the reference, which produces something in your voice rather than a generic one.
Getting unstuck on a standard. Asking what a standard covers, in general terms, to orient yourself before going to the actual text. This is genuinely useful and it comes with a hard rule attached, below.
Spreadsheet work you can check. Writing a formula, explaining what an inherited spreadsheet is doing, restructuring data. The reason this is safe is that you can see immediately whether it worked.
What that looks like in practice
The client-explanation use is worth showing rather than describing, because it is the one most people skip and it is the highest return.
The weak version is asking for the answer: "how should I treat this". You then have to verify everything it said, which takes longer than working it out yourself, and the risk is that a plausible wrong answer survives the check.
The strong version is asking for the translation: you write two sentences of what you have decided and why, in your own shorthand, and ask for a version a client will understand, in a specified tone, no longer than a paragraph. You already own the correctness. All the model is doing is the part that takes you fifteen minutes and does not need your expertise.
The difference between those two prompts is most of the difference between accountants who find this useful and accountants who tried it once and concluded it was unreliable. Both conclusions are correct about the prompt they used.
What needs real software
Categorisation and reconciliation at volume is a genuine machine-learning problem and it has been solved reasonably well inside the accounting platforms rather than by chat. If you are doing bookkeeping at any scale, the feature you want is in the tool that already holds the ledger, not in a separate AI subscription. The same goes for invoice capture: dedicated tools that read a scanned invoice into structured data are mature, and pasting invoices into a chat window is a worse version of a solved problem.
Anomaly detection across a full year of transactions is the one where dedicated tools are meaningfully ahead of a general assistant, because the assistant cannot hold a year of data and the software can.
The practical rule: if the task needs to touch your ledger or produce an audit trail, it belongs in software that already does. If it produces words for a human to read, a general assistant is usually enough.
The parts that go wrong
It will state a rule that is not the rule. This is the failure that matters most in this profession, and it is not occasional. Models are trained on a snapshot of text and will describe a threshold, a rate or a filing requirement from whatever version was in that snapshot, in exactly the same confident tone as something current. There is no signal in the output that distinguishes the two.
So: anything a model tells you about a rule, a rate, a threshold or a deadline gets verified against the primary source before it goes anywhere near a client. Use it to find out what to look up, never as the thing you looked up. This is not caution for its own sake, it is the specific way this technology fails, and it is described in more detail in why AI makes things up.
Client data and where it goes. Before pasting a client's financial information into anything, know what the terms say about training on your inputs. Business and team plans generally have stronger commitments than consumer plans. Check the current terms rather than trusting a summary, including this one, because these change.
Confidence is not calibrated. A model that is wrong sounds exactly like a model that is right. In most jobs that costs you an awkward correction. In this one it can cost a client a penalty, which is why the bucket-three list exists.
Review time is real time. Anything you have to check carefully has not saved you as much as it appears to. The tasks worth automating are the ones where checking is fast, which is why translation and drafting are on the good list and judgement is not.
Where to start, in order
If you want a sequence rather than a list:
- This month: use a general assistant for client correspondence and explanations. No new spend, immediate time back, and nothing at risk if it is wrong because you read it before sending.
- Next: first-pass document review on files you were going to read anyway. You will find out quickly how much you trust it, at no cost.
- Then: look at what your existing accounting platform already includes. Most of the categorisation and capture features people go shopping for are already sitting in a tab they have not opened.
- Only then: consider dedicated tools, and only for a task you have measured. Buy for a bottleneck you can name in a sentence.
The order matters because it is cheapest first, and because most practices find that steps one to three cover more than they expected.
What to ask a vendor before buying anything
Every product in this category demos beautifully, because the demo uses clean data. Five questions that get past that:
"What is your accuracy on categorisation, measured how, on whose data?" A number without a method is a marketing number. Ask what counts as correct and who decided.
"What happens to a transaction the model is unsure about?" You want a queue for review, not a confident guess buried in the ledger. A product without an uncertainty path is a product where errors are invisible.
"Can I see what it changed, and undo it?" An audit trail per automated action, and a way back. Non-negotiable in this profession and surprisingly often missing.
"Is my client data used to train your models?" Get the answer in the contract, not on the call.
"What does this do that my existing platform does not?" Ask last, after the demo has shown you the features. A meaningful share of purchases in this category duplicate something already included in the software the practice runs on.
The uncomfortable bit
The tasks AI handles best in accounting are the ones junior staff learn on. Reading a file to find what matters, drafting the client explanation, summarising a standard: that is how judgement gets built. A practice that automates all of it will be faster this year and will have a training problem in five.
Nobody selling this software will raise that with you. It is worth deciding deliberately rather than discovering it later.
For a broader view of automating a small practice rather than the accounting work specifically, see AI automation for small business. For what an agent can and cannot be trusted to do on its own, the agents guide covers the permission question in detail.
Questions people ask about this
- Can AI replace an accountant?
- No, and the split is not close. AI handles explanation, drafting, and first-pass review well. It cannot take a filing position, sign anything, or make a judgement call a regulator would ask you to justify, because it has no way of knowing when it is wrong and no accountability when it is.
- Do accountants need special AI software, or is ChatGPT enough?
- For anything producing words a human reads, a general assistant is usually enough and you probably already pay for one. Dedicated software earns its cost where the task touches your ledger or needs an audit trail: categorisation at volume, reconciliation, invoice capture, anomaly detection across a full year.
- Is it safe to put client financial data into AI tools?
- It depends entirely on the plan and the current terms. Business and team plans generally carry stronger commitments about not training on your inputs than consumer plans do. Check the terms as they stand today rather than trusting any summary, because these change.
- Can AI be trusted on tax rules and thresholds?
- No. This is the specific failure that matters most in this profession: a model will state a rate, threshold or deadline from whatever version was in its training data, in the same confident tone as something current. Use it to work out what to look up, then verify against the primary source before it reaches a client.