AI Tax Preparation and Circular 230: The Signature Nobody Reassigned
A model can draft the return, propose the election and write the memo. It cannot hold a preparer identification number, cannot be disciplined, and cannot sign. Every duty the tax practice rules impose stays exactly where it was — on the individual who signs, and on the practitioner who owns the firm's procedures.
The short version. Tax practice is one of the few areas where the regulatory framework already assumed that most of the arithmetic would be done by software, so it never allocated responsibility to the software. It allocated responsibility to people — the signing preparer, the practitioner giving advice, the principal responsible for firm procedures — and it did so in language that is entirely indifferent to how the draft was produced. Introducing an assistant therefore changes the work and changes nothing about the accountability, which produces a specific and predictable gap: review calibrated for human error, applied to a different error distribution, evidenced by nothing.
Where the Duty Actually Attaches
Trace a single position from generation to filing. The regulated moments are not where firms instinctively look — the interesting one is step two, which almost never leaves a record.
The model drafts
No status, no dutyThe assistant classifies expenses, proposes an entity-level election, drafts a memo, or fills a schedule. It holds no professional status, carries no identifying number, and owes no duty to the client or to the tax authority. Nothing about this step is regulated — which is exactly why nothing about it is documented, and why the record of what the model was asked and what it returned usually does not survive the engagement.
A staff preparer accepts
Duty attaches on acceptanceThe moment output is carried into a workpaper or a return, it stops being a suggestion and becomes the firm's work product. Acceptance is a professional act performed by a person, and the person who performs it is exercising — or failing to exercise — the due diligence the practice rules require. In most firms this step has no artefact at all: no note on what was checked, no record that the underlying authority was opened.
The reviewer signs off
Supervision standard appliesA reviewer who assumes the draft came from a trained associate reviews it as they would that associate's work — looking for the errors humans make. Model errors have a different shape: locally fluent, structurally plausible, wrong at the citation. A review calibrated for one failure mode is not a review of the other, and the reviewer's sign-off is nonetheless an assertion that the work was reviewed.
The practitioner signs the return
Full responsibility, no delegationThe signature is a representation of primary responsibility for substantive accuracy. It does not distinguish between a position the signer derived, a position an associate derived, and a position a model derived. Every downstream penalty analysis starts here, and 'the software produced it' has never been a defence for any prior generation of tax software either.
The firm principal owns the system
Independent, structural exposureSeparately from any individual return, the practitioner with principal authority is answerable for whether the firm's procedures were adequate. A firm where assistants entered the workflow informally, tool by tool, without an approved list or a data rule, has a procedures gap that exists whether or not a single return was wrong.
The Error Distribution Your Review Was Not Built For
Firm review procedures encode decades of knowledge about how humans get returns wrong. A first-year associate transposes a figure, misses a state filing requirement, applies last year's threshold, or misunderstands a client's facts. These errors have tells: they cluster in unfamiliar areas, they are accompanied by hesitation in the workpaper, and they usually leave the rest of the return internally inconsistent.
Model errors are shaped differently. Output is uniformly confident regardless of whether the underlying question was routine or unsettled. It is internally consistent, because consistency is what the generation process optimises. And the failure concentrates in exactly the place a busy reviewer is least likely to check: the authority. A citation to a ruling that exists but stands for something adjacent, a regulation section number that is one digit off, a threshold from a prior year presented without qualification, a provision described accurately but applied to a taxpayer it does not reach.
The reviewer who signs off has made a professional assertion that the work was reviewed. If the review looked for transposition errors and the error was a fabricated authority, the assertion is still on the file. This is the mechanism by which competent firms produce unsupported positions without anyone behaving carelessly by the standards they were trained on.
Taxpayer Information Is Not Ordinary Client Data
Most vendor-risk analysis treats client data as a confidentiality and contract question. Return information is not in that category. Disclosure or use outside the permitted scope is a federal offence for a preparer, the consent regime has strict format and timing requirements that a clause in an engagement letter does not satisfy, and cross-border processing carries an additional requirement on top. That framework was written for a world of paper and fax, and it applies unchanged to a paste into a chat window.
Approved enterprise assistant, contract in place
Conditionally usableRequires a written agreement covering confidentiality, no training on your data, retention and deletion, subprocessors and processing location — plus a determination of whether the arrangement fits an exception or needs consent. Contract terms do not by themselves answer the consent question; they only make the answer reachable.
Consumer chatbot, personal or free account
NeverNo agreement, no processing-location control, no retention control, and an interface designed to make pasting a document the fastest path to an answer. This is where the busy-season violation happens, and it is the single control most worth enforcing technically rather than by policy.
Assistant embedded in your tax software
Verify, do not assumeBeing inside the tool you already trust says nothing about where inference runs, which model provider is behind it, or whether an existing agreement was amended to cover it. Read the AI-specific terms as a separate document, and ask specifically whether prompts and returns are retained.
Any assistant that processes outside the country
Additional consent requiredCross-border processing of return information carries its own heightened consent requirement, and the answer depends on the vendor's actual inference geography, not on their headquarters. Get the answer in writing and re-ask it when the vendor changes models.
Client-facing intake bot collecting documents
Disclosure and consent by designAn intake assistant is collecting return information in a channel most firms never mapped. Whatever consent the engagement requires has to be obtained in the required form before collection, and the bot has to be able to record that it was.
Advice Is a Separate Exposure From Returns
Firms concentrate their AI governance on return preparation because that is where the volume is. The larger professional exposure is usually written advice. Practice standards require a practitioner giving written advice to base it on reasonable factual and legal assumptions, to consider all relevant facts the practitioner knows or reasonably should know, to relate applicable law to the facts, and not to rely on representations the practitioner knows are unreasonable — and they prohibit taking the likelihood of audit into account.
A model asked to write a client memo will do several of those things badly by default. It will assume facts not in evidence in order to produce a complete answer. It will generalise from the most common fact pattern rather than the client's. It will present a contested area as settled because settled language is more fluent. And it will not tell you which of its assumptions are load-bearing. Written advice that leaves the firm on letterhead carries the practitioner's professional standard regardless of who typed it, which makes memo generation a higher-supervision activity than return drafting, not a lower one.
Controls That Actually Change the Outcome
- A citation-verification step that is a step. Every authority in model output gets opened and read before it reaches a workpaper, and the workpaper records that it was. This single control removes the majority of the penalty exposure, because it converts fabricated authority from an undetected defect into a caught one.
- An approved-tool list with a technical boundary. Policy alone does not survive the last week of an extension season. Blocking consumer assistant domains on firm devices, and providing an approved alternative that is genuinely faster, is what actually prevents the paste.
- A consent posture decided before the season. Determine for each approved tool whether the arrangement needs taxpayer consent, in what form, and whether processing crosses a border. Decide it in October, not while a client waits.
- Review instructions rewritten for model error. Tell reviewers explicitly what to look for: unverified authority, unstated assumptions, prior-year thresholds, and confident treatment of unsettled areas. A reviewer who does not know the error distribution changed cannot compensate for it.
- Prompt and output retention on engagements. If a position is later challenged, the ability to show what was asked, what came back and what the preparer did about it is the difference between a documented process and an assertion. Almost no firm retains this today.
- Marketing claims run past the same standard. Any accuracy, savings or audit-risk claim about your automated review needs substantiation you could produce on request, and no claim may promise a result.
Frequently Asked Questions
Our AI only does classification and data entry, not positions. Is that lower risk?
Lower, but not low, and the boundary is softer than it sounds. Classification decisions are positions wearing different clothes: whether an expenditure is repair or improvement, whether a worker is an employee or a contractor, whether an activity is passive, whether a cost is currently deductible or capitalised. Each of those is a legal characterisation that drives a number on a return, and a model performing it at volume is making thousands of small determinations nobody reviews individually because each looks like data entry. The useful test is not whether the tool writes prose — it is whether reversing any single output would change a line on the return. If it would, the output is a position, and it needs the same verification posture as one.
Can we rely on the AI vendor's accuracy claims as our due diligence?
No, for the same reason you cannot rely on a research publisher's claim that its treatise is correct. Due diligence is an act you perform, and reliance on the work of another is only available where you exercised reasonable care in engaging, supervising and evaluating that person — a framing that assumes a supervisable party. A vendor benchmark is a marketing artefact about aggregate behaviour on a test set; it says nothing about the specific return in front of you. Vendor quality is worth diligencing when selecting a tool, and it is worth documenting that you did. It is not a substitute for verification at the engagement level, and a firm that has substituted it has no evidence of review to produce.
What if the client used AI to prepare their own records before sending them to us?
This is the emerging version of a very old problem and it deserves an explicit intake question. Preparers may generally rely in good faith on information furnished by a taxpayer without independently verifying it, but that reliance ends where the information appears incorrect, incomplete or inconsistent, and it never extends to ignoring what you actually know. Client-side AI produces a characteristic pattern worth learning to spot: categorised books with suspiciously uniform classification, reconstructed records with implausibly round numbers, mileage or expense logs generated rather than kept, and summaries that no longer match source documents. Add a question to the organiser asking whether AI tools were used to produce the records, and treat a yes as a reason to sample source documents rather than accept the summary.
Does an AI assistant that answers client tax questions constitute practice before the IRS?
Practice generally means matters relating to a taxpayer's rights, privileges or liabilities before the tax authority — preparing and filing documents, corresponding, representing at conferences and meetings, and rendering written advice with respect to certain transactions. A public-facing bot answering general questions is closer to publishing than to representation, but the line moves fast once the tool becomes taxpayer-specific: intake, notice interpretation, correspondence drafting, or advice on how to respond to an examination all move toward the regulated core. The safer architecture separates the two clearly, keeps taxpayer-specific advice behind a practitioner, and does not let the interface imply the bot is your firm speaking. An unlicensed answer that a client reasonably attributes to the firm is the firm's answer.
State boards regulate CPAs separately. Does any of this reach them?
Yes, and firms tend to run the federal analysis and stop. State accountancy boards impose their own competence, supervision, confidentiality and advertising rules, and several have taken an interest in technology-assisted work. Independence and professional-standards questions arrive from a third direction: if the same platform both generates entries and produces the analysis used in an attest engagement, that is a self-review concern independent of anything the tax practice rules say. Multi-state firms should also assume the answers are not uniform, because the boards are not uniform. Build the federal posture first, then check the states where the firm and its individual licensees are registered, and treat the strictest as the operating rule.
How do preparer penalties interact with the client's own penalties here?
They are separate assessments that can both land on the same position, and the AI angle affects each differently. Taxpayer-level accuracy penalties can be mitigated by reasonable cause and good faith, of which reliance on a competent professional adviser is the classic instance — but that reliance argument depends on the adviser having been given complete information and having actually exercised judgement. If the position came from a model the preparer did not verify, the story that supports the client's defence is precisely the story that supports the preparer's penalty. Preparer-level penalties turn on the support for the position and on the preparer's knowledge, with a reasonable-cause and good-faith exception of their own that verification evidence feeds directly. Both roads run through the same artefact: proof of what was checked, by whom, before signing.
The Question That Reveals the Gap
Pick one filed return from last season that an assistant touched. Ask who verified the authority behind its largest discretionary position, on what date, and where that verification is recorded.
If the answer is that the reviewer would have caught anything wrong, you have a review culture rather than a review record — and a review record is the only part of this that helps you after a notice arrives.
Related Reading
- AI tools and attorney-client privilege — the parallel problem in the other advice profession, where the artefact at risk is the protection rather than the position.
- The AI vendor procurement questionnaire — the diligence questions that determine whether a tool can lawfully receive regulated client data at all.
- Shadow AI and unapproved employee tools — why the busy-season paste into a consumer chatbot is a structural problem, not a training problem.