Evaluating AI Correspondence Tools Against Real Tax Firm Regulatory Writing Standards
AI-drafted tax correspondence must clear Circular 230, not just pass a vendor demo.

AI correspondence tools can draft client memos, penalty abatement requests, and IRS notice responses for you, and the leading platforms ground every citation in a traceable primary-authority database, not generated text. But you need to ask whether the output survives scrutiny under Circular 230, IRS Alert 2026-19, and a firm's own sign-off process, standards that existed well before generative AI arrived and have not been loosened to accommodate it.
Most firms still evaluate these tools the way they evaluate practice-management software: a demo, a price comparison, a speed test. That approach misses what makes a tool usable in a tax practice. Mapping the actual regulatory standards and using them as evaluation criteria gives a firm something more durable than a vendor's sales pitch: a test any tool must pass before it touches a client file.
What Circular 230 requires when AI drafts the correspondence
Circular 230 does not contain a carve-out for AI-assisted work, and nothing in Alert 2026-19 creates one. These obligations apply with identical force whether a human drafted every sentence or a language model produced the first version. Four provisions matter most directly when AI is doing the drafting.
IRC Sections 6713 and 7216 govern confidentiality, and they are implicated the moment client return data goes into a platform that isn't built to protect it. If you upload return data to unsecured or public AI tools, you risk violation, because that data may be retained, processed, or exposed to other users in ways you never authorized. That obligation isn't confined to the licensed preparer who signs the return. It runs to every staff member who touches client data, including administrative staff who might paste a client's financial details into a free chatbot to save time on a routine letter.
Alert 2026-19 closes a specific loophole before firms can fall into it: vendor language calling a product "secure," "closed," or "enterprise" is not self-certifying. A marketing label does not answer whether a tool's use satisfies Circular 230, IRC § 6713, or IRC § 7216. That determination requires looking at what the tool actually does with client data, not what its product page says it does. None of this represents a new burden grafted onto tax practice. It is the same confidentiality obligation practitioners have carried for years, now applied to a workflow that happens to involve a model.
The AICPA and FTC standards that sit alongside Circular 230
Circular 230 is the most visible regulatory layer, but it is not the only one a tool has to clear. Other professional standards and data-security rules sit alongside it, each with its own requirements, so a tool that satisfies one but not the others still leaves a firm exposed.
SSTS Section 1.4 requires members to exercise due professional care even when a tool is doing the drafting, and the standard is explicit that using an AI tool does not absolve the member of professional obligations under AICPA or other applicable ethical standards. SSTS ¶1.4.6 adds a distinction with real consequences for how firms should treat different categories of AI tool: subscription-based tax research tools and resources may carry more weight than articles pulled from independent internet sources. A general-purpose chatbot trained on open internet text sits closer to the independent-source end of that spectrum than to the subscription tax-research end, so if you build workflows around AI, you need to weigh that distinction and not treat all AI output as equivalent.
The FTC Safeguards Rule requires every tax professional to maintain a Written Information Security Plan, and updated IRS Publication 5708 resolved a prior ambiguity in that requirement: multi-factor authentication is now required for all users accessing systems containing customer information, regardless of whether they're connecting from inside the office or remotely. The security architecture needs to be in place first, and the tool evaluated against it, not the other way around.
AI hallucinations as a regulatory problem, not just a quality problem
Fabricated citations are not a drafting imperfection that a careful editor smooths over before the letter goes out. They are the specific mechanism by which AI-drafted correspondence turns into a Circular 230 violation, a court sanction, or a disciplinary referral. The causal chain is mechanical: a model hallucinates a citation, a practitioner fails to independently verify it under the Section 10.22 due diligence standard, the fabricated authority goes out under the firm's name, and the consequences follow from there.
The record of what happens next is no longer theoretical. In one Tax Court case, the U.S. Tax Court saw Judge Ron Buch strike a petitioner's pretrial memorandum in 2024 after discovering citations bearing the hallmarks of AI-generated hallucination. In Tax Court Order 2026-16, Judge Mark Holmes confronted fictitious case citations in a brief and wrote that "submitting a brief with fictitious caselaw is a recipe for sanctions and a clear violation of Rule 11(b)," adding that the attorney's argument "collapses like an overmixed soufflé when one looks at the citations used to prop it up." Alert 2026-19 itself points to two other incidents as its cautionary examples, Whiting v. City of Athens and the 2025 Deloitte Australia matter, in which Deloitte agreed to partially refund a government report after reviewers found references to nonexistent academic papers and a fabricated quote attributed to a federal court judgment. A parallel incident surfaced at EY Canada in May 2026, when GPTZero published an investigation finding that most citations in an EY Canada report on loyalty-program safeguards were hallucinated; the Financial Times later reported that EY withdrew the study, which contained fake footnotes, fabricated data, and a citation to a McKinsey report that did not exist.
These are not edge cases involving careless practitioners at disreputable firms. Under IRC Section 6662 and Treas. Reg. Section 1.6662-4(d)(3)(iii), "substantial authority" for a tax position can only be built from recognized sources, and a citation invented by a general AI model carries no defensive weight whatsoever. Hallucination, in this context, is a regulatory exposure measured in sanctions and censure.
Source architecture and whether a tool can meet the standard
The single most important technical variable separating a defensible AI correspondence tool from a dangerous one is where its output comes from: a curated database of primary legal authority, or a prediction of likely text sequences drawn from broad training data. That distinction maps directly onto Circular 230 compliance risk, and it explains why two tools that look similar in a demo can produce wildly different exposure for the firm that adopts one of them.
General-purpose large language models, including consumer tools like ChatGPT, Claude, and Microsoft Copilot, carry a categorically different risk profile for tax correspondence. Asked to produce a memo on a specific IRC section, they generate text that reads fluently like a memo while potentially citing Treasury Regulations that do not exist, inventing revenue rulings, and presenting fabricated case law with total confidence. Alert 2026-19 states this directly: these systems predict likely text sequences rather than retrieving factually verified information, they carry a training data cutoff date that may not reflect current law, and the confidence with which a model presents an answer does not correlate with whether that answer is accurate.
Purpose-built tax research tools operate on a different principle. When one of these tools cites a Treasury Regulation, that citation traces to the actual regulatory text sitting in its database, and a practitioner can click through to verify it directly. Defensible correspondence depends on that traceability existing.
A fair objection to this framework deserves acknowledgment: tool category alone does not guarantee a good outcome. Even the most rigorously built, citation-grounded tool does not eliminate the Section 10.22 obligation for independent verification. That objection is correct, but it does not erase the underlying distinction. Source architecture sets the floor for how much independent verification a piece of correspondence requires before it can go out the door. If a tool is anchored in primary authority, it demands less remediation work from the practitioner reviewing it than a tool generating unanchored text from broad training data, and a firm's evaluation process needs to measure exactly that difference.
Evaluating AI correspondence tools against the regulatory criteria
With the regulatory framework established, a firm can assess any candidate tool across five concrete axes: the nature of its source database (primary authority versus generated text), whether its citations are traceable back to source documents, how its confidentiality architecture handles client data under IRC §§ 6713 and 7216, whether its security certifications fit inside the firm's WISP, and how much independent verification it still leaves to the practitioner under Section 10.22. Measured against those axes, the tools available to tax firms today differ meaningfully.
Marble is purpose-built for tax practitioners rather than retrofitted from generic accounting software, and its design is oriented around automating the routine backend of a tax engagement, including client intake, document review, and compliance checks, so that practitioners spend their time on the advisory judgment calls Circular 230 actually demands of them. Marble handles document triage and intake automation instead of generating unanchored correspondence pulled from broad training data, which narrows the surface area where hallucination risk can enter a firm's workflow. It fits best for firms that want to satisfy the Section 10.36 supervisory obligation by building AI automation into documented, auditable procedures rather than leaving individual practitioners to experiment with consumer tools on their own.
TaxGPT is SOC 2 Type II certified, with automatic PII redaction and encryption applied to client data. Its agents work on that data to complete assigned tasks but do not use it to train the underlying model, so it addresses the IRC §§ 6713/7216 confidentiality threshold directly. Its AI Tax Writer drafts IRS correspondence, client memos, penalty abatement requests, and engagement letters, with citations pulled from the IRC, Treasury Regulations, court decisions, and official IRS guidance. A companion product, TaxGPT Cowork, deploys agents across tax preparation, bookkeeping, payroll, and advisory workflows inside a firm's existing software stack, with practitioner review built into each step, which supports the Section 10.36 supervisory-procedures requirement. Built-in hallucination controls and a cited-answers-only policy cover part of the Section 10.22 verification burden, but you still have to review each citation independently. If a firm wants one AI writing and research layer with documented security credentials attached, this suits it.
Blue J applies machine learning trained on tax case law, IRS rulings, and regulatory materials to produce predictive tax research and outcome analysis, drawing on thousands of cases and rulings, with one-click memo generation available at preferred AICPA/CPA.com pricing. Its outcome-prediction capability is particularly suited to controversy work and cross-border matters, where Section 10.37's requirement to consider contrary authority is hardest to satisfy manually.
The fee disclosure and billing obligations Alert 2026-19 adds to the evaluation
Circular 230 Section 10.27(a) prohibits charging an unconscionable fee, and OPR Alert 2026-19 extends that prohibition squarely into AI use. The alert warns that billing clients for time the practitioner didn't actually spend, without adjusting for AI-driven time savings, may itself constitute an unconscionable fee, and it advises that cost savings generated by AI should be passed on to clients, with AI use disclosed.
That has a direct bearing on how a firm evaluates a tool, not just whether it should adopt one. No single uniform federal rule forces practitioners to disclose AI use to clients today, but over 300 federal judges and dozens of districts have issued standing orders or local rules that do just that. So many clients have no reliable way to know whether AI touched their correspondence. A firm evaluating a correspondence tool needs to ask not only if it produces compliant drafts, but also what its adoption means for engagement-letter language, fee schedules, and the billing review process that catches discrepancies before they reach a client.
Building a firm-level evaluation protocol that uses these standards as the actual test
If a firm has mapped the standards above, it can run a structured evaluation whose real test is whether sample outputs survive scrutiny under each standard, not whether a vendor's marketing claims match what the firm hopes is true.
A workable protocol has six components. First, a source architecture audit: ask each vendor to demonstrate, for a specific IRC section, how a citation in its output traces back to the primary authority in its database; a vendor that cannot do this fails the Section 10.37 threshold before any other criterion matters. Third, supervisory procedure mapping against Section 10.36: identify every point in the tool's workflow where a practitioner reviews and approves AI-generated content, and document that process in the firm's written procedures before the tool goes live. Fourth, a disclaimer and written-advice standard check against Section 10.37's requirements, an area where an SSRN study evaluating AI outputs against a Circular 230 compliance rubric found the weakest performance across the board, with only 26.7 percent of outputs rated fully compliant. Fifth, a fee and billing policy established in writing before deployment, so the Section 10.27(a) obligation is satisfied before the first AI-assisted engagement closes. Sixth, update the WISP to confirm MFA is implemented for every user who accesses systems containing customer information, as updated Publication 5708 requires, and name the tool itself in the firm's written plan.
The goal of this protocol was never to find a perfect tool. Marble's underlying design premise, automating the repetitive intake and document-triage work of a tax engagement while keeping practitioners in the verification and sign-off role, lines up with what a Section 10.36-compliant procedure actually looks like in practice: automation handles the volume, and the practitioner's judgment remains the thing the client, and the regulator, is ultimately relying on.


