Est.
FeaturesLong read

How AI Accounting Software Actually Works for Tax Firms

Break AI tax software into its four layers to see what you're actually buying.

Contributing Editor · · 13 min read
Cover illustration for “How AI Accounting Software Actually Works for Tax Firms”
Features · September 15, 2026 · 13 min read · 2,865 words

AI adoption among tax firms jumped from 9% in 2024 to 41% in 2025, according to Wolters Kluwer's Future Ready Accountant report. That's a profession rewiring its core workflow inside a single filing season, and most practitioners doing the adopting still can't say what they're actually buying. Here's the useful move: break the technology into its layers, and the picture gets clearer fast. Once you can see the layers, you can tell the difference between a feature list and a system that holds up under review, and you stop paying for marketing copy dressed up as engineering.

Sentiment is already ahead of deployment, and the gap deserves sitting with. Thomson Reuters Institute's 2025 Generative AI in Professional Services Report found that 68% of tax and accounting professionals feel excited or hopeful about generative AI, but 52% of firms using it lean on general tools like ChatGPT, and only 17% use software built specifically for tax work. Most of the profession is improvising with a consumer chatbot instead of running a system built for statute, citation, and audit trail. That's a mistake, not a neutral choice, because a general-purpose chatbot was never trained on the specific statute or jurisdiction a return depends on. Document triage, data entry, and compliance checking eat the hours that could go toward advisory work. A firm that doesn't know what a tool is doing under the hood can't tell whether it's closing that gap or just moving the risk somewhere less visible.

The four-layer stack that most AI tax tools are built on

Vendors sell this software as a feature list: smart extraction, auto-categorization, "AI-powered insights." That framing hides more than it shows. Strip the marketing language off any of these products and four distinct layers make up their architecture: document ingestion, pattern recognition, compliance logic, and workflow automation.

Each layer does a different job, needs different training data, and fails in its own way. Ingestion breaks on a bad scan. Pattern recognition breaks on a client situation it's never seen before. Compliance logic breaks when nobody updates it after a law changes. Workflow automation breaks when its exception-handling rules are too narrow for the case sitting in front of it. Mixing up these failure modes is how a firm ends up trusting a tool in the wrong place, or distrusting it in the right one, and most vendor demos are built specifically to blur that line.

Products bundle all four into a single interface, which is exactly why the layers are hard to see individually. At the low end sit narrow machine learning classifiers, trained to do one job like sorting document types or flagging an outlier. At the high end sits agentic orchestration, where a system starts and coordinates multi-step work across an entire engagement on its own. No firm needs every layer running at full depth to get value out of a tool. What matters is knowing which layers a given product actually implements, and at what depth, rather than what a slick demo implies.

Diagram: The Four-Layer Stack Inside AI Tax Software. Visualizes: Visualize the four distinct architectural layers that make up AI tax tools, showing each layer's job and its characteristic failure mode.

How document ingestion works: from client upload to structured data

Client documents don't arrive in any consistent shape. A W-2 comes in as a PDF from a payroll portal. A 1099 shows up as a phone photo of a printed page. A K-1 lands as a scanned fax with a coffee ring on it. Document intelligence systems handle this by pairing optical character recognition with classification models, and before any data gets pulled, the system first has to figure out what kind of document it's even looking at.

Once the type is identified, extraction logic goes to work locating specific fields: W-2 box values, 1099 interest or dividend amounts, K-1 allocation percentages, mapped to the correct line on the return. This is where the real labor savings show up. According to Filed.com, manual data entry runs an error rate around 10%, while machine-learning-driven extraction brings that down to roughly 1%. Review doesn't disappear, though. The bottleneck just moves. Instead of keying in numbers, the preparer validates numbers the system already pulled, which is a faster job and a genuinely different one.

Some systems add adaptive organization on top: trained on a firm's own filing history, they learn what a "normal" document set looks like for a given client type and flag anything that deviates, a missing schedule, an unfamiliar form. Multi-layered validation adds another check, cross-referencing figures across supporting attachments so a mismatch gets caught immediately instead of surfacing three weeks later when a reviewer happens to notice it.

None of this is magic, and it breaks in specific, predictable places. A low-resolution scan, handwritten margin notes, or a document format the model has never seen will degrade extraction quality fast. What this layer hands off, when it works, is structured, validated data ready for the next layer. Nothing more.

What pattern recognition actually does inside a tax workflow

Pattern recognition is where the machine learning does its most visible work: expense categorization, transaction classification, account reconciliation. These models train on large sets of labeled prior returns and transactions, and the core task always reduces to the same question: does this new item resemble something the model already learned to recognize?

The mechanic that makes this layer improve over time is correction feedback. Every time a preparer overrides a categorization, that correction becomes a training signal. Season over season, the classifications sharpen for that specific firm's client base, and review time shrinks with it. Anomaly detection runs the same logic in reverse: instead of matching a transaction to a known category, it flags line items that deviate from the historical norm for that client or industry.

Audit risk scoring is a related but separate application. Models trained on IRS audit trigger patterns assign a risk signal to line items on a completed return and recommend the preparer gather supporting documentation before filing. Blue J applies this kind of predictive modeling further upstream: a practitioner enters the facts of a case, and the model predicts likely outcomes based on existing statute and judicial precedent, effectively simulating how a court or the IRS might treat the situation.

The limitation that matters most here is structural. Pattern recognition is interpolation. It performs well on situations resembling what it's already seen, and it gets shakier as the facts drift from that training data. A client with an unusual asset structure, a new business line, or a cross-border arrangement generates more flags, not because the software is broken, but because there's less precedent to lean on. That's the mathematical nature of the tool, not a temporary gap in coverage, and no amount of additional training data fully closes it for a genuinely novel fact pattern.

How compliance logic is embedded, and what keeps it current

Compliance is where tax software stops being generic machine learning and becomes tax software specifically. Compliance logic depends on more than the data-driven learning categorization uses. It's coded directly from statute, regulation, and administrative guidance, then layered on top of whatever pattern recognition produces.

Two approaches dominate in practice, and they trade off against each other in a fairly obvious way. Rules-based engines run on explicit if/then logic pulled straight from the tax code: predictable, auditable, but dependent on someone manually updating the rules every time the law changes. The other approach is the fine-tuned large language model, trained specifically on tax law rather than general text. Sphere's proprietary TRAM (Tax Review and Assessment Model) indexes and interprets tax law to determine taxability across states, countries, and other jurisdictions. TaxGPT trains on the specific nuances of US tax law, including state department of revenue content and IRS materials, rather than a general web corpus.

A live compliance function has emerged on top of both approaches: regulatory monitoring, where the system tracks legal changes and alerts a firm to how a new rule affects the client profiles already on file. That's a real shift from reactive to proactive, and it lines up with where the profession says it actually wants help. Thomson Reuters' 2026 AI in Professional Services Report found that 69% of tax firm respondents named tax research as the leading potential AI use case, ahead of return preparation itself at 57%. Tracking the law eats more hours than most practitioners like to admit.

Bloomberg Tax's AI Assistant, recognized in CPA Practice Advisor's 2025 Technology Innovation Award, delivers citation-backed answers through chat-based, cross-jurisdictional search. Thomson Reuters' Checkpoint Edge AI integrates with UltraTax CS, connecting research directly to the return workflow. The distinction that matters most in this layer, more than any other, is between a model fine-tuned on current, jurisdiction-specific law and a general-purpose LLM asked to answer a tax question. A generic chatbot produces a confident, citation-shaped answer that's simply wrong, because nothing in its training ever pinned down the actual statute or jurisdiction in play. Pick the wrong one here and the gap becomes genuine risk. It's the difference between a defensible position and a malpractice exposure.

Workflow automation: how the layers connect into an end-to-end process

Put the four layers together and a full engagement looks like this: a client uploads documents, ingestion extracts and validates the data, pattern recognition classifies transactions and flags anomalies, compliance logic checks the result against current law, and the preparer gets a review-ready draft with every exception already surfaced.

According to OpenLedger, most organizations now rely on AI to manage 60 to 70% of routine tax workflows, with accountants shifting into an oversight role, handling exceptions and supplying the judgment a machine can't. That's a redistribution of labor, not a replacement of it, and firms that market it as the latter are overselling their own product.

The current frontier is agentic AI, and the term deserves a precise definition instead of buzzword treatment. Earlier automation waited for a human to give it an instruction. Agentic systems start actions on their own, monitor conditions as they change, and push work forward within a defined set of rules, with nobody clicking "go" at every step. Wolters Kluwer lays out a practical hierarchy for these systems. "Taskers" automate narrow, repetitive jobs like document classification. "Automators" run a full process end-to-end, flowing categorized transactions straight into a trial balance. "Collaborators" offer guidance at complex routing decisions. "Orchestrators" sit above all of them, coordinating multiple agents toward one outcome.

Thomson Reuters introduced CoCounsel as an agentic system for tax, audit, and accounting professionals, built for multi-step execution across those workflows rather than answering one question at a time. Thomson Reuters added further agentic capability in November 2025, and early users have reported measurable time savings on straightforward returns. That's a modest number on any single return, but it compounds fast across a season's volume. Gartner projects that 40% of enterprise applications will embed AI agents by the end of 2026, up from under 5% in 2025, a curve steep enough that firms evaluating tools today should assume the ground keeps shifting under them.

Accounts payable offers a clean, non-tax example of what full agentic automation looks like: an AI agent verifies vendor details, routes an invoice through the correct approval path, and schedules payment without a human touching it, cutting processing time by around 75% and escalating to a person only when something falls outside its rules. That's the model tax workflows are converging toward. The preparer's role shifts from doing the work to reviewing it: checking flagged exceptions, applying judgment to the edge cases the system couldn't resolve, and signing off before anything goes out the door. That sign-off isn't optional, and it isn't a temporary artifact of immature technology either. Human oversight is legally required, and even the most advanced automation layer produces a review-ready draft, never a filed return.

What the major platforms have built and where they differ

The Tax Adviser's 2026 Tax Software Survey names four platforms as dominant in professional tax prep: Wolters Kluwer (ATX, CCH Axcess Tax, CCH ProSystem fx), Intuit (Lacerte, ProSeries), Thomson Reuters (UltraTax CS), and Drake Software (Drake Tax). All four have spent the past two years layering AI on top of platforms that, in some cases, have been the industry standard for decades. They are not converging on the same product, though, and the differences matter more than the shared branding suggests.

Thomson Reuters has moved fastest on the agentic side. CoCounsel launched as embedded agentic AI for tax, audit, and accounting workflows. In October 2025, the company announced three AI additions to its ONESOURCE line: AI-powered Indirect Tax Product Classification, Tax Regulatory Insights for Direct Tax, and Global Classification AI for trade customers. CoCounsel picked up further agentic capability in November 2025, and Checkpoint Edge AI's integration with UltraTax CS connects research to the return workflow without requiring a separate tool.

Wolters Kluwer has pushed on a parallel track instead of chasing the same lead. CCH iFirm was upgraded in July 2025 with an AI-powered virtual agent, a cloud-based workflow dashboard for workpapers, and integration with a tax news and research app. TaxWise Online, powered by Expert AI, was announced on November 25, 2025, aimed at tax preparers and Electronic Return Originators, with error reduction and firm growth as the stated goals. CCH Axcess Expert AI extends that same intelligence across the broader CCH Axcess suite, tying tax, audit, and firm management into one workflow. Wolters Kluwer's own Future Ready Accountant report found that 73% of regular AI users report performance better than expected, and 77% of firms plan to raise AI spending by 2028. Those are numbers that suggest the investment case is closing.

Where these platforms actually differ comes down to a few concrete questions: how deeply the research layer sits inside the return-preparation interface itself, whether the compliance model is purpose-built on tax-specific data or adapted from a general LLM, and how mature the agentic capability is versus a simple assistant that answers a question but doesn't act on it. On compliance and research specifically, purpose-built tax tools, trained on tax law and wired directly into the return workflow, beat general-purpose tools bolted on after the fact, and that isn't close. It follows directly from how the compliance layer works: a model has to train on the actual statute to apply it correctly, and a system added on afterward rarely gets that training right.

What AI tax tools genuinely cannot do, and where practitioner judgment is irreplaceable

Pattern recognition fails predictably on novelty. A client with an unusual asset structure, a new business line, or exposure across multiple jurisdictions generates more exceptions precisely because the model has seen fewer examples like it before. The next model update leaves that unresolved. It's inherent to how the layer works, which means the situations most in need of expert judgment are exactly the ones where the AI is least reliable.

Compliance logic goes stale without active maintenance, whether it's a rules engine or a fine-tuned model. Whether a tool stays current with the most recent law changes is a critical factor buried on a features page. It's one of the most important questions a firm can ask a vendor before signing a contract, and vendors that dodge it are telling you something.

Agentic systems only operate inside the rules they've been given. When a situation falls outside those parameters, the system escalates instead of deciding, and the quality of that escalation depends entirely on how well the firm configured its exception handling in the first place. A poorly configured system either escalates too much, burying the preparer in noise, or too little, letting a real problem slip through as a false negative.

The advisory layer isn't automated at all, and shouldn't be mistaken for something on a near-term roadmap. Predictive analytics can surface an opportunity, a planning strategy, a risk to flag, but interpreting that signal for a specific client's actual circumstances, explaining it in a way the client trusts, and making a judgment call under real uncertainty stay entirely human. Thomson Reuters' 2026 AI in Professional Services Report found that only 14% of tax firms are using specifically agentic AI so far. Most firms evaluating these tools today are assessing capability that hasn't fully matured in production yet.

Given all that, a firm evaluating any AI tax product does better asking a short set of concrete questions than reading a features page. Which layers does the tool actually implement, and at what depth, versus what it just claims? Is the compliance and research layer trained on current, jurisdiction-specific tax law, or adapted from a general-purpose model? Does the system surface exceptions clearly and prominently, or bury them three clicks deep? Does it plug into the return preparation software already in use, or demand a second, parallel workflow? Does the model actually get sharper from this specific firm's corrections, or does it stay static out of the box?

Understanding the stack doesn't make the technology less useful. It makes the evaluation honest. A practitioner who knows what each layer does, and doesn't do, can ask a vendor sharper questions, set client expectations that won't need walking back later, and spend their own hours where the software genuinely can't go, on the judgment calls, the client conversations, and the advisory work no amount of pattern matching will replace.

Sources

  1. How are different accounting firms using AI in 2025?
  2. Artificial Intelligence (AI) in Tax and Accounting | Wolters Kluwer
  3. Using AI to Win Tax Season
  4. tax.thomsonreuters.com
  5. dualentry.com
  6. filed.com

More in Features