Automating Compliance Exception Flagging in Tax Reviews
Machine learning catches tax errors before filing by spotting patterns humans miss.

Tax compliance in the United States runs $546 billion a year, and most of that isn't legal fees or software licenses. It's time. The IRS expects Americans to spend 7.1 billion hours on filing and reporting in 2025, which works out to 3.4 million people doing nothing else, all year, but paperwork. Automated exception flagging, catching anomalies and mismatches in tax data before they turn into filing errors, is how a growing slice of that time gets clawed back. Here's how the technology works, what it actually catches, what it means that the IRS itself is now using it, and what a firm needs to have in place before any of this is worth trusting.
What compliance exception flagging actually means in a tax review context
An exception is any transaction, entry, or document that breaks from an expected pattern, threshold, or rule badly enough to need a human look before the return goes out. That's really the whole definition. The word itself doesn't pass judgment; it just means something's worth a second glance.
Three kinds of exceptions get treated differently downstream, so they're worth pulling apart. Anomalies are statistical outliers: a withholding rate applied inconsistently across transactions that otherwise look the same. Inconsistencies are mismatches between systems, like a TIN on a withholding certificate that doesn't match what's sitting in the payee database. Risk signals are the slippery ones. They don't mean an error happened; they correlate with exposure down the road. A client edging toward an economic nexus threshold in a new state hasn't broken any rule yet, but the flag fires anyway because that pattern reliably predicts trouble later.
Flagging isn't resolving. The system's job ends at the surface; a practitioner, or an approved correction workflow, picks it up from there. Catching any of these three categories by hand requires someone who both knows what to look for and has the hours to look. In most practices, those two things rarely show up together. Automated flagging just separates the two jobs cleanly: the machine catches, the person decides.
How automated exception flagging systems work under the hood
Older rules engines only fire on conditions someone thought to write down. They miss anything novel and need constant retuning every time a jurisdiction changes a rule, which, given how often that happens, is often. The shift underway now layers machine learning on top of those rules: the rule set handles the deterministic stuff, and the ML layer picks up the anomalies nobody bothered to code for. Thomson Reuters calls this generation agentic, meaning the system validates data, flags missing fields, and flags outliers on its own, rather than sitting there waiting for someone to run a query.
The workflow looks pretty similar across vendors once you dig into it. Data comes in from ERPs, invoicing platforms, exemption certificate repositories, payee databases, structured and unstructured alike. Each transaction gets sorted: reportable or not, covered or noncovered, taxable or exempt. From there the system checks payee data against certificates on file, validates TINs, and checks withholding rates against treaty schedules. An anomaly-scoring layer scores how far something deviates from historical norms, and flagged items land in a queue with context attached: what got flagged, why, and what the resolution options look like.
The accuracy numbers hold up better than I expected, honestly. A 2025 peer-reviewed study running 3,232 tax records across manufacturing and services pitted Random Forest, XGBoost, and SVM against each other. Random Forest won, at 92.00% accuracy in manufacturing and 93.39% in services, while the other two sat in the 85 to 90% range. That's a classification result from a controlled study, not a sales pitch, and it says ML-based risk prediction has cleared a real bar rather than staying a proof of concept somebody points to in a deck.
Explainability isn't a nice-to-have here. Systems like ONESOURCE pair AI output with human controls specifically so the result can be audited: every flag needs a traceable reason, not a bare number with nothing behind it. This matters more in tax than in almost any other domain, because the U.S. runs more than 12,000 taxing jurisdictions, each drifting its own rules over time. A rules-only engine snaps under that kind of fragmentation. What makes ML worth the trouble is that it generalizes across scenarios nobody explicitly wrote a rule for, which is the difference between automation being viable at scale and automation just being a convenience.
What automated exception flagging actually catches: four high-frequency use cases
Exemption certificate gaps show up constantly in B2B sales tax review, probably more than anything else. Certificates go missing, expire, or get attached to the wrong transaction, and the exposure grows with every new customer a business adds. Avalara uses AI to digitize, validate, and store these certificates, replacing what used to be somebody's spreadsheet. Kintsugi, which took a $15 million strategic investment from Vertex in April 2025, builds certificate validation straight into its compliance layer. The flag reads simply enough: no certificate on file for a customer with a history of exempt sales, or one on file that had already expired by the transaction date.
Withholding certificate validation is a cousin problem, not the same one. Specialized AI tools read withholding certificates in a range of formats, extract key data points such as income codes and rate information, and check all of it against the applicable withholding rules. The flags tend to be pointed: a TIN that doesn't match the payee database, a claimed treaty rate that doesn't line up with the payee's actual country of residence, an income code that doesn't match the payment type.
Economic nexus detection is its own animal, mostly because the rules won't sit still. States have moved nexus thresholds and guidance repeatedly since Wayfair, faster than any manual tracking process can realistically keep pace with. Automated systems watch transaction volume by state close to real time, flag threshold crossings before they become unregistered obligations, and in some setups trigger the registration workflow directly. The flag: cumulative sales in a state closing in on, or crossing, the nexus threshold for the first time.
Digital asset reporting is the messiest of the four, and it rounds things out that way for a reason. Basis and acquisition date are often just unknown for noncovered lots coming out of self-custody wallets or non-reporting brokers, so automation has to take customer-supplied data at face value and flag the discrepancies rather than pretend the data set is complete. The IRS and Treasury keep issuing fresh guidance here, Notice 2025-23 being one of the more recent, and any system that doesn't get updated in step with it develops gaps that are silent rather than obvious. The flag: missing cost basis on a disposal, or a transaction type that doesn't match the reported holding period.
What ties all four together is high volume, rules that won't stop moving, and documents arriving in a dozen formats. That's precisely the combination manual review handles worst.
What the IRS's own AI adoption means for practitioners running exception-flagging programs
The IRS isn't watching from the sidelines here. As of June 2025 the agency had 126 active AI use cases running, up from 10 in August 2022, covering audit selection, fraud detection, identity verification, and compliance scoring. The IRS has moved to formally govern AI use in audit selection and exam support, spelling out what systems can do unsupervised, where mandatory human review sits, and what documentation has to travel with an AI-generated referral. This isn't a pilot program anymore; it's an institution writing the rulebook.
Enforcement revenue climbed 12% in the first five months of fiscal year 2026, and IRS leadership has pointed to technology-driven productivity as the reason. This happened while the workforce shrank hard: between January and May of 2025, headcount dropped 25%, from roughly 103,000 to 77,000. Put those two together and there's only one reading. The agency is leaning on AI-powered audit selection harder precisely because it has fewer people reviewing returns, not more.
For practitioners, the takeaway is blunt. If the IRS runs exception flagging against submitted returns to find anomalies, the practitioner who runs that same process before filing catches those anomalies first. Thomson Reuters research puts a number on skipping this step: manual compliance processes carry roughly three times the audit risk exposure of automated ones.
None of this means IRS AI is finished. As of June 2025, 61% of its use cases were still listed as "in development," and TIGTA found the agency hasn't yet built a feedback loop for feeding examination outcomes back into model refinement. The tooling is getting better fast, but it isn't catching everything yet, and practitioners shouldn't assume today's blind spots stay blind for long.
How much of the compliance cycle automation can realistically compress
Thomson Reuters' ONESOURCE reports compliance cycles compressing from 30 days to 11, errors cut by as much as 50%, and real-time anomaly detection pushing error rates under 1%. Take those numbers seriously. Also take them for what they are: vendor-reported figures from deployments that already had clean data pipelines and mature workflows before automation showed up. Most firms are not starting from that place.
Data infrastructure is usually the first wall a deployment slams into. Many senior finance executives name inadequate data infrastructure as their biggest barrier to adopting AI, and a stronger data foundation is widely cited as what would speed things up considerably. This makes sense: exception flagging is only as good as what feeds it. Bad or scattered input doesn't produce fewer flags, it produces more false ones, and that just shifts the wasted time from filing over to triage.
The return on investment is uneven too. Gartner reports that strong AI returns in finance functions remain uncommon even with adoption climbing steadily. That gap has more to do with implementation than with whether the technology works at all. It tracks with the academic research: the 3,232-record study that clocked low-90s accuracy for Random Forest was working off labeled training data in a research setting. Production accuracy at an actual firm rides on how well the model gets trained against that firm's own transaction history, not on some ceiling baked into the technology itself.
So the honest version, for anyone evaluating this: real time savings and real error reduction are on the table, but the ceiling is set by data quality and implementation care, not by any limit in the underlying tech.
What it takes to implement automated exception flagging in a tax practice
Four things have to be in place before flagging pays off, and skipping any one of them is where most deployments quietly stall out.
Data consolidation comes first, and it's the one everyone underestimates. Exception flagging needs a single, or at least reconcilable, view of transaction data across the ERP, invoicing system, certificate store, and payee database. Fragmented data produces fragmented flags, and fragmented flags don't earn anyone's trust.
Rule and threshold configuration is next. Even the ML-augmented tools need tuning to a firm's specific jurisdictions, entity types, and filing calendars. Out-of-the-box rule sets handle the common cases fine; they miss the edge cases that actually make up a given firm's risk profile.
Exception workflow design gets skipped more than anything else on this list. A flag with no defined resolution path just turns into a queue nobody owns. Somebody has to decide in advance who reviews what, on what timeline, and how the resolution gets documented well enough to survive an audit later.
Human review integration closes the loop, and it's really just the model's core premise restated: touchless by default, manual by exception, only works if the exception queue gets actually reviewed instead of piling up in a corner. The human isn't a fallback tacked on at the end. It's a design requirement from day one.
Tool choice matters here as well. Generic accounting platforms can surface some anomalies, sure, but they weren't built around treaty schedules, nexus logic, or certificate validation, so they miss the tax-specific stuff by construction. Tax-specific platforms, tools built for tax practitioners from the ground up rather than bolted onto general accounting software tend to fold flagging directly into the existing review workflow instead of adding a separate layer on top. The difference shows up in onboarding time, in how well jurisdiction-specific edge cases get covered, and in how naturally exceptions surface inside a workflow people already use.
Adoption still lags the technology's maturity, for what it's worth. Gartner's 2025 survey of 183 CFOs and senior finance leaders found error and anomaly detection adopted by 34% of AI-using finance functions. Meaningful, but still a minority. Gartner projects 90% of finance teams will run at least one AI-enabled technology within two years. Exception flagging is moving from edge to baseline, and the firms running it now are just early.
What practitioners need to keep doing themselves once exception flagging is automated
Automated flagging redirects attention. It doesn't remove the need for it. The system surfaces the exception; a person still has to decide whether it's a data entry slip, a real compliance gap, or a perfectly legitimate transaction the model happened to misclassify. That call takes actual tax expertise, and no queue makes it by itself.
Three things stay a human job no matter how good the tooling gets. Ambiguous exceptions, novel transaction structures, cross-border arrangements, guidance that changed last month and hasn't been coded into the rules yet, need someone to interpret them, not rubber-stamp them. Model calibration is another: if the same exceptions keep surfacing, or keep getting waved through the same way every time, that's a sign the configuration needs adjusting, and somebody has to be watching for it. Client communication matters most of the three. When a flagged exception has real filing or advisory consequences for a client, the conversation that follows is advisory work. It was never compliance processing, and automating the flag doesn't change that.
The Thomson Reuters Institute's survey of indirect tax professionals found 89% naming filing accuracy as their top success metric, 63% prioritizing minimized penalties, and 51% wanting to cut time spent on compliance generally. Automated exception flagging can serve all three, but only if the hours it frees up actually go toward the advisory and review work those goals were about in the first place. The technology moves where the hours land. Where they go from there is still on the practitioner, same as it always was.


