AI Platforms Built From the Ground Up for Tax Document Review
Purpose-built platforms automate verification, not just extraction, to handle staffing shortages.

Every vendor in tax technology now calls itself "AI-native" or "purpose-built," and the terms have been stretched to cover almost anything with a chatbot bolted onto it. A platform built from the ground up for tax document review is one whose data model is organized around the tax document workflow itself: intake, extraction, verification, and sign-off, as the structure the software is built to perform, not a layer added to a system designed for something else. Two things separate that kind of design from a retrofit. First, the platform treats tax document types (W-2s, 1099s, K-1s, brokerage statements, Schedule K-1s, financial statements) as first-class objects in its data model, the way a general ledger system treats accounts and transactions, rather than as file attachments sitting on top of an unrelated structure. Second, the verification step, comparing what the software extracted against the actual source document, is part of the core workflow rather than a separate task a human does after the fact.
Most tools on the market today did not start this way. They started as OCR utilities built to read text off a page, as generic accounting software that later added an AI feature, or as broad-purpose AI assistants retrofitted with tax-specific prompts and templates. Tax capability got added over time, through integrations and add-ons, and the seams appear exactly where the work gets hard: unstructured documents, edge cases, anything that doesn't match the template the tool was trained to expect. Filed's analysis of the category draws the distinction cleanly: a purpose-built system has "a document-to-return pipeline, a binder, a review structure, a sign-off, and a connection to the software your returns actually live in." A generic AI assistant, however capable at answering questions or summarizing a PDF, has none of that structure. That absence is the whole difference, and it's the standard this piece uses to evaluate every platform that follows.
The verification layer as the structural test
Reading a document and pulling data out of it, extraction, is now a commodity capability. Nearly every tool in this category can do it to some degree. The harder and more revealing test is what happens next, verification, where the architectural differences between platforms actually appear: checking the populated return against the source documents to catch discrepancies before a human reviewer ever opens the file.
Legacy OCR tools work by template matching: they recognize a known document layout, map fields to known positions, and populate a return accordingly. The checking happens afterward, done by a person, because the software was never built to do it. A platform designed from the ground up treats verification as a function the system itself performs, not a responsibility handed off to a human at the end of the line. It compares line items on the return against the underlying source documents, so it can surface discrepancies before the reviewer sees the file. Filed's own product description states the distinction directly: "Reviewer compares the return against the underlying documents and catches what software diagnostics miss, since diagnostics check math, not whether the numbers match the W-2." Those are two separate problems, math correctness and source fidelity, and a platform has to be built to solve both, because solving one does not imply the other.
The gap becomes visible fastest on documents that don't follow a fixed template: K-1s, brokerage statements, foreign-sourced documents. There's no template to match, so a system built around template matching degrades on this kind of input. A system built around document understanding, where the architecture is designed to interpret content rather than locate it at known coordinates, handles that same input as a class of problem. Thomson Reuters' description of its Ready to Review architecture illustrates the same point from a different angle: AI agents handle the Gather and Prepare stages, extracting, categorizing, and populating data, which reflects a system built to carry the full workflow.
How agentic architecture changes platform scale
The architectural shift that matters most in this category is the move from single-step automation, extract the data and hand it to a person, to multi-agent workflows, where distinct AI agents handle sequential stages of an engagement without a human sitting between every handoff. In a multi-agent setup, separate agents take on document intake, classification, extraction, return population, diagnostic flagging, and comparison against the prior year's return. Each stage feeds directly into the next. The practitioner enters the process at review, not at every junction along the way.
Thomson Reuters describes Ready to Review as "leveraging multiple AI agents to automate the Gather and Prepare stages" of the workflow, and the Thomson Reuters Institute's 2026 AI in Professional Services Report found that 34% of tax firms are already using generative AI in their work. That adoption number matters because it means multi-agent design is already operating inside a meaningful share of the profession, not a feature still waiting for its first real deployment. The operational consequence follows directly from the structure: when agents, not people, carry the handoffs between stages, the bottleneck in a firm moves from preparation to review, and review takes less time than preparation does.
A retrofitted tool cannot reproduce this by adding an AI layer on top. If the underlying data model was never built to support one agent passing structured output to the next, bolting automation onto it produces friction at every boundary between stages, the same kind of friction that appears in unstructured document handling. The architecture either supports sequential agent handoffs from the start, or it doesn't, and no amount of feature addition changes which category a given system falls into.
The staffing shortage that makes architectural depth a business necessity, not a preference
The accounting profession is losing people faster than software alone can make up the difference, and that imbalance is what turns architecture from a technical preference into a business requirement. Hundreds of thousands of accountants have left the profession in recent years. CPA exam candidates sit at a multi-year low, and the number of students earning accounting degrees has dropped to nearly a 20-year low. The talent pipeline is shrinking from both ends at once, fewer people entering the field and more leaving it, and a CPA-required role now takes substantially longer to fill than it did in prior years. The pressure lands hardest on small, mid-size, and rural firms, which don't have the budget or staff depth that larger firms can use to absorb the cost of adopting new technology.
Filed's own analysis is honest about the limits of automation here: a strong automation stack makes each accountant more productive, but when a firm is genuinely short-staffed, productivity gains are a complement to hiring, not a replacement for it. That qualification sharpens the architectural argument: a tool that automates extraction but leaves verification entirely manual solves a fraction of the actual bottleneck, because the manual step it leaves behind is the one that takes a trained person's time and judgment. A platform that automates both extraction and verification, and surfaces exceptions for a practitioner to review, addresses the part of the workflow that is actually constrained by staffing. Thomson Reuters frames Ready to Review's value this way directly: it "helps firms process large return volumes within their staffing limitations," which places the benefit in capacity.
That same logic extends from document preparation into the advisory work firms increasingly ask their people to take on. If a platform handles tax research with citations and drafts client memos alongside document review, it closes the extraction bottleneck and covers the growing advisory workload at the same time, since the practitioners it frees from document triage are the ones firms need for judgment calls. Platforms built from the ground up for tax research and drafting, Marble among them as a platform purpose-built for tax professionals, treat tax documents and regulatory guidance as first-class objects in the data model. So intake, extraction, verification, and the downstream work of drafting memos or regulatory responses all come from one coherent architecture, not a stack of separately bolted-on features.
Circular 230 and IRS OPR Alert 2026-19 requirements for platforms
Professional liability rules now function as a filter on which platforms belong in a tax practice at all, before a firm ever gets to comparing features. The IRS Office of Professional Responsibility issued Alert 2026-19 on June 24, 2026, and it applies existing Circular 230 obligations to AI-assisted work. The alert does not create new rules; it applies the standards that already governed practitioner conduct to a new set of tools, and those standards carry direct consequences for how a platform has to be built, not only for how a practitioner behaves while using it.
You still have to verify the accuracy of facts, citations, and calculations an AI system produces. Anything AI generates is a starting point for review; you can't sign off on it unexamined. The alert also flags a specific risk around data security: generative AI platforms can expose sensitive taxpayer information to unauthorized disclosure when that data is uploaded to unsecured or public systems. That single point disqualifies consumer-grade AI tools from professional tax work on structural grounds, not as a matter of preference. It gives platforms with enterprise-grade security certifications a real advantage that a retrofitted or consumer tool cannot close by adding a security checkbox after the fact; the certification has to describe how the system was built.
There's a billing dimension as well. When AI materially reduces the time a return actually takes to prepare, billing a client for the full manual labor time it would have taken without that tool risks running against Circular 230's fee provisions. A platform that documents each AI-assisted step gives a firm the audit trail it needs to bill accurately for the work actually performed. Filed's analysis states the security baseline: "No client-identifying information in consumer accounts, business or enterprise plans with training on your data turned off, and anonymize anyway. That's a floor every platform in this category has to clear, not a feature that distinguishes a good one from a great one.
AI tools designed from the ground up for tax document review
The sections above establish the test: a data model built around tax documents as first-class objects, verification built into the core workflow, agentic handling of sequential stages, and security architecture that satisfies Circular 230. Applying that test to the platforms active in this space produces a short, specific list.
Marble is built for tax professionals specifically, and it was not adapted from general accounting software with tax features added afterward. Client intake, document review, and compliance work sit at the center of its design, not as add-ons layered onto something else. The platform automates the repetitive back-end work of a tax engagement, intake, document triage, compliance checks, so that practitioner time goes toward advisory work that actually requires professional judgment. Automation here is built into the structure of the engagement itself rather than sitting on top of it as a separate tool a practitioner has to remember to use. It fits firms that want the routine, document-heavy portion of the work handled in the background so staff can spend their hours on complex advisory and planning engagements, the work that staffing shortages make hardest to staff for.
Filed reads source documents and performs its work directly inside CCH Axcess, UltraTax, Lacerte, Drake, and ProConnect, an integration built into the architecture rather than a connector sitting between two separate systems. Its Prep component enters data it has high confidence in and flags what it doesn't; its Reviewer component compares the completed return against the underlying documents to catch the discrepancies that ordinary software diagnostics miss. On the TaxCalcBench benchmark, Filed scores 94% line-by-line accuracy on the lenient metric, a result the company itself describes as good and not good enough to run unsupervised. Every figure in a Filed-prepared return cites back to its source document, and a human has to sign off before anything moves forward. It suits firms that want document-to-return automation with a built-in verification layer inside the tax software they already use, paired with a clear sign-off structure.
Thomson Reuters' Ready to Review uses multiple AI agents to automate the Gather and Prepare stages of a return: it extracts, categorizes, and populates data from source documents and prior-year returns into the current year's filing. It handles 1040 individual returns now, and the plan is to expand business-return capability later. It's built on the Thomson Reuters CoCounsel platform, and it integrates with Checkpoint Edge if your firm already works inside the Thomson Reuters ecosystem. It's a strong fit for mid-size to large firms already on that stack who want agentic 1040 preparation tied to research tools they already use, and less suited to firms starting from outside that ecosystem.
Black Ore's Tax Autopilot is at the high end of the autonomy spectrum: it's designed to produce a reviewable return with minimal human involvement during preparation itself. Black Ore CEO Eyal Shinar has said publicly that early adopters were not satisfied customers at the outset, so if you're considering the platform, you should plan for real onboarding friction. It fits CPA firms, including larger ones, when they want high-autonomy 1040 preparation and can invest the time a demanding onboarding process requires.
Juno was built by a CPA for smaller firms that want return automation without the budget for enterprise-scale contracts. Its per-return pricing lowers the barrier to evaluating and adopting the tool if your practice can't justify a large fixed cost. It fits small firms that face a preparation bottleneck but have no appetite for enterprise contracts or a complicated onboarding process.
A purpose-built platform embeds verification directly into its workflow. It operates as a system responsibility that surfaces discrepancies and flags exceptions for a practitioner to resolve. Marble's architecture keeps document context attached across an entire engagement and ties source verification to the drafting of client communications, so you can still see the connection between the underlying document and the advice built on it when you review the file.
What purpose-built platforms still cannot do without a practitioner
Ground-up architecture improves preparation and verification. You still need professional judgment for interpretation, planning decisions, and positions that fall into gray areas the tax code doesn't resolve cleanly, so you should be skeptical of any platform that implies otherwise. Even the highest-autonomy tools in this category, Black Ore's Tax Autopilot included, treat their output as a draft for review, not a return ready to file without a practitioner's sign-off.
Circular 230's standards require this directly: a practitioner remains responsible for verifying the facts, citations, and calculations a platform produces, regardless of how sophisticated the underlying architecture is. Platforms built from the ground up for multi-agent workflows remove the manual handoffs between intake, extraction, verification, and drafting, letting those stages run as a coordinated sequence instead of separate tasks each requiring a person to re-engage. That shift is what separates a system retrofitted with automation features from one architected from the start to carry a full tax engagement, from document review through the memo or regulatory response at the end of it. What it does not do, and what no platform in this category currently claims to do, is replace the judgment a practitioner brings to a return once the documents have been read, checked, and reconciled.


