AI in Financial Reporting

It has been a while since I posted, but I have been getting a lot of questions lately from clients and from my accounting students about the same topic: 

Where does AI actually belong in the financial reporting process? 

Everyone wants to know if they should be using it, and the honest answer is that it depends a lot on which part of the process you are talking about.

I have been working with several clients on this over the last few months, plus I have an upcoming speaking session on AI in reporting, so I thought I would put some of these thoughts together here. For each point, I am also going to push back on myself a little, because I think that is a more useful way to think through this than just picking a side.

Where AI is actually helping right now

The clearest win I have seen is variance analysis. AI tools are good at scanning through a large number of GL lines and flagging the ones that look different from what history would predict. Think of a vendor payment that is 40% higher than normal, or an account that usually nets close to zero but did not this month. This does not replace the accountant figuring out why the variance happened, but it does replace a lot of the manual scanning that used to take up time during the close.

Pushback: threshold alerts and standard deviation bands have done this job for two decades without needing AI. The real bottleneck was never spotting the variance, it was investigating it, and AI does not investigate for you. So the time saved may be smaller than it sounds.

Narrative drafting is another good use case. MD&A sections, board commentary, variance explanations built off the numbers in the system, these are a good fit for AI to draft a first version. The accountant still owns the final language and the numbers behind it, but starting from a draft instead of a blank page saves real time.

Pushback: once a draft exists, people tend to anchor to it instead of rethinking it from scratch. If the AI frames a variance a certain way, the human editing it may not catch a wrong framing as easily as they would if they had written it themselves.

Document and data extraction is a third one. Pulling structured data out of contracts and invoices, for example when you are working through a revenue recognition analysis under ASC 606, is something AI handles well. We are already seeing this show up inside Copilot for Dynamics 365 Finance.

Pushback: contract language around variable consideration and performance obligations is exactly where extraction tools still make errors, and those errors are silent unless someone independently rereads the contract. This is the one use case where I would want a stated error rate before trusting it.

Reconciliations are also a good fit, especially the fuzzy matching kind. Bank recs and intercompany recs where the transaction descriptions do not line up exactly are a good example. AI models are better at "this is probably the same transaction described a little differently" than the old rules-based matching was.

Pushback: fuzzy matching algorithms have existed for years and already handle this well. If AI tools are better, that needs to be shown, not assumed, and I would be cautious about trading a matching method you can fully explain for one that is more of a black box.

Where I would be more careful

The audit trail is the first thing I think about. If an AI tool suggests a journal entry, a reserve estimate, or drafts language for a disclosure, can you show an auditor exactly what data went into that suggestion and how it got there? A lot of organizations that are turning these tools on have not thought through how they are going to document AI-assisted judgment the same way they would document a person's judgment.

Pushback: to be fair, human judgment is not always well documented either. Nobody logs a controller's full reasoning for picking an estimate. A well-built AI tool might actually log its inputs more completely than a person ever does. The real problem may be documentation discipline in general, not something unique to AI.

The second thing is relying too much on pattern matching for the calls that require real judgment. AI is good at finding patterns in historical data. It is not as good at knowing when this quarter is different from history, like a one-time acquisition or a new revenue arrangement with no precedent. The risk is not that the tool gets it wrong, it is that a preparer trusts a generated explanation without applying the same skepticism they would apply to a junior staff member's work.

Pushback: that is really a statement about the preparer, not the tool. People can trust a spreadsheet formula or a colleague's number just as blindly. Blaming the AI tool lets weak review standards off the hook.

The third is what I would call stacked vendor risk. If your reporting process runs through D365, plus a Copilot layer, plus a separate analytics tool with its own AI features, you now have several tools contributing to one reported number, each one a bit of a black box.

Pushback: finance has always run on stacks of black boxes. Nobody fully audits Excel's internal math or the ERP's posting engine either. What actually makes AI different is that its output is probabilistic instead of deterministic, and that distinction is worth spelling out rather than just calling it a black box.

The last one is governance lagging behind adoption. I see this constantly with Copilot features in D365. They get turned on because they came with the platform you already pay for, not because IT, internal audit, and controllership sat down and agreed on how the feature should be used and monitored.

Pushback: this happens with every new feature, not just AI ones. New Excel functions, new Power BI visuals, new D365 workspaces all get adopted the same ad hoc way. The question worth asking is whether AI actually deserves a higher governance bar than other features, or if we are just noticing the general problem because AI is the current headline.

How I am framing it

When I look at a new AI use case in reporting, I ask three questions. Is this task pattern detection or judgment? Can I reconstruct how the output was produced? Who approved turning this on?

Pushback: the pattern versus judgment split is cleaner in theory than in practice. Deciding which variances are even worth flagging is itself a judgment call, it just got pushed into the tool's design instead of the accountant's head. A framework like this can create false confidence that something is safely automated when the judgment simply moved upstream.

Summary

AI is already useful in specific parts of the reporting process, and it is going to keep showing up in more places as Microsoft keeps investing in Copilot across D365 and Power Platform. My honest take after arguing both sides here is that none of this is really new. We asked the same questions about RPA, cloud ERP, and even spreadsheet macros a decade ago. The organizations that will do well with AI are the ones matching the use case to the right amount of human oversight and building the documentation trail before an auditor asks for it, not after.

I will be talking through more of this, including where I disagree with myself, at an upcoming session on AI in reporting at the Solver Ascend conference in Nashville, TN. More Info Hope to see you there.

Comments

Popular posts from this blog

Showing the last 12 months based on a date slicer

Copilot Dynamics 365 Finance & Supply

RPA with Dynamics 365 Finance & Supply Chain