A businessman in a suit works on a laptop at a dark wood desk in a modern glass-walled office, with a city skyline visible through floor-to-ceiling windows at sunset.
No items found.
Blog

Guardrails, not gimmicks: why assurance work needs governed AI, not a general-purpose LLM

The five gaps in generic AI that no prompt can close and what governed AI does instead

General-purpose AI can likely pass a CPA exam question. It can't tell you whether your engagement file is ready for sign-off. And it certainly can't be part of your evidence trail. Here's why the difference matters more than the demo.

Somewhere in your firm right now, a second-year associate has a general-purpose chatbot open in another tab.

They're not being reckless. They're being efficient. They've got a lease agreement that needs summarising, a disclosure they can't quite place, and a manager asking for a first draft by lunchtime. The tool is fast, it's articulate, and it's free.

It also has no idea what engagement they're on, no idea what their materiality threshold is, no memory of the risks the team logged three weeks ago, no view of the prior-year control environment, and, critically, no obligation to tell them when it's guessing. Whatever it produces will land in a workpaper with no source, no trail, and no record of who checked it.

That's the real AI risk in assurance: governance, not capability.

Capability was never the bottleneck

Let's dispense with the idea that frontier models aren't smart enough for professional work. They clear graduate-level benchmarks. They reason through multi-step problems. If the question is raw capability, that argument ended a while ago.

Clearing a benchmark and being fit to inform a professional conclusion are different things, and the distance between them is where most AI in audit quietly fails. Caseware's own VP of AI and Methodology, a former standard-setter, regulator and inspector, framed it about as bluntly as it can be framed: a baseline model with no tuning will sometimes do an OK job, sometimes a great job, and sometimes a terrible one. The inconsistency is the problem, more than the quality of any single output.

Auditors are trained from day one to be sceptical. You cannot build a review process around occasional brilliance. You need something you can calibrate for.

The five things a generic LLM structurally cannot give you

This is a description of what general-purpose models are built for, not a knock on how well they do it. A general-purpose assistant sits outside your engagement, and that architectural fact produces five gaps that no amount of prompt engineering closes.

  1. It has no engagement context. Ask a generic model to identify risks from a trial balance and you'll get something plausible: inventory is up 11%, margins have compressed. Both true, both useless on their own, because the model is missing the synthesis a competent auditor performs almost automatically: connecting that movement to the board minutes, the prior-year control environment, the specific risk profile of this client at this point in time. Caseware Verity, an agentic AI layer built into Caseware Cloud, reasons across the actual engagement: trial balance data, financial statements, logged risks, controls, materiality, engagement documents, checklists and your firm's own knowledge bases, before it says anything.
  1. It doesn't know how standards connect. Public training data contains a great deal about auditing standards. What it doesn't reliably contain is the interconnective judgment that makes them function. Take SAS 145: a dense web of requirements that reference one another constantly. Very few standards operate in isolation, and that interconnectivity is often invisible in a base model's training. Caseware Verity's domain intelligence is delivered deliberately, through structured methodology and instruction sets that encode standards interconnectivity and firm-specific methodology into the platform itself.
  1. It can't cite itself. This is the one that should end the debate for anyone who has sat through an inspection. Every AI-assisted output in Caseware Verity is citation-backed and traceable before it enters the engagement file. You can see the engagement data, the standard, or the firm guidance that informed it, and check the source before you act on it. A chatbot's confident paragraph sometimes has no provenance at all. An unsourced assertion isn't audit evidence, it's an opinion you've inherited from a stranger.
  1. It has no concept of your permission model. Your engagement file has roles, restrictions and segregation of duties for good reasons. Caseware Verity operates inside those firm-configured permissions and controls, with engagement-level data isolation. A browser tab respects nothing, because it was never told anything existed.
  1. It changes underneath you without telling you. Public models get silently updated. Behaviour drifts. The prompt that worked in March returns something different in July. Caseware runs a continuous evaluation framework that re-runs assessments as a regression test each time a new model is released, measures what changed and adjusts, precisely so firms can build a dependable review process on top. In beta, Caseware Verity hit 94% accuracy against a golden dataset drawn from real usage traces, with indicative manual workflows dropping from 15 to 20 minutes to under two.

The guardrail that matters most: the human stays in the loop by design

Here is the design principle that should reassure every partner who has to put their name on an opinion.

Nothing gets written into the engagement file without a human reviewing and accepting it. That's built into the product itself, not a policy staff have to remember to follow. Caseware Verity surfaces context, proposes risks, drafts suggestions, flags gaps and provides source references. The reviewer decides, refines, accepts or overrides. "Human judgment is sacrosanct," Caseware's team said at CwX 2026.

Caseware's stated test for defensible AI in a regulated environment is that outputs must be authorised, bounded, traceable and explainable. It's worth holding your current AI usage against that four-part test:

Authorised: is the tool operating within permissions your firm actually configured? Bounded: does it work within a defined task and a defined methodology, or can it wander anywhere? Traceable: can you show where an output came from, months later, to someone who isn't inclined to take your word for it? Explainable: can the reviewer understand the reasoning well enough to accept responsibility for it?

A general-purpose chatbot can fail all four, sometimes by a large margin.

This is also why the agentic suites are built as specific agents rather than one all-purpose oracle: a Disclosure Checklist Agent that generates citation-backed suggestions you review, refine, accept or override inside the workflow; a Document Intelligence Agent for extracting from source documents into workpapers; a Risk Suggestion Agent that proposes engagement-specific risks with supporting rationale. Narrow scope is a feature. It's what makes the output reviewable in the first place.

Security and confidentiality: the part your clients care about

Your client's confidential financial information does not belong in a consumer chatbot's input box, and the professional confidentiality obligations you signed up to don't have a "but it was quicker" exemption.

Caseware Verity is built into Caseware Cloud, so it inherits an enterprise security posture rather than bolting one on: SOC 2 Type II and ISO 27001 certification, AES-256 encryption at rest and TLS in transit with managed key rotation, multi-factor authentication and role-based access controls, engagement-level data isolation, logical separation of client data, defined data-residency regions, continuous monitoring with endpoint detection and centralised logging, annual independent audits, penetration testing and dual peer review of code changes. Prompts, documents and client data are not used to train the underlying large language models.

Compare that to the honest answer most firms would have to give a client who asked what happened to their information last busy season: we're not entirely sure.

There's a broader point here. Giving your people a secure, approved, engagement-aware tool is one of the most effective ways to eliminate shadow AI. Staff don't reach for unsanctioned tools out of carelessness. They reach for them because nothing sanctioned is good enough yet. Fix that second problem and the first one largely disappears.

The uncomfortable question

Recent research suggests AI is now effectively core to US audit and accounting firms, and that the conversation has shifted from adoption to control. That reframing is the whole ballgame. The question in front of you isn't whether your firm uses AI. It already does, whether or not you've approved it.

The question is whether that usage is governed, documented and defensible, or whether you'll find out how it was being used during your next inspection.

Caseware Verity exists because the answer to that question shouldn't be left to chance, or to whichever tab happens to be open.

See what governed AI actually looks like inside an engagement. Request a Caseware Verity demo →

No items found.

Latest news and insights.

Explore expert perspectives to help your firm stay ahead.

From Technology Adoption to Business Value: What Accounting Firms Need to Get Right

Lead your firm to accuracy, efficiency, and growth with Caseware.

The authority in AI-powered audit.

Contact Us
Caseware Launches Verity AI Agents for Assurance
Discover how Verity brings workflow native AI agents directly into assurance engagements to enhance efficiency, transparency, and professional judgment.
Read the IDC study

By registering, you agree to our Privacy Policy.