Short answer

For most Australian accounting firms, Microsoft 365 Copilot is the default because it lives inside the Microsoft tenancy you already govern. ChatGPT and Claude are stronger as standalone reasoning and drafting tools. The decision is driven less by model quality than by where your documents live, what your engagement letter permits, and who administers it.

Firms usually ask this question expecting a capability comparison. Capability is rarely the deciding factor at this point. All three are competent at the tasks an accounting practice would use them for, and the gaps between them close every few months.

What does not close is the governance difference. That is what this article is actually about.

The comparison at a glance

Microsoft 365 CopilotChatGPTClaude
VendorMicrosoftOpenAIAnthropic
Sits insideYour M365 tenancyStandalone, with connectorsStandalone, with connectors
Sees your firm's documentsYes, within tenant permissionsOnly what you connect or pasteOnly what you connect or paste
Business tier trains on your contentNoNoNo
Consumer tier trains on your contentN/ADepends on settingsDepends on settings
Admin controlsMature, via M365 adminAvailable on business tiersAvailable on Team and Enterprise
Typical firm useDocuments, email, Teams, ExcelDrafting, analysis, researchLong-document work, drafting, analysis
Australian data residencyAvailable via M365 configurationCheck current termsAvailable on some enterprise arrangements

The row that matters most is "consumer tier trains on your content". Across all three vendors, the free and personal tiers handle data differently from the business tiers, and staff using personal accounts is the single most common compliance exposure we find in Australian firms.

Why Copilot is usually the default

Not because it is the best model. Because of where it sits.

If your firm runs Microsoft 365, your documents, email and Teams conversations are already inside a tenancy you control, with permissions you have already configured. Copilot operates within those permissions rather than requiring you to move data anywhere new.

That has three practical consequences.

The diligence is shorter. You have already assessed Microsoft as a provider. Adding Copilot extends an existing relationship rather than introducing a new third party.

Data residency is a configuration question. Microsoft offers data residency options that most Australian firms have already set up for their core tenancy.

Administration is where your IT already is. Whoever manages your M365 environment manages this. For firms without dedicated IT, that matters more than it sounds.

Audit your permissions first

Copilot inherits your permission model, including its faults. If your SharePoint permissions are loose, and in a lot of firms they are, Copilot will happily surface documents to staff who technically had access but never went looking. Firms should audit permissions before enabling it, not after.

Where ChatGPT and Claude are stronger

Two areas, consistently.

Long-document work. Reading a lengthy agreement, a set of financial statements or a technical ruling and producing a structured summary or comparison. Standalone tools tend to handle large volumes of pasted or uploaded text more comfortably than an assistant embedded in an Office application.

Open-ended reasoning and drafting. Working through a problem where the shape of the answer is not known in advance. Explaining a technical position several different ways for different audiences. Producing a first draft of something structurally new.

Between the two, firms tend to report Claude as stronger on long-document analysis and careful writing, and ChatGPT as stronger on breadth of integrations and general familiarity. That is a soft preference and you should test both against your own work rather than take it from an article.

The data handling question, properly

This is where firms need to be precise, because vendor terms differ by tier and change.

The general shape across all three vendors is the same. Business and enterprise tiers operate under commercial terms where customer content is not used to train models, and come with a data processing agreement, administrative controls and configurable retention. Consumer tiers operate under different terms, and depending on the provider and your settings, inputs may be retained and used for model improvement.

Anthropic's commercial terms cover Claude Team and Enterprise and the API, with no training on customer content. Enterprise adds SAML and SCIM, audit logs, custom retention and data residency options. Consumer plans on Free, Pro and Max have a user-controlled setting for whether conversations contribute to training.

Microsoft's position for M365 Copilot is that tenant data stays within the tenancy boundary and is not used to train foundation models.

OpenAI's business and enterprise tiers similarly do not train on customer content by default, while the consumer product's position depends on account settings.

What to do with this: get the answer in writing from the vendor for the specific plan you are buying, and record it. Do not rely on this article, or any article, for a claim you would need to stand behind with a client or the TPB. Terms change. TPB(GS) 55/2026 expects you to complete an appropriate review of commercial AI tools yourself.

The five questions to put to any of them are in AI tools for Australian accounting firms in 2026.

What none of them are good at

Australian tax technical work.

All three are trained overwhelmingly on international material. They will answer a question about Division 7A, small business CGT concessions or trust streaming fluently, confidently and sometimes wrongly. The confidence is the problem, because it does not degrade as accuracy does.

Under Code items 9 and 10 you must take reasonable care to ascertain a client's state of affairs and to ensure the taxation laws are applied correctly. TPB(GS) 55/2026 is direct that AI output is not a substitute for your own analysis. Treat every technical output from any of these three as an unverified assertion.

Use them for how to explain a position. Do not use them to determine the position.

Do you need to pick just one?

No, and most firms end up with more than one for sensible reasons.

A common pattern in firms we work with: Copilot as the default because it is in the tenancy and covers document and email work, plus one standalone tool with a small number of licences for the people doing heavier analysis and drafting.

What you should not have is three tools, no policy, and staff choosing between them on personal accounts. Pick a default, name it in your policy, and give people a legitimate option for the work the default does not handle. Our AI use policy template has a table for exactly this.

A note on building on top of these

If your firm goes beyond chat and starts building actual automations, the platform question changes shape. The underlying models from all three vendors are available through APIs, and a well-built automation should not be locked to one of them.

We build agents that run across Claude, ChatGPT and Microsoft 365 Copilot, and we do that deliberately. Model quality rankings change every few months. A workflow built so that swapping the underlying model is a configuration change rather than a rebuild ages considerably better than one wired to whichever model was best in the quarter it was built.

If you are evaluating this for your practice, our automation audit covers which platform fits your stack and where the compliance boundaries sit.

Frequently asked questions

For most Australian firms on Microsoft 365, Copilot is the practical default because it works within a tenancy you already govern. ChatGPT and Claude are stronger for long-document analysis and open-ended drafting. Many firms run Copilot plus one of the others.
It operates within your Microsoft tenancy under permissions you control, which is a better starting position than a standalone tool. You still need client permission under Code item 6 and you should audit your SharePoint permissions before enabling it.
On business and enterprise tiers, no. On consumer tiers it depends on your account settings. This is the distinction that matters most, and it is why staff on personal accounts is the exposure to fix first.
Firms tend to report Claude as stronger on long-document analysis and careful writing. That is a preference, not a benchmark. Test both on your own work before deciding.
Not reliably. All three are trained largely on international material and produce confident errors on Australian tax law. Use them to explain a position you have already established, not to determine one.
You need a tier governed by commercial terms rather than consumer terms. For most firms that means the business tier rather than the top enterprise tier, unless you need SSO, audit logs or data residency guarantees.

Related reading: can Australian tax agents use ChatGPT with client data and AI tools for Australian accounting firms in 2026.

Vendor terms and model capabilities change frequently. Verify current data handling terms directly with each provider for the specific plan you are considering. Last verified August 2026.