De-identification of client data before AI — why it usually doesn't work
Removing the name is not de-identification. Removing the name is removing the name.
Published 26 July 2026
One of the more common workarounds firms adopt for AI use is asking staff to “de-identify” client data before pasting it into an AI service. The theory is that if the information is no longer about an identifiable individual, the privacy analysis changes and the disclosure becomes acceptable.
This is well-intentioned and it usually does not work. This paper explains why.
What de-identification actually means
The Privacy Act defines personal information as information about an identified individual, or an individual who is reasonably identifiable. De-identification is the process of removing or altering information so that the individual is no longer reasonably identifiable — either directly, or through combination with other available information.
That “reasonably identifiable” test is doing an enormous amount of work.
An individual is reasonably identifiable if the information could be used to identify them by someone with access to relevant other information, either now or in the future. It does not require identification to be certain, or easy, or by everyone — only that a reasonable person with reasonable access to other data could work it out.
The bar is much higher than most people assume.
Why “removing the name” is not de-identification
The single most common de-identification practice is to replace names with initials, letters or a generic placeholder. This is almost never sufficient, for four reasons.
1. Direct identifiers are more than names
Names are one direct identifier. Others include:
- Dates of birth
- Addresses
- Phone numbers and email addresses
- Employer or workplace
- Vehicle registration
- Medicare numbers, TFNs, ABNs, drivers’ licences, passport numbers
- Distinctive titles (“the CEO of…”, “the founder of…”)
- Uncommon dates that would appear in a professional context
Removing the client’s name from a paragraph that still contains the address, date of birth and reference to their employer has removed the name and nothing else.
2. Quasi-identifiers combine
A quasi-identifier is a piece of information that does not identify someone on its own but does so in combination with others.
Postcode + gender + date of birth uniquely identifies most Australians. That is not an exaggeration — it is a well-documented result from population re-identification research.
A “de-identified” summary that says “our client, a 47-year-old male director of a listed company based in Toorak, seeking advice about…” has narrowed the population to something in the low double digits. Add a mention of the specific industry or dispute and it is a single person.
3. The context around the query re-identifies
Even where the substantive text is genuinely stripped of identifiers, the context in which the query occurs is not.
If you routinely ask an AI service questions about clients, and one of your queries relates to a particular kind of matter that only one of your clients is currently in, the query itself narrows down who you are talking about — to anyone with visibility of both the AI service data and your client list.
For a firm handling a small number of high-profile matters, this re-identification-by-context is close to unavoidable regardless of the text content.
4. Free-text material contains identifying detail everywhere
Attempting to de-identify a substantial document by hand is much harder than it looks. Names appear in headers, footers, footnotes, signature blocks, embedded metadata, exhibit references, court numbers, addresses of properties in dispute, transaction reference numbers, dates that anchor a specific event. Missing one is easy; missing several across a long document is close to inevitable.
If a lawyer or clinician is doing the de-identification manually while trying to save time by using AI in the first place, they are being asked to do the more painstaking task specifically so they can do the easier task safely. The reason people paste raw documents is that this is exactly the extra work they were trying to avoid.
The specific problem for professional firms
Two structural features of professional practice make this worse.
Small client populations. Most firms have far fewer clients than they think of themselves as having. A “confidential” reference to “our client in the pharmaceutical space” or “our high-net-worth client with a family trust structure” narrows down to a small set of people rapidly — sometimes to one.
Novel or distinctive facts. The matters that are most interesting to summarise with AI are often the ones with distinctive facts. Those distinctive facts are the identifiers. A garden-variety dispute is genuinely hard to identify from a de-identified summary; a novel and unusual one is not.
Where de-identification does work
To be fair, there are situations where genuine de-identification is achievable and appropriate.
- Large aggregated datasets with statistical de-identification techniques (k-anonymity, differential privacy) — used in genuine research and epidemiology, not in a paste-into-chatbot workflow.
- Genuinely generic queries — “what does the case law say about X” — where the query has no client-specific fact pattern at all.
- Public information — where you are querying about publicly available material, no de-identification is needed because there is no personal information at stake.
If your AI use is confined to those categories, de-identification is not an issue. But if that is genuinely your AI use, you probably do not need to solve the AI-confidentiality problem in the first place.
The regulator’s view
The OAIC has issued guidance making clear that de-identification is a process, not a status. Information described as “de-identified” but not actually de-identified is still personal information for the purposes of the Privacy Act, and the entity holding it is subject to the APPs.
That means calling data “de-identified” in your AI policy does not make it so. The test is whether the individual is reasonably identifiable from what is actually transmitted, and the answer is usually yes.
The alternative that actually works
The reason de-identification is popular is that it appears to preserve the productivity benefit of AI without the compliance cost. If it actually worked at scale, it would be the right answer.
Because it does not, the honest alternatives are two:
Use AI on genuinely non-personal work only, with the personal work handled without AI. This is defensible and constrains the risk cleanly, at the cost of losing the AI benefit for the highest-value use.
Use AI on personal information in an environment where the transmission problem doesn’t arise. An on-premise system doesn’t require de-identification, because nothing is disclosed. Staff can use the AI on the real, identifiable content and get the full productivity benefit, because there is no third-party recipient to protect the data from.
The second option is more expensive up front and removes the analysis rather than complying with it.
The one exception to insist on
Regardless of which route you take, one form of de-identification is worth maintaining: do not include information about identifiable individuals other than the client in AI queries.
The counterparty, the witness, the neighbour, the tenant, the beneficiary who has not consented to your firm’s AI use — none of them signed up for your AI workflow, and they have privacy rights you must protect even if your own client has consented to something. Redact them, and redact them properly, whichever AI tool you use.
That specific redaction is achievable, defensible, and worth building into the workflow. General de-identification of the client’s own file, sadly, is neither of the first two of those things.
This article is general information about common obligations under Australian privacy and professional conduct rules. It is not legal, medical or financial advice and does not account for your circumstances. Obtain your own advice before acting on it.