What ChatGPT Can And Can't Do With Your Taxes (And When It Becomes Dangerous)

What ChatGPT Can And Can't Do With Your Taxes (And When It Becomes Dangerous)

Gary Fong

Part 4 of a 5-part series on how AI changed the IRS audit landscape. Catch up on Part 1, Part 2, and Part 3 if you haven't.

"I just asked ChatGPT."

That's the sentence we keep hearing from self-employed people who tried to DIY their tax research this year. It's also the sentence that's going to put a meaningful number of returns into IRS correspondence audit queues over the next eighteen months.

Not because ChatGPT is bad. It isn't. As a tool for understanding what something means, it's remarkable.

The problem is what people do with it.


What ChatGPT actually does well with taxes.

Let's be honest about the upside. There's a real one.

Explaining concepts. "What's the difference between Section 179 and bonus depreciation?" — ChatGPT will give you a clearer, more digestible answer than most CPA-prepared FAQ pages.

Brainstorming categories. "What deductions might I be missing as a freelance graphic designer?" — it will surface ten plausible categories you can then investigate.

Plain-English translation. "What does Treas. Reg. §1.274-5T actually require for mileage logs?" — translating IRS-speak to actionable rules is a strength.

Document drafting. A mileage log template, a home office worksheet, a 1099 explanation letter — ChatGPT will produce a usable first draft in seconds.

Methodical walkthroughs. "Walk me through Schedule C line by line" — the AI is genuinely patient in a way that a $400/hour CPA at tax time isn't.

These are real wins. None of this is what we're warning about.


Where it goes wrong.

The trouble starts when "ChatGPT explained it" becomes "ChatGPT said I should claim it." Four specific places where reliance on a single AI shifts from helpful to dangerous.

1. Hallucinated citations.

Ask ChatGPT for the legal authority supporting a tax position and there's a measurable chance you'll get a citation that doesn't exist. The case name, the IRC section, the Treasury Regulation, the IRS publication — all confidently formatted, all sometimes invented. Recent academic research on LLM hallucination rates in legal citations puts the false-citation rate for general-purpose chatbots somewhere between 17% and 33% depending on the model and prompt.

If you cite a fake Treasury Regulation on your audit defense file, the examiner doesn't pause to admire your effort. They invalidate your defense on that point. And if the same pattern shows up elsewhere in your documentation, your credibility on the entire return takes a hit.

2. Stale knowledge.

Every AI model has a training cutoff. ChatGPT-4o's was October 2023. Claude Sonnet 4.5's is January 2025. Gemini 2.5 Pro's is January 2025. Tax law changes constantly — TCJA sunsets, IRC amendments, inflation adjustments to standard deductions and phase-outs, new revenue procedures every few months.

If you ask a model trained in October 2023 about the 2024 vehicle mileage rate, it may give you the 2023 number. It may not tell you it's giving you the 2023 number. The model doesn't always know what it doesn't know.

3. Confidence on edge cases.

Ask any single AI a hard question, you'll get a confident answer. Ask three AIs the same hard question, and you'll watch them disagree about 15-20% of the time. The disagreement itself is the signal: that's the edge case, that's where the answer isn't obvious, that's where you need a human professional.

A single AI doesn't surface the disagreement. It just picks one answer and presents it like it's settled. You don't know which 15-20% of its answers are in the disagreement zone — because nothing in its response tells you.

4. Agreement bias.

Phrase a tax question with a desired outcome embedded ("can I deduct this $4,500 in home office expenses?") and the model is more likely to find a way to say yes than it would be from a neutral framing ("evaluate whether this $4,500 home office deduction is supportable"). This isn't malice — it's a known property of how RLHF-trained chatbots respond to leading prompts.

When you're using AI to research aggressive tax positions, this bias works against you. The model nudges toward "yes." The IRS examiner nudges toward "no." Guess who wins on audit.


The fix isn't to stop using AI. It's to use three.

The single-AI failure modes have a common shape: a model says something confident, but you have no second source to check it against. The solution is structural, not philosophical.

Three-Pass cross-verification. Run every substantive finding through Claude, then ChatGPT, then Gemini — three models from three different labs, trained on different data with different cutoffs, with different post-training tuning. Lock the finding only when all three converge on the same answer with citations to the same authority.

This is the Pass Three discipline that AuditClaude is built around. The mechanics:

  • If all three AIs agree with the same supporting citation that you can verify on irs.gov or in the IRC: the finding is locked. High confidence.
  • If two of three agree and one objects: the objection is the signal. Investigate the disagreement. Usually the holdout has a real point — often a recent IRC change one model missed, or a different reading of an ambiguous reg.
  • If all three disagree or hedge: don't claim it. This is the territory where you absolutely need a CPA or EA before putting it on a return.

The output isn't a smarter chatbot. It's a filter — one that converts the AI's confident-but-occasionally-wrong output into a smaller pile of positions that have actually been stress-tested.


Why this matters for your audit defense file.

Pass Three isn't about being thorough for its own sake. It's about what happens when the IRS letter arrives.

Imagine an examiner looking at your audit defense documentation. Two scenarios:

Scenario A: "ChatGPT said I could claim this."

Examiner: nope.

Scenario B: "This deduction is supported by Treas. Reg. §1.183-2(b), confirmed by IRS Pub. 535 page 47, independently surfaced by Claude, ChatGPT, and Gemini using the relevant facts of my return. Here's the contemporaneous documentation. My CPA reviewed and signed the return."

Examiner: okay.

The difference between Scenario A and Scenario B isn't the answer — both might land on the same deduction. The difference is defensibility. Scenario B has a paper trail. Scenario A has an opinion from a chatbot.


What we're not saying.

We're not saying don't use ChatGPT. ChatGPT is fine. Brilliant, even, for the first 80% of tax research.

We're saying: don't use it alone. The last 20% — where edge cases live, where the IRS examiner is going to look, where confident-but-wrong answers do the most damage — is the part where a single AI fails you and a multi-AI consensus protects you.

That last 20% is the part AuditClaude's Three-Pass System is built to handle. Pass One finds. Pass Two attacks. Pass Three verifies — across Claude, ChatGPT, and Gemini, with every finding cross-checked before it makes it into your audit defense file.

Keep your CPA.

Same standing reminder: AuditClaude does not replace your CPA. Even three AIs converging on the same answer doesn't equal reasonable-cause defense under Treas. Reg. §1.6664-4 — that protection comes from a CPA or EA filing on your behalf. The Three-Pass System produces the file your CPA reviews. Your CPA still signs the return.

One AI is a guess. Three AIs is a position.
Your CPA is what makes it a defense.


AuditClaude — The Three-Pass System for self-employed Schedule C filers.

The methodology playbook, the Maya Parker case study (with every Pass Three cross-verification example), and twelve copy-paste prompts that take you from "what did I miss" to "here's the audit defense file — Claude, ChatGPT, and Gemini agree." Digital download.

Get AuditClaude — $97 →

Back to blog