Skip to main content

4 posts tagged with "Federal Court"

Discussion of U.S. federal AI-related cases.

View All Tags

It’s Not Just GIGO: Don’t Dunk on the Dakotas and the Data Was Not the Problem

· 13 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC

TL;DR

  • N.D./S.D. can mean “Northern District” and “Southern District” for federal cases, overlapping with “North Dakota” and “South Dakota” postal abbreviations; other states have possible issues here like MD and OK
    • “CA” (Court of Appeals? California?). I addressed this in my “Writing Help” game.
    • “SC” is a triple threat: South Carolina, Supreme Court, or Superior Court?
  • I looked at the data. Yes, there were some human errors, but it is clear that Charlotin has put in the work to create and maintain this valuable resource. That’s why I think people find it worthwhile to send him cases. He’s provided something of value and put in the time. Anyone who looks closely at the data can see that. People who churn out slop reports, on the other hand, are wasting your time. If it wasn’t worth their time to write, why would it be worth your time to read?
  • Once you account for relative population, the map is much more of a dog bites man story. To me, the more interesting part is the specifics of each case and the citation graph of cases that are frequently cited by later cases in these disciplinary decisions.
    • For example, I thought this may have been a mistakenly coded jurisdiction but it was not: Lowery Wilkinson Lowery, LLC, et al. v. State of Illinois, et al. (E.D. Oklahoma 2025). Instead, it’s an interesting attempt by the plaintiff to apply the “Indian Country” McGirt ruling and the plaintiff also misattributed that U.S. Supreme Court ruling to the Eastern District of Oklahoma Judge.
    • “There is no one righteous, not even one.” After data cleanup, New Hampshire alone seemingly had no hallucination cases, but I found both that New Hampshire had had its own hallucination case around July 2025 per reporting last fall, and that a New Hampshire attorney had been responsible for hallucinations (actually from the client) in a case in the Vermont Supreme Court.

“N.D.” and “S.D.”: Did Not (Usually) Mean the Dakotas

Damien Charlotin maintains a well-known public database of AI hallucination cases around the world. It has even been cited in a legal case here in Iowa. I’ve written about his work before, including his point about low false-positive rates potentially creating automation bias, lulling users of legal AI tools into a false sense of security. Recently I saw a post on LinkedIn, in which Charlotin, reacting to a report by Straight Arrow News on a study by “Laine AI,” criticized the claim that AI-driven legal hallucinations are especially concentrated in North Dakota and South Dakota. Charlotin pointed out that this was most likely due to a misinterpretation of “N.D.” and “S.D.”—common abbreviations in federal cases indicating “Northern District” and “Southern District.”

The June 19, 2026 article noted how surprising the result was, with the headline: “Where does the most AI legal slop come from? Not from the states with the most lawyers.” Laine AI, according to the article, stated that North Dakota courts had 109 cases with AI errors, South Dakota had 82, and California had 59. Supposedly, this was from Charlotin’s website, but as I often advise clients, this is not a good use of LLMs: pivot tables already exist!

info

I previously viewed the page on Laine’s own website with the statements about the Dakotas. It appears that as of July 2, 2026, they’ve corrected the numbers and they now reflect a more realistic breakdown of the states and have reworded the conclusions. I cannot speak to the accuracy of the revised version.

After I manually reviewed and coded the states for 1,145 USA cases and cleaned up data from Charlotin’s website downloaded on June 21, 2026, I found that the actual count was: California 121, North Dakota 6, and South Dakota 1. California had some “E.D. Cal” and similar, as well as two Ninth Circuit cases that were appeals from California cases, a case that was “CA SC” (Court of Appeals, South Carolina? California Superior or Supreme Court?) which was actually California; and one “CA 8th Circuit (Bankruptcy)” a federal Eighth Circuit case that was an appeal from District of North Dakota Bankruptcy.

None of these minor ambiguities would explain the large discrepancy in counts. My review of the numbers suggests that Charlotin was correct about the likely cause of the hallucinations.

You can read Charlotin’s own account of the mix-up on LinkedIn.

tip

As a general observation, LLMs can be very helpful for speeding up human review of typos and spelling variants in data like this (e.g., “E.D. Cal,” “California,” “EDCA,” and “ED Cal.”). For this situation, I did not use LLMs and only did manual review, but I wanted to make that general observation. Good for finding typos and variants. Bad for replacing pivot tables.

That being said, 6 cases in North Dakota relative to its population size and number of attorneys does seem to be somewhat disproportionate. There, I’d argue that almost every state is probably undercounting/not catching AI hallucinations, or the hallucinations have been caught, but not reported and added to the database. In fact, this exercise led me to find some information about New Hampshire as the only (apparent) state without hallucination cases.

So, Midwest’s honor defended. Let’s move on to some other lessons.

It’s not just “garbage in, garbage out”

Sometimes you’ll hear about “garbage in, garbage out” (GIGO). As it is applied to practical deployment of generative AI in businesses and other organizations, there’s a common refrain that “the AI is only as good as the data you give it.”

I’d somewhat reframe that to say “the data quality limits how well your AI can perform.” Even with good data—company policies, solid accounting spreadsheets, detailed written reports—the LLMs can still hallucinate. LLMs can make ungrounded statements even when they have access to the correct answers, they can misunderstand details, and they can improperly combine text, among other failure modes.

Charlotin’s dataset is carefully maintained as far as I have seen, and I’ve looked through a lot of it. His database appeared to be the Laine study’s sole source of data. The Laine summary stats quoted in the aforementioned report, on the other hand, suggest AI hallucinations. These uses of AI: “AI can write this research paper,” “AI can write that quarterly report for the board,” “AI dashboards give the CEO unmediated access to perfect visibility into the organization” are actually organizationally destructive uses of LLMs. They are truthiness machines, but once you look under the hood at the data, you can tell that the summaries bear no relationship to the underlying data. Since there is no relationship, correcting the data is pointless. That’s the difference between a human-curated database—like Charlotin’s—and a slopped together dashboard or report or slide deck. Of course humans make mistakes too, and human-curated databases will have rough edges or errors, but it is worthwhile to look at them closely and help fix them because somebody who cares about the truth is on the other end.

Abbreviations Are Not Always OK

North Dakota and South Dakota are not the only examples. There is a whole class of collisions between legal shorthand and two-letter state postal codes. Here are the ones I’m aware of:

Literal StringWhat it really meansCollidesImpact
‘N.D.’Northern DistrictNorth DakotaLow population, many federal northern districts, dramatic false spike
‘S.D.’Southern DistrictSouth DakotaLow population, many federal southern districts, dramatic false spike
‘M.D.’Middle District (e.g., M.D. Fla.)MarylandMore populous, fewer middle districts, mixed effect
‘Ct.’Court (Sup. Ct., Ct. App., Dist. Ct.)Connecticutpossible, but less frequently seen “in the wild”
‘CA’Court of AppealsCaliforniaLarge population, hides in the noise
‘S.C.’Supreme Court or Superior CourtSouth CarolinaMid-spectrum

Supreme Court or Supremest Court?

‘S.C.’ for Supreme Court raises a deeper problem. You still don’t know which court is highest in the given state without domain familiarity, because state court naming is not uniform across jurisdictions and “Supreme Court” is not the highest court everywhere:

  • In New York, the “Supreme Court” is the trial court and Court of Appeals is the highest.
  • Texas and Oklahoma each have two highest courts: Supreme Court (civil) and Court of Criminal Appeals (criminal).
  • Maryland changed the naming to be more familiar in December 2022.

If there are other examples, please let me know. But my main point is that you can’t just assume based on the name.

Divided Islands

When parsing strings for the name of a state, “Hawaii” or “Hawai’i” can be written with or without the apostrophe, which can also be written as ʻokina (Unicode U+02BB), a real letter in the Hawaiian alphabet. This is actually an example of an area where appropriate use of LLMs can help, because it can be exasperating for regular expressions, especially with an unpaired apostrophe.

Look at the Data and Learn About the Domain

Anyone who has spent time with both LLMs and legal citations should recognize that “N.D.” could have multiple, conflicting meanings. Also, you should assume every map is secretly a map of population until proven otherwise. So when “North Dakota” has far more AI misuse than California, look at the rows in your spreadsheet first.

My own background is in financial-crimes investigation, OSINT, and GIS. I have hit similar failures with AI, law, and geography before, as I wrote about in December 2025 (NOTE: correction, I did this in December, but I wrote about it in January). I rebuilt my hallucination-case map with Claude Code, and at first it produced a patchwork quilt because it matched county names across states without using the state as part of the key, and there are a lot of Polk Counties. The output was obviously wrong to me because I looked at the output, and I also knew why it failed, because I’m familiar with the domain. Claude Code also tried to fabricate an Eastern and Western District of South Carolina, which is a more specific hallucination that would require domain knowledge (which I had developed), to recognize that the output was incorrect.

Nice Try, New Hampshire

After manually reviewing and cleaning/hand-coding the state data, the only state that (apparently) had no attorneys with hallucinated cases was New Hampshire.

However, I found that a New Hampshire attorney included a fake quotation from a real case in a Vermont Supreme Court divorce case provided by his client. This touches on multiple topics I want to write about, including: a) that LLMs can hallucinate quotations, so checking that a case merely exists is not sufficient; b) that client-provided AI-generated fake citations are a problem for attorneys who “don’t use AI,” so they need a reality check about Shadow AI use; and c) family law is coming up frequently for AI hallucinations (e.g., early Iowa misuse and a recent Nebraska Supreme Court case).

It wasn’t adding up. A New Hampshire attorney representing a divorce client filed a legal brief that quoted from a previous court case. But when Vermont Supreme Court justices went to the case he referenced, the quote was nowhere to be found. In a November hearing, they asked the attorney where the quote came from. “Your honor, my client used an AI, um, helper” said the attorney. Justice William Cohen, now retired, responded: “The secondary source was AI? And you didn’t identify it?” “I’m not familiar with what’s involved with it and so forth,” the attorney said. He claimed that his client offered to write the brief using artificial intelligence. The quote came from “AI GPT or something like that,” he said later on. “I didn’t use it exactly but it’s a common one, I believe.” After the hearing, justices on the state’s highest court chastised the attorney for his mistakes in a court filing, requiring him to file a copy of the write-up in all of his pending cases in Vermont Superior Court. Valley News

Later, I found a more direct example of a New Hampshire case. The “Windham case” with Judge Lisa English, reported in October 2025 involving a NH/MA attorney. I unfortunately could not track down the name of the case, despite having the attorney’s name and the judge.

The Laine report produced unsupported conclusions about supposedly excessive AI misuses in North Dakota and South Dakota. However, my analysis of the data allowed me to find the one missing gap and show that there is no state without any AI hallucination cases.

Charlotin’s Map on July 2, 2026

As of today, there is no state with 0 hallucination cases on the map of USA cases on Charlotin’s website. US AI-hallucination case map by state from Charlotin’s database.

Indian Country Jurisdiction: Plaintiff also claims that Judge White wrote the McGirt opinion (he did not)

Because I was manually reviewing the cases for state coding, I thought that a case with “State of Illinois” as Defendant, but in the Eastern District of Oklahoma, may have been mistakenly coded. It was not. Instead, it was interesting case and a great example of why you should look at the data and not just let AI summarize everything in the aggregate for you.

Lowery Wilkinson Lowery, LLC, et al. v. State of Illinois, et al. (E.D. Oklahoma 2025).

Multiple times, the plaintiffs attempted to raise several jurisdictional arguments under McGirt that are so inadequate and clearly unresearched that any reasonable attorney would know they lack merit. Plaintiff repeatedly asserts that if this court fails to direct bar disciplinary proceedings in Illinois, that it is effectively “overturning” McGirt v. Oklahoma, 140 S.Ct. 2452 (2020), a case where the United States Supreme Court held that the entirety of the Eastern District of Oklahoma is “Indian Country” for the purposes of the Major Crimes Act. Dkt. No. 74 at ¶44. Plaintiff also claims that Judge White wrote the McGirt opinion (he did not), indicating that counsel has not made any attempt to research the case. Id. In addition, plaintiffs further allege that “the Defendants accused Plaintiff Lowery of a violation of the Major Crimes Act while on an Indian reservation land under McGirt … Therefore, if the Court wants to exercise jurisdiction under McGirt it would be Mandatory jurisdiction.” Lowery III, Dkt. No. 22 at p. 5. OMNIBUS ORDER

Takeaways

  • Even if an LLM feels like it gives you an overview of the data, look at the data. There’s fun stuff in there and you’ll catch more hallucinations.
  • But you should also ask if you should even be using LLMs for a given data analysis task. Pivot tables and SUM functions still exist.
  • Your starting assumption should be that every map is a population map.
  • Damien’s database is legit.
  • Come at the Midwest, better not miss.

Stop Saying 'AI Is Like a Junior Associate': Looking at the Sullivan & Cromwell Hallucinations.

· 19 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC

In the wake of the Sullivan & Cromwell AI hallucinations in a New York Chapter 15 Bankruptcy case, I am once again seeing a common trope that AI is “like a junior associate.” No, it isn’t!

“Supervising AI is no different from supervising junior lawyers or allied professionals.” Clio, ABA So says an article by Clio on the ABA website in Law Technology Today from January 2026. This is not correct! Supervising AI is very different! They are hardly the only source for this claim. Go look at the X/Twitter or LinkedIn discourse around Sullivan & Cromwell and you’ll see it everywhere.

Why do I care? Because metaphors shape how we operationalize our knowledge. This is a bad metaphor that fundamentally misunderstands how LLMs go wrong. In this post, I will give you some better metaphors to remind you of the real risks of using LLMs in legal work and explain why using the “Junior Associate Metaphor” can make those risks harder to spot.

note

Terminology: large language model (LLM) not the legal degree; “AI” discussed here as shorthand is more accurately “generative artificial intelligence,” one type of the broader category of AI; “frontier models” are the most advanced generative AI models.

If you rely on the Junior Associate Metaphor, you are likely to repeat the same mistakes as Sullivan & Cromwell at some point if you use AI regularly. Just look at the mistakes in their Schedule A: changed verbatim quotations, including major addition of wording, minor deletion or addition of words, unnecessary use of ellipses and brackets, and changes in capitalization; apparently hallucinated citations (e.g., “In re Three Arrows Cap., Ltd., 2022 WL 17985951”: more on this in “Mutant Citations” section below); and arbitrary changes to reporter numbers or years in otherwise correct case citations.

This was a process failure across several documents, not just a simple mistake limited to one filing:

  • Errors in Motion for Provisional Relief
  • Error in Verified Petition
  • Error in Motion for Joint Administration
  • Errors in Motion for Entry of an Order Scheduling the Recognition Hearing
  • Errors in Declaration of Andrew Chissick
  • Errors in Declaration of Paul Pretlove

I explicitly warned about this and even have an educational game that I created to educate attorneys about these exact failure modes.

Silently Making Material Changes In the Last Edit

If you hand a document to a junior associate for a final read-through to check wording, flow, formatting, typos, grammar, that is what they will do. The cites have already been checked. The quotations are done. The content is there. It is just cleanup. You cannot trust an LLM to follow those instructions, and this is one of the biggest problems with the “junior associate” analogy. I made a game to illustrate this problem; the AI writing problem was relevant in a case in a state Supreme Court case; and a recent research paper, “LLMs Corrupt Your Documents When You Delegate,” confirms what I have been teaching in my CLE courses.

info

For more, see this interactive game I made to demonstrate this exact failure mode: The AI "Writing Help" Trap. I also discuss this in my CLE On Demand catalog.

Better Metaphors: Japanese point-and-call and the Glass Donut Machine

Here is my metaphor: AI is like a Japanese train and you are the operator. You point-and-call everything. Instead of pointing at your instrument readings, you are pointing out and saying the material details like quotations, parties, and citations out loud and navigate to each source one last time. This is the very last thing a human does before the document is submitted (I really mean last as in the final person touching the document, no Copilot or Claude “just clean this up” afterwards LAST MEANS LAST).

Now, the Clio ABA article I mentioned earlier also talks about pilot checklists, but it barely develops this as a metaphor. Even though the article is ostensibly about the checklist, it leans heavily on the inapt Junior Associate Metaphor. And the Clio checklist itself points out things that human junior associates do not do, like blending jurisdictions (more on that in “mutant citations” below).

Do you find this level of deliberate, checklist-based review too tedious and time-consuming? Then here is my second metaphor. I have a machine that makes delicious donuts instantly and for free. Only problem is, I think one in every few batches has broken glass in it. That’s OK, “you just have to check…by tearing apart every single donut to look for the tiniest bits of glass?” But that is very time-consuming and ruins the donuts. “I guess we’ll just keep eating them, since no one has died yet.” Or maybe (crazy thought) you stop using that machine and go back to making donuts the old-fashioned way.

Now I’m absolutely not saying AI has to be perfect, nor am I anti-AI. But AI workflows can be considered to save time if and only if they save time after you account for the time it takes to properly check outputs, especially for catastrophic failures. In the case of Sullivan & Cromwell, the case In re Prince Global Holdings Limited is a multi-billion dollar pig butchering scam bankruptcy. It is possible that the case could be delayed due to the hallucinated citations (DOJ indictment of Prince Group chairman; S&C filing, April 18, 2026).

You Cannot Just Check Formatting As a “Dipstick Method”

If the associate is mistaken and doesn’t know what to do because they’re out over their skis, they are probably also going to make formatting errors, or factual errors, or come and ask questions. Because they are not going to know what they should do. Their draft is not going to look right. It is not going to sound right. Because the associate does not have the knowledge or the skill. The associate is going to hit certain limits.

To check your car's oil, you use a dipstick. It doesn’t directly show you the level of the oil, but rather how much oil is on the dipstick. Usually, this is a good proxy for what is in the car. Before LLMs, a senior attorney might have gotten away with skimming a human-prepared document for formatting. The outward signifiers of correctness (italics, special characters, spelling Latin words, using legal terms of art) would all correlate closely with the associate’s actual background knowledge and the diligence that went into preparing the document.

That formatting check might have worked (most of the time) prior to ChatGPT for human-prepared documents. That does not mean that was the right way to do it. It is not a defense of that method. I am merely describing what may have happened. But you cannot use this Dipstick Method with LLMs. It is as if the LLM regularly dunked your dipstick in a tub of oil before putting it back in your car. Now, the dipstick is back in your car, but it is not a reliable indicator of your oil levels. And having someone remind you “you have to check it” is not helpful advice.

Unlike the associate, generative AI has a jagged frontier of knowledge, where LLMs excel in some difficult areas, and fail foolishly in other simple areas. Therefore, if what the senior attorney is actually checking for is not the correct thing itself (e.g., quotation, citation), but the proxies for the correct things, like formatting—Is this italicized? Is the “section” symbol used? Did they have reporter numbers?—we see senior attorneys in these AI hallucination cases time and again signing off on stuff that looks good. Sometimes not just passable but looking really good. Again: “looks good,” but it is actually not correct.

Sometimes people will defend LLMs by saying lawyers have been cutting corners on review since before LLMs with these cursory formatting reviews. Or lawyers have been citing cases without having read them closely since before LLMs. This is the wrong conclusion. These tools are becoming ubiquitous and are causing what Damien Charlotin aptly called “Polluting the epistemic commons.” The rate at which these errors spread is growing and they will compound as the LLMs cite each other's errors, embellish errors, and create their own novel errors. This is a positive (but bad) feedback loop of polluted legal information as these citations get baked into ostensible “good” case law.

Before LLMs, the “Dipstick Method” was probably good enough—or at least good enough that the people doing it did not get caught. But, LLMs are laying bare some old bad practices that were not OK to begin with and are much more harmful now. Now that we have LLMs, it is more important to hold the line on the older rules that should have been followed all along. To use a different car analogy, drunkenness was a problem before widespread private automobile ownership, but drunk driving significantly compounded the damage a drunk could do to others.

S&C apparently had the right policy, but...

S&C stated that its “training repeatedly emphasizes the risk of AI ‘hallucinations,’ including the fabrication of case citations, misinterpretation of authorities, and inaccurate quotations. It instructs lawyers to 'trust nothing and verify everything’ [emphasis added] and makes clear that failure to independently verify AI-generated output constitutes a violation of Firm policy. The training also reviews the significant consequences of AI-related errors in various cases.”

Further, “the Firm’s Office Manual for Lawyers...provides that lawyers 'must independently check all answers, case citations, and other information or work product received from an AI Program for both substantive and non-substantive accuracy.’ [emphasis added] The policy further states that no communication may be sent to a court, regulator, client, or other external party without the exercise of appropriate professional judgment and oversight. Notwithstanding these safeguards, the Firm’s protocols were not followed here.

Trusting nothing and verifying both substantive and non-substantive accuracy seem to fit my advice. They seem to be saying not to let proper formatting and fancy wording trick you. Yet that was not enough. Why? I think the reasons are likely because a) attorneys are still not exposed to the specific ways that LLMs can hallucinate, because without exposure to convincing hallucinations that are hard to spot, naive users can be overconfident in their ability to spot hallucinations; and b) we default to metaphors, and the dominant metaphor for supervisors is the “intern” or “junior associate,” which is wrong in specific ways that lead attorneys to miss LLM hallucinations.

Mutant Citations

Schedule A from the Sullivan & Cromwell filing enumerating AI citation errors across multiple documents in In re Prince Global Holdings Limited

LLMs will hallucinate based on nothing (ungrounded hallucinations). A human might do this occasionally. They might bluff. But they won't be able to sustain it with endless detail. An LLM can indefinitely and arbitrarily extend a fake scenario: because that's what they do. So if you keep pushing the point, the associate will give up. The LLM won't.

Most of the attention gets focused on purely hallucinated cases. Cases that simply do not exist. But a bigger risk is what I refer to based on the words of a judge in the Eastern District of Michigan as “mutant citations”. These are amalgamations of real cases stuck together by the LLM. Maybe the name of a real party from one real case and a second party from a different real case. Maybe it is a federal case mixed up with a state case in the same state or region. Maybe it is the name of a real case but the wrong jurisdiction and the wrong year.

When a human cites a case and the senior attorney sees it has Bluebook formatting, a Westlaw reporter number, proper italics for Latin, uses legal terms of art, has quotations, etc., the assumption is that the document must be correct. “The associate must have actually looked at the cases. Nobody would just make that up.” A human wouldn't.

Remember that hallucinated case I mentioned from Sullivan & Cromwell: In re Three Arrows Cap., Ltd., 2022 WL 17985951”? Well Three Arrows Capital was a real cryptocurrency company that collapsed, along with several others in 2022. It is a realistic party with a realistic year that would be relevant to a bankruptcy involving pig butchering scams, which typically transact in cryptocurrency.

And, of course, it made up Westlaw numbers. For comparison, take a look at this case from December 27, 2022 on the DOJ website: Woodward v. USMS, No. 18-1249, 2022 WL 17961289 (D.D.C. Dec. 27, 2022) (Contreras, J.).” See the DOJ OIP summary. As far as I can tell, there is no real case with “2022 WL 17985951”; the results on Google show Damien Charlotin’s AI Hallucination Database and a link for the Sullivan & Cromwell filing in PACER. But it is pretty close to the numbers at the very end of 2022. It would be impressive, even shocking, for a junior associate in 2026 (who, for all we know, might’ve been in undergrad in 2022) to shoot from the hip with a WL number in the right ballpark for an end-of-2022 case, but not an actual case.

In fact, it gets more specific. The real Three Arrows Capital case was a Chapter 15 Bankruptcy case under the same judge (Chief Bankruptcy Judge Martin Glenn) as the Prince Group case, both in the Southern District of New York.

Remember that Clio ABA article I mentioned earlier? The checklist has reminders that hint at the possibility of mutant citations, even as the article keeps pushing the unhelpful analogy of supervising a human employee.

6. Confirm the correct jurisdiction General AI frequently blends jurisdictions. Ensure the content reflects the correct federal, state, provincial, or local law, including terminology and standards. 7. Look for bias or mischaracterization Check that cases are accurately described and not selectively framed. Generative AI can reflect bias from training data if left unchecked.

But does such a checklist give you the remotest sense of how almost-correct the AI can be when it hallucinates? You could “confirm the jurisdiction” and still not catch that the particular case citation as written was a hallucination.

If humans erred the way AI does, we’d call it “fraud”

A human associate typically does not say, “Oh, I found an on point case, but it's in the Fifth Circuit. We're in the Seventh Circuit. So what I'm gonna do is I'm just gonna change the citation, to this real case, and say that it was in the Seventh Circuit. And wherever it says, ‘Texas,' I'm gonna say, 'Illinois.' And wherever it says, ‘Louisiana,' I'm gonna say ‘Wisconsin.' And I'll make up some names and change the details to be consistent with that new geography.” A junior associate does not do that. An associate typically does not have the capacity to fabricate that many details. For LLMs, it’s what they do.

However, if the associate did do that, you would take that as alarming, fraudulent behavior. Most likely, you would not call that a simple “mistake.” It is probably something you would want to discuss internally at your firm for disciplinary measures or potentially even the Bar. That is not the kind of deceptive behavior you want from your employees.

You probably don’t need to worry about that, though. Somebody who is phoning it in and not really doing the work is probably not going to go through the trouble of writing a detailed fake scenario like this. An associate will more likely do a bad job that looks like a bad job. Or ask for help or do a good job.

Arbitrarily detailed: Surprisingly voluminous falsehoods

But for the LLM, it is trivial to make up details. That is why in the infamous Mata v. Avianca case, the attorneys did not only cite fake cases but provided the supposed text of those fake cases to the Court after the opposing counsel (Avianca) said they could not find the fake cases supposedly about airlines (their area of expertise).

If, like the Mata attorneys, you ask an LLM to provide the text of a case that does not exist, it might continue to bluff and hallucinate the entire text of the case. The more you drill down within the LLM itself, the more detail it gives. But it’s all fake! You have to remove yourself from the AI tool and actually navigate to the primary source, not a chatbot or AI summary: and I do include Google and the legal AI tools in this list. A junior associate simply won’t spit out pages and pages of six fake cases off the top of their head when pressed to defend a bad citation.

Verisimilitude: Surprisingly accurate falsehoods

Another surprising thing I have observed is when LLMs make up cases, the reporter numbers can be shockingly close to the right thing. They might be close to an unassigned number. It doesn't correspond to a real case. But it would fall into that range for the jurisdiction, for the state, for the year, if that case had existed.

A junior associate is not going to be able to do that off the top of their head. They are not going to be able to make up a number. They are not going to be able to say—and I'm going to make this up off the top of my head, so it's not going to be right—“847,” and know that that 847 is going to fall between 845 and 850, which are both real cases from the same state and year.

An associate is not going to have the encyclopedic knowledge and recall to do that. And like I said before, it is not a human mistake. If a human were capable of doing that, we would call it deception or lying, not a mistake. But this is the kind of mistake that frontier LLMs make.

The time it takes to check

It can take a lot of scrutiny to verify or negate these types of hallucinated citations. It will not necessarily be obviously wrong because it may not be absurd. The hallucination could be very close to the real thing, which is bad. Obviously wrong things in some sense are actually better.

This is why I frequently say that I don't like it when people say “the AI models are getting better.” Rather, I reframe it as “the models are getting more capable.” If a model is old, an incorrect case citation might have said something like Smith v. Jones (2009). That looks hallucinated. It could be a real case, but it is vague: two extremely common surnames and no jurisdiction.

If instead, it had the name of a real individual (first and last name) versus a real car dealership in Indiana, and it said that the case was in the Court of Appeals of Indiana, 2014, with a reporter number that fit with that year, it could take a lot of work to determine that citation was actually an AI hallucination.

When checking, watch out for Doppelgänger Hallucinations

Lastly, I need to talk about my Doppelgänger Hallucination tests, which I first warned about in October 2025. Basically the LLM is sycophantic and wants to tell you what you want to hear. If you ask it a question, even the question as simple as just giving it a bare case citation with no other wording, only a properly formatted fake case, may result in the LLM providing a description of that case.

If you use other AI tools to try to cross-reference your original source, you may become more convinced that the fake case is real if you are following the Junior Associate Metaphor. Maybe you google it (AI Overview), or check Gemini, ChatGPT, Claude, Perplexity, etc. Maybe some of those AI tools are telling you it is a real case. Maybe not all of them, but maybe multiple of them are saying it is real. The details they tell you about that case are probably not consistent with each other (since they are hallucinating independently). But you might go even further down the rabbit hole of “what's going on with this case?” And in the end, it's not real.

A lot of times, the LLM likes to tell you it was an “important” case or a “significant” case. It will make up descriptions of the fake case and tell you it was a major case that perhaps “set a precedent” or “is frequently cited.” All of that is not true.

In my Kruse v. Karlen test, I used the 22 known fake citations from the reference table. Searching for only the citations in quotes, the Google AI Overview gave an inaccurate answer describing the fake case as real—despite having access to a source explicitly calling it fictitious—roughly a quarter of the time; this rate rose to over half of the time if the user opted for “AI Mode.”

This phenomenon of another LLM corroborating a fake case is not a hypothetical. This has happened to real attorneys where they've double-checked and triple-checked using multiple LLM-powered tools.

An associate does not usually do that. Even if they did, they cannot sustain that level of embellished fabricated detail off the top of their head. It's just not how people work.

AI is not like a junior associate. Treating it like that can result in a three-page single-spaced list of corrections, with possibly even more errors being spotted by opposing counsel.

Three Ways AI Can Make Things Up. How True But Irrelevant Can Be Harder to Correct Than Pure Nonsense.

· 5 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC

More Than One Type of Hallucination

ChatGPT sometimes makes things up. For example, ChatGPT famously made up fictional court cases that were cited by attorneys for the plaintiff in Mata v. Avianca. But totally made up things should be easy to spot if you search for the sources. It’s when there’s a kernel of truth that large language model (LLM) hallucinations can waste the most time for lawyers and judges or small businesses and their customers.

  1. A “Pure Hallucination” is something made up completely with no basis in fact.
  2. A “Hallucinated Summary” has a footnote or other citation referencing a real source, but the LLM’s description of what that source says has little if anything to do with the source.
  3. An “Irrelevant Reference” is when an LLM cites a real sources and summarizes it fairly correctly, but the citation itself is not relevant to the purpose of the citation. This might be because the information is outdated, because the point only tangentially refers to the same topic, or for other reasons.
info

These examples were derived by actually reading the sources and were not written by LLMs. All of the written content on our website and social media is human-written, unless it is an example of AI-output that is clearly labelled.

danger

AI can help people summarize or rephrase content they know well. But Midwest Frontier AI Consulting strongly encourages AI users not to rely on AI-generated overviews of content they are not already familiar with precisely because of the subtler forms of AI hallucinations described below.

Scenario 1: You Got Your Chocolate In My Case Law

  • Pure Hallucination: ** The LLM says: “Wonka v. Slugworth clearly states that chocolate recipes are not intellectual property.” ** In reality: No such case exists.

  • Hallucinated Summary: ** The LLM says: “NESTLE USA v. DOE clearly states that chocolate recipes are not intellectual property.” ** In reality: The case involves a chocolate company but is not about intellectual property rights.

  • Irrelevant Reference:

    • The LLM Says: ‘HERSHEY CREAMERY v. HERSHEY CHOCOLATE involved two parties that both owned trademarks to “HERSHEY’S” for ice cream and chocolate, respectively. This supports our assertion that chocolate recipes are not intellectual property.’
    • In reality: The facts of the case do not support the conclusion.

Mata v. Avianca Was Not Mainly About ChatGPT

· 11 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC

Mata v. Avianca: The First ChatGPT Misuse Case

The case Mata v. Avianca was a personal injury lawsuit against an airline in the U.S. District Court for the Southern District of New York (SDNY). However, the reason it became a landmark legal case was not the lawsuit itself, but the sanctions issued against the plaintiff’s lawyers for citing fake legal cases made up by ChatGPT. At least that was the popular version of the story emphasized by some reports. The reality, according to the judge’s opinion related to the sanctions, is that the penalty was about the attorneys doubling down on their misuse of AI in an attempt to conceal it. They had several opportunities to admit their fault and come clean (page 2, Mata v. Avianca, Inc., No. 1:2022cv01461 - Document 54 (S.D.N.Y. 2023)).

Take this New York Times headline “A Man Sued Avianca Airline. His Lawyer Used ChatGPT,” May 27, 2023. This article, written before the sanctions hearing in June 2023, focused on the ChatGPT-gone-wrong angle. By contrast, Sarah Isgur of the Advisory Opinions podcast had a very good breakdown noting the attorney’s responsibility and the back-and-forth that preceded the sanctions (episode “Excessive Fines and Strange Bedfellows,” May 31, 2023). However, in that podcast episode the hosts questioned the utility of ChatGPT for legal research and said “that is what Lexis and Westlaw are for” but as of 2025 both tools have added AI features including use of OpenAI’s GPT large language models (LLMs).[^1]

caution

I am not an attorney and the opinions expressed in this article should not be construed as legal advice.

A surrealist pattern of repeated dreamers hallucinating about the law and airplanes.

Hallucinating cases about airlines.

Why Care? Our Firm Doesn’t Use AI

Before I get into the details of the case, I want to point out that only one attorney directly used AI. It was his first time using ChatGPT. But another attorney and the law firm also got in trouble. It only takes one person using AI without proper training and without an AI policy to harm the firm. It seems that one of the drivers for AI use was access to other federal research tools was too expensive or unavailable, a problem that may be more common for solo firms and smaller firms.

Partner of Levidow, Levidow & Oberman: “We regret what's occurred. We practice primarily in state court, and Fast Case has been enough. There was a billing error and we did not have Federal access.” Matthew Russell Lee’s Newsletter Substack

You might say, “Fine! We just won’t use AI then.” Do you have a written policy stating that? Do you really not use AI? I have two simple questions:

  1. Do you have Microsoft Office? (then you probably have Office 365 Copilot)
  2. Do you search for things on Google? (then you probably see the AI Overview) If the answer to either is yes (extremely likely), are you taking measures to avoid using these AI features? If not, how can you say you don’t use AI? Simply put, avoiding AI is not the default option. It requires conscious effort to avoid the features being added to existing software, from word processors to specialty legal research tools.

Overview of Fake Citations

The lawyers submitted hallucinated cases including the court and judges who supposedly issued them, hallucinated docket numbers and made up dates.