A lawyer turned to ChatGPT for help with a murder appeal, and the AI got him into trouble.
New Mexico defense attorney Stephen Aarons used ChatGPT while preparing an appeal for a client convicted of murder and sentenced to life in prison. But when the material entered the court record, several claims were clearly inaccurate. The New Mexico Supreme Court said the brief contained fabricated witnesses, false testimony, and other inaccurate claims.
The bigger problem was not simply that ChatGPT got the facts wrong: It generated details that were not in the case record, and Aarons submitted them without checking. The court held him in contempt and imposed a $5,000 sanction.
AI hallucinations go legal
According to a Reuters report, Aarons, a private attorney based in Santa Fe, was representing Oscar Rene Sandoval in an appeal of his murder conviction and life sentence in the killing of the mother of his children when he filed an AI-assisted brief in court without verifying its contents.
The New Mexico Supreme Court later found that the filing contained a string of false claims. According to the court order published by Justia, the brief named people who had not testified, including Officer Michelle Amarillo, Officer Sanchez, Manal Al-Jibury, and Teresa Marquez. It also attributed statements to Danny Stanton and Linda Stanton that did not appear in the record.
The filing contained false testimony attributed to Mariah Chavez and Marquez concerning the shooter’s clothing and appearance. It also incorrectly stated that the shooter wore dark pants and a white shirt, Reuters reported.
Aarons admitted during an Aug. 21 hearing that he had fed a computer-generated transcript of the underlying proceedings into ChatGPT while preparing the brief. He said he believed the chatbot would produce a “bulletproof” summary. However, he signed and filed the brief without verifying its factual claims or legal authorities.
The court found Aarons in contempt, referred him to the state Disciplinary Board, barred him from appearing before the New Mexico Supreme Court while the disciplinary case is pending, and ordered him to pay a $5,000 sanction.
The court also struck the briefs and said Aarons had shown no remorse for the harm caused to his client. Aarons disputed that characterization, telling Reuters that he was remorseful and that it was “an honest mistake.” A public defender has been appointed by the court to represent Sandoval going forward.
More must-read AI coverage
- SS&C Intralinks DealCentre AI vs. Datasite: Which platform is built for the future of dealmaking?
- SS&C Intralinks FundCentre AI vs. Juniper Square: Which platform better supports modern private markets fund managers?
- Why Data, Not Models, Determines AI Success
- The Rise of the AI-Native Factory: How Physical AI Is Transforming Manufacturing
AI hallucinations are bigger than this case
The problem behind Aarons’ filing is not unique to ChatGPT or legal work. Generative AI systems can produce convincing statements that are unsupported or entirely fabricated.
Large language models generate probable responses rather than guaranteed facts. Their training and evaluation can also reward guessing instead of acknowledging uncertainty. OpenAI confirmed this in a September 2025 report, noting that “language models hallucinate because standard training and evaluation procedures reward guessing over acknowledging uncertainty.” A response can therefore sound confident and authoritative while still containing a person, quotation, event, or citation that never existed.
That distinction becomes more important as people use AI for work that depends on factual accuracy. A wrong answer in a casual chatbot conversation may be easy to ignore. However, the same error inside a legal filing, medical document, financial analysis, or security report can carry consequences well beyond the original AI response.
Aarons’ case shows what that risk looks like when generated text is allowed to pass through without a human fact-check.
What this means for anyone using AI at work
Aarons’ case reminds us that using AI at work does not transfer responsibility to the chatbot.
That does not mean professionals should stop using tools like ChatGPT where their organizations permit them. It means AI output should be treated as an unverified draft, especially when the work involves legal, medical, financial, security, or personnel decisions.
Before relying on an AI-generated claim, users should trace it to the original source, confirm quotations and names, and check that every cited authority exists and supports the argument being made. Aarons’ case shows that when this review is skipped, responsibility remains with the person who submits the work — not the chatbot that generated it.