An engineer hired to give an independent opinion typed the conclusion he wanted before he did the analysis. That is not the story. The story is that he typed it into something that kept the receipt.
Read only the first prompt in each thread. If it names the conclusion instead of the evidence, that thread is a liability no amount of editing the output can fix.









404 Media has published the ChatGPT logs of an expert witness in the litigation over the Watson Grinding explosion in Houston. The prompt everyone is quoting asked the model to "show how 3M is 0% at fault for the explosion at Watson Grinding."
The reaction has settled into a familiar shape: AI has contaminated expert testimony, courts must respond, another hallucination scandal. I understand why it is landing that way. The quotes are lurid, the deaths are real, and we have spent three years learning to blame the chatbot.
But that reading gets the news backwards. Retained experts have been reverse-engineering conclusions for their clients since the invention of the retained expert. What changed is not the bias. It is that the bias now arrives in discovery with a timestamp on it.
The report that wrote itself
The underlying facts are not in dispute and they are ugly. At 4:24 in the morning on January 24, 2020, a building at Watson Grinding and Manufacturing in northwest Houston exploded. Two employees, Gerardo Castorena Sr. and Frank Flores, were killed. A neighbor, Gilberto Mendoza Cruz, later died as well. Roughly 200 nearby houses and businesses were damaged. The U.S. Chemical Safety Board's final report traced the propylene release to a degraded, poorly crimped welding hose, a manual shutoff valve left open, and an inoperative gas detection alarm.
That last item is why 3M was in the case. Homeowners alleged the company failed to properly service the facility's gas detection system. 3M retained Josh Autenrieth of Knighthawk Engineering to write an expert report on the standard of care, an engagement court transcripts put at roughly $90,000, billed at $475 an hour.
In discovery, the plaintiffs' attorney Will Moye found a five-page document titled "Citation Overlay" with the unmistakable formatting of a chatbot. He demanded the prompts. The deposition paused three hours while 3M's lawyers gathered them, and Moye came back with 350 pages of ChatGPT conversations. At trial, he and Autenrieth agreed on the record that the filed report was "90 to 85 percent ChatGPT."
Somewhere in those pages, the expert on gas detection uploaded a photograph of a gas detector and asked the model, "what am I looking at?" Moye's assessment of that device: it is the subject of the whole case.
The machine played opposing counsel
Here is the detail that should keep engineering managers up at night, and it is not the one being quoted.
ChatGPT produced a roughly 30-page draft containing the sentence "From a technical and standard-of-care standpoint, 3M is 0% responsible for the January 24, 2020 explosion." That sentence never reached the court. Autenrieth asked the model to review the draft as opposing counsel would. The model flagged "0% responsible" as an easy target, one of several phrases that would let the other side paint him as an advocate rather than an expert. He cut it.
Read that sequence again. The model did not challenge the analysis. It did not ask whether a hose failure at four in the morning could plausibly be nobody's fault but the operator's. It found the sentence that made the bias legible and removed it. The claim survived intact into a document filed with a court.
This is the amoral-engine problem from LLMs Have No Intent landing somewhere with real consequences. Ask a model to defend a position and it will, then coach you on sounding less like someone defending a position. It has no opinion about 3M. It has an excellent statistical sense of what advocacy sounds like, which makes it very good at removing the smell without removing the thing that smells.
Same pattern I wrote about in The Number Nobody Measured: the machine that launders a claim and the machine that checks it run the same weights. Polish is not verification. It never was.
The rule that made it findable
None of this would matter if prompts were protected. Lawyers have spent decades building a doctrine around what an expert must hand over, and drafts, notes, and communications with counsel sit largely on the protected side of that line under Rule 26(b)(4).
Prompts used as part of an expert's methodology may not.
On May 18, 2026, Magistrate Judge Thomas O. Farrish of the District of Connecticut ordered the Conservation Law Foundation to produce the AI prompts its expert, Dr. Naomi Oreskes, used in preparing her report in Conservation Law Foundation v. Shell Oil (No. 3:21-cv-00933). Read the order rather than the headlines, because it is narrower than they are. An expert witness's methodology is fair ground for discovery, it says, citing Macchia v. ADP for the point, "and under the facts of this case, the process by which Dr. Oreskes culled down the defendants' document production into a subset to be worked with is an aspect of that methodology."
Under the facts of this case. That is a ruling about one expert's document filtering, not a holding that prompts are never protected.
CLF argued the prompts were "notes" under a Rule 29 agreement; the court found that classification not "quite clear" enough to shield them. CLF also argued that Oreskes had used search terms rather than prompts. The court disbelieved that one, because her assistant had already written the word "prompt" into a declaration. The tell was in somebody else's filing.
Two cases, two different jobs
The cases are not the same animal:
- Oreskes supplied the rule. She used AI to filter documents. Ordinary, defensible methodology, and nothing in the reporting suggests otherwise.
- Autenrieth shows what the rule finds. Same discovery obligation, pointed at a thread that opens by naming the answer.
- Neither is settled law yet. Per Mayer Brown, CLF objected on June 2 and the order is stayed pending review.
Call it direction of travel rather than precedent. Every litigator I can find writing about it has reached the same conclusion.
The constraint underneath is not really about AI. Discovery has always demanded that methodology be inspectable, because an opinion you cannot audit is not evidence. Prompting is now where a meaningful share of professional analysis happens. The doctrine did not have to stretch to reach it.
We learned this once already, with email
Twenty-eight years to the day before Farrish signed that order, on May 18, 1998, the Department of Justice and twenty state attorneys general sued Microsoft. I was there then, building the newsroom CMS at MSNBC, close enough to watch a whole company rethink what it put in writing.
What made that case memorable was not the legal theory. It was the email. Executives had typed things into a system they experienced as conversation, and the government read them back as exhibits. The lesson the industry absorbed was not "behave better." It was "do not put that in email." A generation of corporate communication training followed, most of it about phrasing.
We are several years into treating prompts the way that generation treated email in 1996. A prompt feels private. It feels like thinking out loud. It is neither. What you produced is a timestamped transcript of what you asked for, usually held by a third party, on infrastructure and under retention rules your team may not control, with a federal magistrate now on record that it can be fair ground.
Your postmortem is an expert report
If you read this as a courtroom story you will file it away as somebody else's problem. It isn't.
Think about what your organization produces in writing that could later matter to somebody with a subpoena. Incident postmortems. Vendor security questionnaires. Architecture decision records. Compliance attestations. Root cause analyses sent to a customer after an outage. Due diligence memos. Each is a document where a professional asserts a conclusion, and a growing share are now drafted by asking a model for help.
Compliance artifacts are the sharpest case, because this industry already mistakes a generated record for a verified one. I made that argument about electronic signatures in The E-Signature Audit Trail Is Theater. The failure mode transfers cleanly: a document that looks procedurally impeccable tells you nothing about whether anyone did the work.
One correction to that analogy, and it cuts against you. Rule 26(b)(4) is a shield built for retained experts, covering their draft reports and most of what they say to counsel. You do not have it. An engineer's postmortem is not expert work product. It is an ordinary business record, reachable by a plain Rule 34 document request like any Slack thread or email, and the only thing standing in front of it is attorney-client privilege you probably cannot claim.
The expert in Houston at least had a doctrine to argue about. Your prompt log has none.
The question is not whether you used AI. The question is what the first message in the thread says.
Compare two openers. "Here are the logs and the two competing theories, help me structure a postmortem." Versus "explain why our service was not the cause of Tuesday's outage." Both produce a readable document. Both look identical on the page. Only one survives being read aloud by a lawyer who is not on your side.
I have written engineering assessments for clients who could have sued over them. The discipline that keeps you honest is old and boring: write down the method before you know what it will say. That was always good practice. It is now the difference between a prompt log that defends you and one that ends you.
What actually works: the discoverable prompt audit
Take any artifact your team produces that a third party might one day rely on, and check the first prompt against this table. The left column is what a conclusion-shopping prompt looks like in the wild. The right column is the same task, framed so the log is an asset.
| Artifact | Prompt that becomes an exhibit | Prompt that becomes a defense |
|---|---|---|
| Incident postmortem | "Explain why our service wasn't the cause of the outage." | "Here is the timeline and the logs. List every plausible contributing cause, including ours, and what evidence would confirm or rule out each." |
| Vendor security review | "Write a justification for approving this vendor." | "Here is their questionnaire and pen test summary. Identify unanswered questions and the controls we are taking on trust." |
| Architecture decision record | "Make the case for moving to the new platform." | "Here are the three options and our constraints. Give the strongest argument for each, then the failure mode of each." |
| Root cause analysis for a customer | "Help me phrase this so it doesn't sound like negligence." | "Here is what we found. Draft a factual account, flagging anything I have stated as certain that the evidence only makes probable." |
| Compliance attestation | "Show that we meet this control." | "Here is the control text and our implementation. Where does the implementation fall short of the control as written?" |
How to score it: read only the first prompt in each thread. If it names the conclusion, that thread is a liability no matter how good the output was. If it names the evidence and asks what follows, you have contemporaneous proof that you reasoned before you concluded. For audit purposes there is no useful third category, and editing the output never moves a thread between columns.
The exception: when the log defends you
I am not arguing that professionals should stop using these tools, or that every prompt log is a bomb. There are exceptions, most threads are fine, and the Oreskes example is the proof.
The most useful sentence written about this case did not come from the bench. Reading the order, Arnold & Porter's Melissa Weberman and L. Michel Marchand drew the practical conclusion: prompts reflecting a disciplined analytical approach may strengthen the credibility of an expert's work, while poorly defined prompts invite challenges to the methodology itself.
That is a prompt log working for you. An expert who can produce a thread showing she asked the model to surface documents contradicting her thesis, then addressed them, has documented rigor in a way no traditional workflow could. Handwritten notes get lost. Reasoning that happened in someone's head cannot be produced at all. A well-formed prompt log is the most complete record of professional judgment that has ever existed.
The tool is symmetrical. It records the discipline and the shortcut with equal fidelity. What is not symmetrical is that only one of those was ever provable before, and it was never the shortcut.
The Bottom Line
Stop treating prompts as scratch. Assume every thread that touches a professional deliverable will be read by someone hostile, and write the first message accordingly. That single habit costs nothing and converts your biggest new liability into your best new documentation.
Then check whether your litigation hold and retention policies mention chat tools at all. In most organizations I have looked at, they do not. The record exists, nobody owns it, and the first person to ask for it will be opposing counsel.
Find out where it lives while you are in there. A consumer chat workspace keeps history by default; an enterprise endpoint configured for zero retention may keep nothing. That difference decides whether you are producing your own logs or subpoenaing somebody else's, and it is a decision most teams have made by accident.
The expert in Houston did not get caught because he used a chatbot. He got caught because he asked it for an answer instead of an analysis, and because a lawyer recognized the formatting. The bias was always there. What's new is that it has a timestamp.
"The bias was always there. What's new is that it has a timestamp."
Sources
- Order granting motion to compel production of reliance materials (ECF No. 970), Conservation Law Foundation v. Shell Oil Co., No. 3:21-cv-00933 (D. Conn. May 18, 2026) — The decision itself. Holds that an expert's methodology is fair ground for discovery, citing Macchia v. ADP, and that 'under the facts of this case' Dr. Oreskes' AI-assisted document culling was part of it. Also rejects the Rule 29 'notes' argument as not 'quite clear', and disbelieves CLF's search-terms-not-prompts position because her assistant's declaration used the word prompt.
- 'Show How 3M Is 0% at Fault:' Expert Witness Used ChatGPT to Write Report Defending Company in Deadly Explosion Lawsuit — Primary reporting on the 350 pages of ChatGPT logs produced in deposition, including the verbatim prompts and the trial admission that the filed report was '90 to 85 percent ChatGPT'.
- Fatal Propylene Release and Explosion at Watson Grinding and Manufacturing — Federal investigation establishing the January 24, 2020 cause chain: a degraded, poorly crimped welding hose, an open manual shutoff valve, and an inoperative gas detection alarm.
- AI Prompts Used by Expert Are Subject to Compelled Discovery — Analysis of Judge Farrish's May 18, 2026 order in CLF v. Shell, quoting the holding that an expert's methodology is fair ground for discovery.
- Court Orders Disclosure of Expert Witness's AI Prompts: What Litigators Need to Know — Litigator analysis of the Rule 26(b) reasoning and the procedural posture, including CLF's June 2, 2026 objection and the stay of the order.
- Complaint: United States v. Microsoft Corp. — The May 18, 1998 antitrust complaint, the case whose email exhibits taught the industry its first lesson about informal writing becoming evidence.
- Court Rules Expert's AI Prompts Are Fair Game Under Rule 26 — Arnold & Porter's Melissa Weberman and L. Michel Marchand on the practical consequence of the Farrish order: a disciplined prompt log can strengthen an expert's credibility, a sloppy one invites challenges to the methodology. Source of the quotation in the exception section.
Is Your AI Paper Trail Defensible?
I'll review how your team uses AI on incident reports, vendor reviews, and compliance artifacts, and where the record works against you.
Book a Prompt Risk ReviewDisagree? Have a War Story?
I read every reply. If you've seen this pattern play out differently, or have a counter-example that breaks my argument, I want to hear it.
Send a Reply →