On Sept. 11, Ars Technica's Jon Brodkin reported that the New Mexico Supreme Court had held a lawyer in direct contempt over a brief containing, in the court's words, "false testimony from wholly fabricated witnesses." The lawyer admitted he had not verified the factual claims in the AI-generated brief before signing and filing it, and told the court how it was made: he fed a computer-generated transcript of the trial into ChatGPT, and it produced quotations that were not in it.

The useful part is the opposite of the headline. The record was there — he handed over the actual transcript — and the answer still contained names the transcript did not.

Chef-owner Margaret Ellison peers at an unchanged cream record card tucked beneath a cardboard box on a storeroom shelf.
Nine o'five: the card under the box nobody had moved

That gap, between having your facts and answering from them, is what three more items describe in the same six days, from three unrelated rooms.

What the researchers measured

On Sept. 16, in Search Engine Journal, the consultant Pedro Dias walked through a study called MemToC, which tests what a model does when its own answer conflicts with what a tool hands back. Researchers took cases where a model had already answered a factual question correctly, then asked again while a tool supplied an incorrect answer. Across four models, the correct answer survived between 6.5% and 17.1% of the time — and in a separate annotation sample of 120 responses to incorrect tool returns, not one explicitly acknowledged that it disagreed with what it had been given.

Dias states the limits in the same breath, and so do we: these were controlled tests, principally on open-weight models of 7 to 9 billion parameters, and, in his words, "the percentages cannot be applied to ChatGPT search or Google AI Overviews." The number is not the point. The shape is: when the data handed to a model is wrong, it will usually go along with it, and it will not mention that it noticed.

Margaret holds a kraft delivery record and a cream card side by side against a tall window to compare them in the light.
Two of her own papers, held up to the light together

What the money did

Two rounds landed the same week, priced against that shape.

On Sept. 16, SiliconANGLE's Paul Gillin reported that TypeSafe AI came out of stealth with $40 million and a model that returns structured decisions instead of conversational text. The company's thesis is worth reading even if you never touch the product: what makes models good at talking to people works against them inside software, because they can "generate plausible but incorrect information, vary their methods between requests and present uncertain answers confidently." Its chief executive, Diogo Almeida, a former OpenAI researcher, put it more bluntly: "We've been optimizing for humans, and we're superhuman at pleasing humans." So the product hands back a labelled answer with a confidence measure, and the software around it decides: proceed above a threshold, ask for more, or route the case to a person.

The same day, covering the same launch, the-decoder's Maximilian Schreiner added the sentence a buyer actually needs. The company markets the model as one that cannot hallucinate; that guarantee, Schreiner notes, covers only the shape of the output — it will not answer outside the allowed options — and "a factually wrong choice within those options is still possible."

Also on Sept. 16, SiliconANGLE's Duncan Riley reported that CADDi raised $114 million at a $1.2 billion valuation to expand in North America. Its case is that general-purpose AI cannot handle an industry's data without domain-specific models; co-founder Yushiro Kato told Fortune, "I've never seen anybody who uses LLMs to do design reviews because it doesn't understand drawings or CAD." One of its investors, Ken Hara of Coreline Ventures, drew the line for everyone: "AI only delivers where a data layer and a semantic layer already exist."

A day earlier, Riley reported a third round: AIUC raised $40 million to extend its certification work to frontier models. Its chief executive, Rune Kvist, said most enterprises have agents that were approved in pilots but "stalled at the security review" — proof of reliability, not capability, is now the bottleneck. Certification runs an agent through roughly 5,000 combinations of risk and attack, hallucinations among them, and one voice AI company used its certificate to obtain insurance that covers an agent giving a customer incorrect information. Co-founder Rajiv Dattani reached for a precedent: "When electricity was burning down houses, the insurers paying the bill funded Underwriters Laboratories to test and certify products."

Say the honest thing out loud: all three sell a cure for the illness they describe, and none of their performance claims are independently verified, as both outlets covering TypeSafe point out. What is a fact rather than a pitch is the money — $154 million in one week, aimed at the same seam.

Ruby, the AI Customer Support Agent, stops in a service-corridor doorway with both hands empty and one palm open.
She stops in the doorway with both hands empty

The same question, asked twice

The fourth reading is the closest to a small business. On Sept. 14, in a contributor column for Entrepreneur — opinions there are the contributors' own, and this is opinion — Meghna Deshraj argued that AI visibility scores are modeled samples rather than ground truth, and brought numbers: in a 2026 crowdsourced study, 600 volunteers ran the same brand-recommendation prompts nearly 3,000 times, and the same list of brands came back in fewer than one run in a hundred. Separate research covering 693,509 repeat answers found that two responses to the same ChatGPT prompt shared only 21.2% of their cited domains.

Her practical advice is the part to keep: the most valuable question set in your category already exists inside your business — in your sales calls and your support tickets — because those are the questions in the words buyers really use.

What this looks like in a business with one owner

Put the four together and those six days say one thing to somebody running thirty seats or one van.

Nobody who writes to you reads your knowledge base. They read one sentence. That sentence is worth exactly two things: the record it was assembled from, and what the thing answering does when the record is silent. The court case is the first failure — the record was fine, the check was skipped. MemToC is the second — the record was contradicted, and the disagreement was never mentioned. The three rounds are the market deciding the second failure is the expensive one, because it is invisible.

A steel rail of cream record cards has one empty gap while the card nearest the open service door lifts in the draught.
One gap in the rail, and one card lifted by the draught

You cannot audit a model. You can audit two much smaller things, and both are in your hands.

Three checks an owner can run in one afternoon

1. Ask the thing that answers for you the five questions you get most — and check the numbers. Not "does it sound right": take each figure in the reply and find the line in your own catalog or policy it should have come from. If you cannot find the line, you have just watched the court filing's failure happen at your own counter.

2. Find the one place your own facts disagree. Your public profile, your catalog, your knowledge base and the promotion you ran last month are four records, and they drift. Read them as a stranger and fix the first contradiction. Whatever answers for you will be confident about whichever one it was handed.

3. Write down the sentence it must say when it doesn't know. Word for word, before you hire anything: what it says, who it hands the question to, and what the customer is told while they wait. In those six days the industry paid for exactly that boundary — a confidence threshold, a domain model, a certificate. Yours costs one sentence. Ours is described below, so you can hold us to it.

How we think about it

We build AI employees, so we are not neutral. Our answer to those six days: the useful question is not whether an AI can read your data, but what it does at the edge of it.

Ruby, our AI customer support agent, answers the people who already bought from you. She starts from your own material, and she repeats the problem back in the customer's own words before she starts helping. Her reply opens with one line saying it comes from your business's AI assistant.

Then the edge. When the answer is not in what she was given (an order she cannot see, for example), she says so plainly and tells the customer how to reach you, instead of guessing. She does not promise a refund, and she does not promise that a problem will never happen again: that is your money and your word, not hers.

A cream record card lies face down on a steel shelf while Margaret and Ruby stand behind it in the service corridor.
Face down on the shelf until the two records agree

That is the whole claim, and it is smaller than a sales page on purpose. You do not have to take our word for it: she answers in public at Tampa Pasta House and Casa Lista, before anything is bought. Setup is a conversation in plain words in the ChatGPT, Claude or Gemini app you already use.

A record that answers everything is not the goal. An answer that knows where it stops is.

Sources

  1. Ars Technica, Sept. 11, 2026, Jon Brodkin — "ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses"
  2. Search Engine Journal, Sept. 16, 2026, Pedro Dias — "It Was There A Minute Ago"
  3. SiliconANGLE, Sept. 16, 2026, Paul Gillin — "TypeSafe AI exits stealth with $40M to build AI for use by software"
  4. the-decoder, Sept. 16, 2026, Maximilian Schreiner — "Former OpenAI researcher builds an AI model that judges options instead of writing text"
  5. SiliconANGLE, Sept. 16, 2026, Duncan Riley — "CADDi raises $114M at $1.2B valuation to bring manufacturing AI to North America"
  6. SiliconANGLE, Sept. 15, 2026, Duncan Riley — "AI agent certification startup AIUC raises $40M to begin auditing frontier models"
  7. Entrepreneur, Sept. 14, 2026, Meghna Deshraj — "The Most Valuable AI Search Data in Your Business Is Already Sitting in Your Sales Calls"

Note on sources 3–4: both describe the same launch on the same day, independently; source 4 is the more critical of the two, and its qualification — that a structural guarantee is not a factual one — is the one we quote.

Note on sources 3, 5 and 6: all three are SiliconANGLE, and all three companies sell products addressed to the problem they describe. Both facts are stated in the text.

Note on source 2: the study percentages are from controlled tests on open-weight models, and the author states they cannot be applied to consumer AI search. We repeat the caveat wherever we repeat the number.

Note on source 7: Entrepreneur states that opinions expressed by its contributors are their own. We quote it as opinion, and the two studies it cites are named as its citations, not as our measurement.

Note on source 1: the court order is used only for the professional facts admitted in it. Nothing about the underlying criminal case appears in this piece.

SMASHONE CORPORATION, Florida, United States