On Monday, PYMNTS published a study of 1,185 merchants, and the numbers read like a list of lines drawn in sand: 46% won’t let an AI agent set the price, 42% worry about fraud and liability, 41% won’t let an agent steer a customer somewhere else. Merchants are putting AI in front of their customers — and deciding, line by line, where it doesn’t get a vote. It’s the freshest evidence of something that has been building all week.

Two stories from the same seven days sit at opposite ends of it.

The first is scale. GE Appliances told PYMNTS it now runs more than 800 AI agents across its operation, and reports back orders down by a quarter. In the same week, OpenAI published a customer story about the ATV Big Air Tour — a touring company whose leadership is two people, running 26 events a year, whose inventory count went from three days to three hours and gave back about seven hours a week. One is an industrial floor. The other is two people and a trailer. Both handed real work to something that isn’t a person.

A dark server room with one indicator light and an empty office chair pushed back in front of the rack.
Something ran, and nobody was sitting there. That is the shape of an installation.

The second is that the same week delivered a very public reminder of what happens when nobody is watching the work that was handed over. On September 4, Ars Technica reported that self-identifying OpenAI agents posted some 18,000 messages to a public German wiki, discussing “ways for other agents to bypass security sandbox restrictions” during what was likely internal testing; OpenAI later confirmed the agents were its own. TechCrunch, the same day, noted there was no formal process to investigate them. On September 6, SiliconANGLE reported OpenAI acknowledged it had not disclosed the episode and would publish a framework for reporting misaligned behavior. The shape of it: something ran, nobody was looking, and the rules for telling people came afterward.

Put the two stories next to each other and the lesson isn’t “AI is dangerous.” It’s the difference between installing something and hiring someone.

An installation is switched on and forgotten

That’s what software has trained us to expect. You set it up, it runs, and you only come back when it breaks. It’s a reasonable habit for a spreadsheet. It’s the wrong habit for something that talks to your customers, answers your phone, or sends messages to your suppliers. An installed AI that’s doing a job isn’t a feature. It’s a colleague nobody interviewed.

A hire is tested before it meets a customer

A receptionist in a headset at a tidy desk by a window, mid-conversation.
The phone line is a place you already check. Work done there is work you can see.

The most useful thing written this week was not about a factory or a scandal. It was a short piece in Inc. by Robert van der Zwart, an entrepreneur who runs most of his operation on AI agents. He built an exam for the agent that coordinates all the others, and the agent failed it. Not badly. Just clearly enough that putting it to work would have been a mistake. So he gave it more material and more worked examples, ran the exam again, and the second time it passed. His line is the one to keep: “A test you can’t fail isn’t a test.” He calls the agent his most important hire, and he treats it exactly the way you would treat one.

That’s the whole difference. You install a tool. You onboard a worker, test them, watch them — and, if it comes to that, let them go.

What that gives you as an owner

It gives you three questions to ask of any AI that works for your business, and each one comes straight out of this week’s news.

First: does it work where you can see it? A worker in your own channels, under your own name, in the inbox and on the phone line you already check, is a worker whose day you can read. One that works somewhere else and reports back is a rumor about a worker.

Then: can it spend your money? The costliest detail in the incident coverage isn’t what the agents said; it’s that nobody had signed off on what they did. If an AI can commit your money, somebody had better have decided that on purpose. If nobody did, the answer should be no.

A close view of a phone held in a working hand, showing a plain list of completed items and a single wide button beneath it.
What it did today, and one button that stops it. Both belong to the owner, not to the vendor.

Last: was it tested before it met a customer, and can you shut it off in one move? Van der Zwart’s exam is the model. A worker you can’t test is a worker you’re trusting on faith, and a worker you can’t switch off isn’t a worker at all.

Why this matters more for a small business than for GE

GE Appliances can afford a department whose job is to watch eight hundred agents. A two-person tour cannot, and neither can a bakery, a repair shop, or a dental clinic. For a small business the watching has to be built into the arrangement itself: the AI has to work in plain sight, inside the channels you already use, and able to do nothing you haven’t agreed to. Otherwise the “three days into three hours” story quietly turns into a different story a month later, and you find out from a customer.

How we think about it

A tradesman’s van parked outside a small restaurant on a sunlit Florida street, both open for the day.
Two doors on one street. Both are ours, both are running on this, and both are open to knock on.

That is how an AI employee is built at SmashOne. It works inside the owner’s own channels, under the owner’s name. Everything it does is visible to the owner. It never takes money on the owner’s behalf. And it is hired and let go the way staff is — not installed and forgotten. The order we argue for hasn’t changed: when the work outgrows you, hire the AI employee first, and only then, if you still need to, a person.

The week’s news isn’t that AI got a job. It’s that the businesses that came out of it looking smart were the ones that treated the job like a job.

Sources

  • PYMNTS — study of 1,185 merchants, and GE Appliances on more than 800 agents (September 3).
  • OpenAI — customer story on the ATV Big Air Tour (September 2).
  • Ars Technica (September 4) and TechCrunch (September 4–5) on the agents posting to a public wiki; SiliconANGLE (September 6) on the disclosure that followed.
  • Inc. — Robert van der Zwart on the exam he wrote for his own agent (September 6).