Adam Mills
Newsweek 26 Jun 2026
Several executives described the same basic standard: Review should rise with risk. A tool that drafts an internal memo does not need the same oversight as one that affects a payment, security decision, health recommendation, legal right or compliance record.
Badhwar said the review threshold should be tied to what the error could touch.
“Anything touching money, security, regulated decisions, or anything that’s hard to reverse gets a human, regardless of how confident the model is or how high the aggregate benchmark sits,” he said.
Expert review is also needed earlier, Dai said, when companies are deciding what to measure. A poorly designed evaluation can make a system look safer than it is.
“The quality of the evaluation framework directly affects whether the deployment decision is valid,” Dai said.
Clark pushed the concern into procurement. As capable, inexpensive models flood the market, he said, many buyers cannot answer basic questions about who built them, what they were trained on, how their behavior was shaped or where company data flows.
“You wouldn’t onboard a supplier into your physical supply chain without knowing its origin,” Clark said. “The same diligence has to apply to your AI supply chain—especially in regulated and federal environments.”
The concern has already moved beyond internal AI committees. This month, the U.S. government ordered Anthropic to block foreign nationals from accessing its Fable 5 and Mythos 5 models, citing national security concerns; Anthropic said it disagreed with the decision but disabled access to comply.
Kurtzig said companies need to draw clearer boundaries around where AI can act on its own and where the consequences of a wrong answer are too high.
“’Mostly right’ is the wrong bar,” he said. “A single bad answer can outweigh a thousand good ones, so the real question is not how often the model is right. It is what happens when it is wrong.”
Enterprise AI adoption is likely to keep moving quickly. The companies in the strongest position may not be the ones that simply pick the highest-scoring system. They will be the ones that can say where a tool works, where it needs review and where it should not act on its own.
Clark framed testing as the thing that lets deployment move faster, rather than the thing that slows it down.
“Speed is the payoff for testing the model against the job it’ll actually do, in the environment it’ll actually run, before you ship it—not after,” he said.