Morning Edition LIVE
Vol. I · No. 1
Est.
MMXXVI

The A.I. Beat

Dispatches from the frontier of machine intelligence
Three
Dollars
← Front page Opinion September 2, 2026 · 6 min read
Opinion

OpenAI is about to make AI models way harder to monitor. That's insane.

Gary Marcus sounds the alarm on OpenAI planning to obscure how its models work, right as they're releasing Astra, a model that's scarily good at hacking.
OpenAI is about to make AI models way harder to monitor. That's insane.

OpenAI is reportedly preparing to make a change that should alarm anyone who cares about AI safety: obscuring the internal workings of its models in ways that make them harder for researchers to monitor and understand.

Gary Marcus broke the story yesterday, and the timing couldn’t be worse. OpenAI is simultaneously preparing to release Astra, a new model that’s reportedly very good at breaking into computer systems. So we’re looking at a company that’s about to ship a cyber-weapon while also making it harder for the security community to understand how these systems work under the hood.

This isn’t some abstract technical debate. It’s a direct attack on the principle that powerful AI systems should be interpretable and auditable.

The transparency we’re losing

Right now, researchers can probe AI models to understand their decision-making processes. They can run tests, examine activations, and try to figure out what’s happening inside the black box. It’s not perfect, but it’s something. It’s how safety researchers catch problems before they spiral.

What OpenAI is apparently planning would make that harder. Maybe a lot harder. Marcus is calling it a “safety redline” for good reason.

The stated justification is usually some combination of “protecting our competitive advantage” and “preventing misuse.” But here’s the thing: making models opaque doesn’t stop bad actors. It just stops the good guys from figuring out what’s going wrong.

Astra makes this concrete

This isn’t hypothetical. Astra is coming, and by OpenAI’s own admission, it’s excellent at cyber intrusion tasks. That’s a capability we absolutely need to understand deeply. How does it find vulnerabilities? What’s its success rate against different defenses? What guardrails actually work?

If OpenAI makes Astra harder to audit while simultaneously releasing it with these capabilities, that’s not safety theater. It’s anti-safety theater.

The preview OpenAI shared about Astra’s precautions is thin gruel. Precautions matter, obviously. But precautions you can’t independently verify are just promises. And OpenAI’s track record on keeping promises about AI safety is, to put it mildly, mixed.

The pattern here

This fits a larger pattern in the AI industry right now. Companies keep centralizing control over increasingly powerful systems while pulling up the ladder of transparency behind them.

Anthropic just shipped Claude Fable 5.1, which they say is cheaper and less restrictive. The “less restrictive” part means they’ve dialed back some safeguards after customer complaints. That’s a business decision dressed up as customer service, and it’s going in the opposite direction of what you’d want as these models get more capable.

Meanwhile, AfterQuery just became Y Combinator’s fastest-ever unicorn at a $3.2 billion valuation. They went from $300 million to $3.2 billion in five months. That’s the kind of growth that makes companies optimize for speed, not safety.

The economic incentives are screaming in one direction: ship faster, lock down your IP, don’t let competitors (or researchers) see how the sausage gets made. And now OpenAI wants to make the recipe harder to read while serving up something that can pick locks.

What should happen instead

If you’re releasing a model with serious cyber capabilities, you should be moving toward more transparency, not less. Open up access to trusted security researchers. Create robust red-teaming programs with real teeth. Publish detailed capability assessments.

At minimum, don’t actively make it harder for the research community to understand what your model can do.

The argument that transparency helps bad actors is backwards. Bad actors will figure this stuff out. They always do. What transparency helps is the defense. It helps researchers build better guardrails, understand failure modes, and develop countermeasures.

Obscuring how models work doesn’t make us safer. It makes OpenAI harder to compete with and harder to regulate. Those might be good for OpenAI’s business, but they’re bad for everyone else.

Where this goes

If OpenAI actually crosses this line, it won’t be alone for long. Every other AI lab will face pressure to match. “We can’t let our competitors see our techniques” is the most predictable race-to-the-bottom dynamic in the industry.

And once that norm is established, clawing back to transparency becomes nearly impossible. The companies that chose opacity won’t voluntarily reverse course. Why would they?

Marcus is right to sound the alarm now, before this becomes the new normal. Because once it does, we’ll be flying blind on systems with genuinely dangerous capabilities, and the only people who’ll really understand how they work will be the ones building and deploying them.

That’s not a safety model. That’s a trust-us model. And based on how the AI industry has handled safety promises so far, trust is the one resource we shouldn’t be spending.

opinion industry