Nine troubling new reports in seven days
By Gary Marcus — via Substack
For something like a week, there been a constant stream of reports about OpenAI’s apparent misconduct: around the hacks, around apparent coverups, possible data fudging — and maybe even intellectual theft with a hint of extortion. Not to mention a bullshit game in which, in (apparent) concert with Jensen, OpenAI tried to pretend Astra was tantamount to AGI.
The story that broke this morning—still waiting for OpenAI’s full response—blew my mind. And yet also left me unsurprised.
Here are nine things I learned just in the last week:
- The widely discussed and worrisome Hugging Face incident wasn’t the only one. Their software also hacked a German website.
- There are indications that they knew about the German incident weeks ago and didn’t report it. Instead, they appear to have covered it up. Per Reuters,

- As Shakeel Hashim summarizes, their choices put others (including Hugging Face) at risk.
I hadn’t quite put together that OpenAI appears to have *known* about the wiki incident before the Hugging Face hack happened.
They knew their models were capable of breaking out and functioning as a swarm with real-world consequences, and evidently didn’t do enough to stop it…
this answer is very unsatisfying
this wiki incident adds crucial context: openai knew about this kind of swarm/message board behavior ~3 weeks before the hf hack as they stopped this swarm of agents as the authors show with openai IP addresses
a big part of the misaligned https://t.co/2cwbF1yZg0
— elie (@eliebakouch)
— Shakeel (@ShakeelHashim) · on X
- It’s come to light that they withheld information from Congress:
32 members of Congress wrote a letter to OpenAI on August 10th after HF, asking OpenAI if they were any similar undisclosed incidents. This seems like it would have been a good time for OpenAI to share information about the wiki incident. OpenAI declined to do so.

When I wrote a letter with @GregCasar to OpenAI last month (post Hugging Face), we asked whether other similar incidents had occurred.
They refused to tell us. Then this comes out.
The House Democratic Majority arrives in January.
I look forward to the public hearings.
— Pat Ryan 🇺🇸 (@PatRyanUC)
— Nathan Calvin (@_NathanCalvin) · on X
- When Altman was asked on July 29 whether there were other incidents, he was coy, saying “there could be” when he likely already knew for a fact that there had been.
- OpenAI’s President, Greg Brockman, has been on a campaign trying to imply that Astra is AGI (or one model away from being so).
we’re now moving into the AGI era (whether you view it as this model, the last one, or the next one), and could not do it without close partners
@ChaseLochmiller @OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
AGI has arrived. Congratulations @OpenAI team.
400K GPUs coming online next.
— Jensen Huang (@JensenHuang)
— Greg Brockman (@gdb) · on X
Multiple reports, even from people who are very AI bullish, like “@scaling_o1” on X, suggest that this is marketing bullshit. In one tweet, he wrote “after having burnt twice through my Pro limits using Astra exclusively, I can say that I don’t feel the AGI”, in another he wrote this:
it’s quite awkward how OpenAI tries to claim that it’s AGI
— Lisan al Gaib (@scaling01) · on X
AI forecaster Peter Wildeford, often more bullish about AI progress than me, wrote this:
My current view is that GPT 6 Astra is not meaningfully better than Fable 5.1 for my personal work, but that using both side-by-side is nonetheless very helpful and additive.
I have been using GPT 6 Astra and Fable 5.1 a bunch over the past two days, largely for policy analysis,…
— Peter Wildeford🇺🇸🚀 (@peterwildeford) · on X
Jun Song, an AI developer, wrote this

Others were downright bitter:

and

Some were just funny, like GPT-6 Astra’s video rendering of a Cheetah in motion.
Artificial Analysis find the whole Claude vs Astra thing a tossup, concluding “[O]verall scores reflect different strengths. Claude Fable 5.1 scores higher on AA-Briefcase and SciCode, while GPT-6 Astra scores higher on Terminal-Bench v4.0 and AutomationBench-AA Score.” The idea that Astra is in some completely new league is preposterous. (Jensen should be ashamed for having fed into it. It will be interesting to see if his credibility takes a hit.)
On a personal note, as a scientist I am especially offended by “we’re in the AGI era” PR blitz because it both distorts and dismisses all the thinking that earlier scholars and technologists like Goertzel, Legg, and Voss (and later folks like Bengio, Hendrycks, and myself) did in formulating the very concept of AGI. More about that in a future essay.
- Much was made of their scoring 99.9% on ARC-AGI-3, but apparently that required an in-house (and likely special-purpose) harness. It’s not a number that ARC-AGI themselves could achieve with the out-of-the-box model:
OpenAI published 99.9%. The benchmark’s own software returned 62.7%. Same model, same test, different scaffolding around it. ARC Prize printed both numbers and says it is not claiming AGI. thenextweb.com/news/openai-as…

— TNW (@thenextweb) · on X
- This is not even the only game with numbers they played in the last few days:
OpenAI quietly altered the performance metrics it reported for GPT-6-Astra in ways that favored the new model, and it continues to change others post-launch. Good reporting here from @EmilyForlini
— Jeremy Kahn (@jeremyakahn) · on X
- They may have ripped off a pair of leading mathematicians, one at NYU Courant, the other at Anthropic, on a Millennium Prize Problem:

The last highlighted bit, a quote from OpenAI to one of the mathematicians, smacks of extortion. You can read Tristan Buckmaster’s full account of his unpleasant interaction with OpenAI here. Katie Miller (wife of the White House’s Stephen Miller) suggests it’s all part a pattern.
We do need to hear OpenAI’s side of this story (which they say is forthcoming) but it is absolutely-no-scientist-will-work-with-them-again level bad if true.
I may even have missed some other recent charges of misconduct. And who knows what else they haven’t disclosed.
And it all reeks of desperation: a desperate, money-burning company running out of time and doing everything it can to manipulate the narrative before its planned IPO.
§
In their lawsuit with Elon Musk, OpenAI was saved from further judgment by a statute of limitations. But the testimony — from all the diverse people who called Sam Altman a liar, from former employees like Murati and Sutskever to the former board members — was damning. So was OpenAI President Greg Brockman’s own diary.
Nobody should be surprised by now that they play fast and loose.
I would not expect to see this much (apparent) foul play at a used car dealership; or least not at one that expected to stay open for long. In a company that could “cause significant harm to the world”, in Sam Altman’s own words, it is unacceptable.
I have said it before and will say it again: OpenAI should be shut down until there are changes at the top; both Altman and Brockman should go.
A further sign of a trouble corporate culture, again pointing to the top, is that more at least sixteen top executives, from the Heads of Science, Robotics, Sora, Safety, and Preparendness, to the AI ethics lead, to the Chief Revenue Office and Chief Operating Officer, CEO of AGI Deployment and Chief Revenue Officer, have departed since January, despite the looming IPO.
In taking no action, OpenAI’s board is being absolutely irresponsible.
P.S. Headline/PSA from the San Francisco Chronicle:

P.P.S. Passing along this AI-generated timeline, perhaps approximately correct but not fully accurate, of the HF incident. Tl;dr: they knew for months and kept going and kept quiet.
@GaryMarcus @_NathanCalvin @GarrisonLovely @ShakeelHashim I made this the other day. AI generated
— GS-InfoSec (@GsInfosystems) · on X
PP
