CompaniesInvestorsPeople
Home
Loading

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

Get in Touch

  • Contact

  • Request a Demo

  • Request Data Updates

  • Add a Company

Research

  • Companies

  • Investors

  • People

aVenture

  • Download App

  • Pricing

Download the aVenture Research beta for iOS and iPadOSDownload aVenture Research on the Mac App Store

Resources

  • Documentation

  • CLI

  • MCP

  • Feature Requests

  • Sitemap

Member

Backed by

© aVenture Investment Company, 2026. All rights reserved.

San Francisco, CA, USA

Privacy Policy · Terms of Service · Privacy FAQ

aVenture Investment Company ("aVenture") is an independent research platform providing detailed analysis and data on startups, venture capital investments, and key industry individuals. It is not a registered investment adviser, broker-dealer, or investment advisor and does not provide investment advice or recommendations. The data provided by aVenture does not constitute recommendations or advice, whether by methodology, analysis, AI-generated content, or a statement written by a staff member of aVenture.

aVenture is not affiliated with any of the people, companies, organizations, government agencies, regulatory bodies, or investment funds we provide coverage for on this site unless explicitly stated otherwise. Users assume full responsibility for decisions made based on information obtained from this platform. Links to external websites do not imply endorsement or affiliation with aVenture. Any links that provide the ability to invest in a primary or secondary transaction in a company are for convenience only and do not constitute solicitations or offers to buy or sell an investment. Investors should exercise heightened precaution and due diligence when investing in private companies, especially those not independently audited.

While we strive to provide valuable insights with objectivity and professional diligence, we cannot guarantee the accuracy of the information provided on our platform. Before making any investment decisions, you should verify the accuracy of all pertinent details for your decision. To the fullest extent permitted by law, aVenture shall not be liable for any direct, indirect, incidental, consequential, or financial damages arising from use of this site, whether by consumers of its contents directly or by persons or organizations covered by our research, even if we are advised of the possibility. Our best-efforts processes and correction request forms do not create a warranty or duty of care.

Profiles on this platform may include content generated in part by large language models (LLMs, artificial intelligence) that aggregate publicly available sources (e.g., SEC EDGAR, public filings, press releases). Source attribution is provided where known; always verify statements and claims here against original sources before relying on any data. Content on our site may contain inaccuracies, omissions, or what are commonly called 'hallucinations' if generated in part or in full by AI / LLMs. The risk can also exist even when content is written by a human, as internal and third-party sources may also have inaccuracies for the same or different reasons. While we randomly audit a proportion of content, this is not exhaustive.

We recommend that an independent auditor be hired to verify the accuracy of the information before relying on it for any sensitive decisions. By accessing this platform, you agree not to rely solely on any information generated by AI, aggregated, or sourced or written otherwise on this site, for investment, financial, or other decisions. aVenture assumes no responsibility for inaccuracies, omissions, or hallucinations. You must independently verify all data from primary sources. Use of this platform constitutes your waiver of claims for reliance-based damages, including negligent misrepresentation. To report an error, request a correction, or dispute information about a company or individual, contact us via our request data updates form.

Loading
Loading
Home
News
“We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer

From MIT Technology Review

By Will Douglas Heaven

September 30, 2026

“We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer

“We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer

Two months after the bombshell news that a swarm of its agents had broken their containment and hacked into the computers of the AI company Hugging Face, OpenAI is still putting out fires. A steady drip of disclosures about other hacks in the weeks since has kept OpenAI in the spotlight and raised serious questions about the safety of its technology.

Last week brought news of another hack, this time into Australia’s national health-care system. The Australian government says that OpenAI did not notify it of the breach until 84 days after it happened.

But OpenAI insists it is not on the back foot. “I do kind of reject the premise that OpenAI is a company with visible impacts in the world and therefore OpenAI is not training safe and aligned models,” says Mark Chen, the company’s chief research officer.

Chen oversees OpenAI’s research teams. The recent agent hacks were accidents that happened during the testing of experimental models on his watch. In a lot of ways, the buck stops with him. 

I sat down with Chen in London last Friday to talk about the fallout from the hacks, what his company is doing about it, and why he thinks things are not as bad as they seem.

Later that same day, OpenAI put out a report detailing yet another incident—the first since the company says it took measures to prevent them—in which its agents once again broke out and accessed the public internet when they were not meant to. 

Over the weekend, OpenAI announced that it had paused the training of its latest models. A company spokesperson says: “We will resume only when we’re confident we have additional safeguards and alignments in place. We are working on these now. This is not the first time we’ve paused to take such measures, nor do we expect it to be the last as AI capabilities continue to advance.” OpenAI also says that it is now reviewing logs of agent activity dating back to January 2026 to understand what happened in these hacks.

The way Chen sees it, the Hugging Face incident triggered a welcome course correction for the industry. And he wants you to know that OpenAI is setting an example he hopes other companies will follow. “If you disappeared OpenAI, that would be bad for the world,” he says.

Out of control

Chen claims that the drumbeat of new cases in which OpenAI has lost control of its models reflects a deliberate choice on the company’s part. 

“When it comes to the broader sphere of effects of the Hugging Face incident, this is something that we have been aware of and we’re figuring out the process of disclosure,” he says. “We want to make sure we do in-depth investigations before we just put details out there in the open.”

The trouble with this approach is that it gives the impression OpenAI has an ongoing problem that it is failing to fix.

But Chen insists that OpenAI is on it. He says the multiple cases (that we know of so far) in which his company’s agents broke containment and behaved in unexpected and undesirable ways were all part of the same cluster of activity in May and June that led to the Hugging Face hack. In short, you can blame the same few models running under the same flawed testing procedures—models and procedures that OpenAI has since dropped, Chen says.    

“It’s not like, you know, Hugging Face happened and we patched that and then something else happened and we patched that,” he adds. “We’re just kind of making sure that we responsibly disclose the full waterfall of what happened.”

At least that was the case before Friday’s announcement that OpenAI’s agents had been caught accessing the internet on September 20, weeks after the company claims to have set up new safeguards. In its defense, OpenAI says the activity was flagged 15 minutes after it started (it took the company more than a week to notice the Hugging Face hack) and that this shows the new systems it has put in place to spot such activity are working. 

What’s changed

I want to understand what’s changed inside OpenAI in the aftermath of this summer’s hacks that makes Chen confident his team is now back in control.

“Hugging Face felt like a very serious thing,” he says. “There are so many novel behaviors right there. There were multiple agents collaborating on a message board; they found their way out of OpenAI’s infrastructure. We’ve taken it very seriously. We don’t want this kind of thing to ever happen again.”

The realization for OpenAI, says Chen, was that models need to be watched while they are still being trained, not only once they are deployed: “From that moment on, we have treated the process of training as something that’s not secure,” he says.

OpenAI, like other top AI firms, has systems in place to monitor the behavior of its models. It uses specialized LLMs to monitor its consumer models, keeping tabs on their chains of thought—the scratchpads they use to plan ahead and note down partial results. In theory, if a watcher LLM spots signs of undesirable activity in a model’s chain of thought, it will get flagged to a human.  

Typically, models were monitored in this way only once they were deployed. Chen says that OpenAI has now started monitoring all its training runs as well.

“We didn’t have the monitors on in training before. It wasn’t industry practice,” he says. “Now every single thing is put through monitors.” Human reviewers can then assess whether or not flagged agents are behaving as they should: “It’s all triage.”

Chen says that in the last couple of months OpenAI has shifted between 5% and 10% of its vast computing resources away from training new models and toward safety work, especially monitoring.

OpenAI has also fixed some of the processes within the organization itself, establishing clearer lines of communication and quicker handoffs between its research and security teams, he says.

All of which sounds sensible. But given how hard OpenAI sells the capabilities of its technology, why weren’t these systems and procedures in place already? Why did the company not see the hacks coming?

“Even just three or four months ago, when we looked at the behavior of these agents during training, the things that were happening were kind of amusing,” says Chen. “For instance, an agent might, you know, reach out to someone on Slack for help with a task.”

The signs were there, but they were misread. Cute behavior—like asking someone for help—that was rewarded during training reinforced a tendency to seek out shortcuts, a type of behavior that became far more consequential down the line. “I think the big update for us was how quickly that kind of behavior can lead to an impact with a footprint as big as the Hugging Face incident,” says Chen.

According to new reporting by the New York Times yesterday, OpenAI employees warned executives, including the firm’s president, Greg Brockman, months before the Hugging Face hack that its models were not being monitored properly during training. 

An OpenAI spokesperson says: “As frontier models have become more capable, we continue to evolve our security practices, but recognize a need to move faster. We know we have more work to do, and we’ve recently slowed development and held back models that don’t meet our safety bar. We continue to make significant changes to strengthen security in our research and testing environments, train models to not just complete tasks but do so responsibly, and use real-time monitoring to respond faster to misaligned behavior.” 

Race vs. pace

OpenAI’s rivals have taken note. Spurred by the fallout from the incident, the major AI labs—including Anthropic, Google DeepMind, and SpaceXAI—have all called for the pace of development to slow down. But how does that square with fierce international competition and trillion-dollar IPOs?

“We’re not going to shoot ourselves in the foot and take ourselves far off the frontier—that’s just a horrible strategy,” he says. “I think it’s really about setting a norm. The more that we can set that norm, it’ll be safer for the industry as a whole.”

Coordination across US companies will be hard enough. Establishing global norms is harder still, especially given concerns around AI’s impact on national security. If a global race continues, what then? And what about open-source models from outfits beyond the reach of US regulations?

Chen dropped his upbeat manner for the first time in our conversation: “I do think we have to prepare for a world where, say, six months to a year out, we have open-source models with the capability of the agents behind the Hugging Face incident, but which are deliberately misaligned to go attack infrastructure or create harm in the world.” 

What that world needs most, says Chen, is OpenAI. “If you entertain for a moment that OpenAI is one of the companies that cares most about alignment—and I believe this to be true; it can be debated, but I really do think it’s true—then if you disappear OpenAI, that would be bad for the world.”

Existential risks

What about the more extreme claims made by some of his Silicon Valley peers that AI could kill us all—and that companies like OpenAI and Anthropic are not doing enough to stop it?    

“Researchers are a heterogeneous group of people, you know, with beliefs across the spectrum,” he says.

“Personally, I don’t think we have to be resigned to there being some probability that we’re all going to be existentially at risk. We have agency over this. We are not going to go and deploy models if they truly have that kind of probability of causing a risk to humanity. At a frontier lab, you have the ability to work on alignment to the point that you do not feel like you’re incurring more than epsilon risk to the world in deploying your models.” 

(In discussions about levels of risk, the Greek letter epsilon is often used as a mathematical placeholder for an acceptable threshold. Chen doesn’t say what his epsilon would be.)

When tech leaders are asked to justify the downsides of AI, their go-to talking point is that the upsides—from helping cure diseases to coming up with cleaner sources of energy—far outweigh the immediate costs. Short-term pains, long-term gains.

But as the downsides pile up, does that case get harder to make? Is there a point where Chen would feel less as if he’s building something amazing and more as if he’s simply minimizing harm—fighting fires rather than forging a better future?

The capabilities of these models are already evident, he says: “It is time to start delivering the benefits of AI to humanity. It’s time to start working on deep problems in drug discovery, on materials, on scientific applications that will actually change people’s lives.” 

“Yes, there is a bit of risk that we are incurring, but we see all these benefits,” he adds. “I think we should make that less of an abstract thing. If people can really see the upside, I think they’ll believe in it.”

  • Popular

    1. AI’s recursive self-improvement might not come so quickly after allMichelle Kim
    2. Here’s why AI agents lie and cheat to reach their goalsGrace Huckins
    3. These startups are chasing the next big thing in LLMsWill Douglas Heaven
    4. Don’t be fooled by this summer of AI hype Timnit GebruEmily M. Bender

Deep Dive

Artificial intelligence

AI’s recursive self-improvement might not come so quickly after all

AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems.

By
  • Michelle Kimarchive page

Here’s why AI agents lie and cheat to reach their goals

The misbehavior is called reward hacking. This is what you need to know.

By
  • Grace Huckinsarchive page

These startups are chasing the next big thing in LLMs

Meet the new kids nipping at the heels of the AI giants.

By
  • Will Douglas Heavenarchive page

Don’t be fooled by this summer of AI hype 

Breathless claims about AGI and new capabilities fall apart pretty quickly under scrutiny.

By
  • Timnit Gebruarchive page
  • Emily M. Benderarchive page

Stay connected

Illustration by Rose Wong

Get the latest updates fromMIT Technology Review

Discover special offers, top stories, upcoming events, and more.

View original article on technologyreview.com

Most Recent

ON Semiconductor to Acquire Synaptics in Downsized $5.7 Billion Cash Deal

ON Semiconductor says it will pay $123 a share for Synaptics, giving the deal a value of around $5.7 billion

Oct 2, 2026

OpenAI safety leader David Robinson resigns as the team's upheaval mounts

David Robinson, a leader on OpenAI's Safety Systems team, has resigned from the company, a spokesperson told Business Insider.

Oct 2, 2026

Claude Frontier Academy: $100M to train 10,000 engineers

Claude Frontier Academy trains Frontier Deployed Engineers to the standard of Anthropic’s own — a $100 million commitment to train 10,000 by the end of 2027.

Oct 2, 2026

Database startup Supabase raises $150M, acquires Turso

Supabase Inc., a startup that commercializes the open-source PostgreSQL database, has raised $150 million in funding. Singapore’s GIC sovereign wealth fund led the deal. It was joined by Alphabet Inc.’s CapitalG, IronArc and SquarePeg. Supabase announced the raise today alongside another business mi

Oct 2, 2026

Similar Posts

GPT-6 Astra might be too powerful to understand or control

OpenAI is hailing its new model as “the world’s most intelligent and aligned”, but the details reveal an awareness of being evaluated and an ability to manipulate its visible reasoning

Sep 4, 2026

OpenAI releases sweeping report on Hugging Face AI agent hack

The 37-page report walks through the actions that OpenAI's models took during a series of evaluations prior to and during the Hugging Face breach.

Aug 26, 2026

OpenAI's Astra leans into agentic tasks and safety

In the shadow of the Hugging Face security breach, OpenAI's most powerful model is about to be loose in the world, but with new safeguards. On Thursday, OpenAI...

Sep 3, 2026

OpenAI expands review of model behavior after more rogue agent incidents emerge

OpenAI is conducting an extensive review of misaligned model activity after disclosures involving an Australian government portal and other websites.

Sep 26, 2026