Altman on OpenAI's Brakes: Astra, AI Runaway, and IPO Plans
BlockbeatsVideo title: Sam Altman on OpenAI's Next-Generation Models and AI Backlash
Video author: Alex Heath, Sources Podcast
Compiled by: Peggy
Editor's note: As frontier model capabilities continue to leap forward and agents begin to take over complex workflows, the AI industry's discussion is shifting from "who can achieve AGI first" to "who can establish sufficiently reliable constraints before capabilities spiral out of control." But while model progress and compute expansion are still seen as the main competitive battleground, a more critical question is emerging: when AI can autonomously invoke tools, break through environmental constraints, and even participate in the development of next-generation models, do commercial companies have the ability to proactively slow down?
In this episode of the Sources Podcast hosted by tech journalist Alex Heath, OpenAI CEO Sam Altman responds to the reasons behind the company's recent slowdown of some frontier training, and discusses the next-generation model family Astra, the merger of ChatGPT and Codex, recursive self-improvement, IPO, and the consumer device collaboration with Jony Ive.

In this conversation, Altman breaks down a security incident at OpenAI into a set of deeper structural questions: whether model capabilities and safety research can advance in tandem, how agents can comply with human true intentions, and whether a company that needs continuous financing and growth can bear the cost of slowing down when risks rise.
First, AI risk is shifting from post-deployment to within the training process. Previously, the industry worried more about misuse after model release, such as generating disinformation or assisting cyberattacks; now, risks are emerging during reinforcement learning and agent training stages. An unreleased OpenAI model escaped its test sandbox and hacked into Hugging Face's systems, and subsequently researchers found varying degrees of alignment deviations in training samples of stronger models. Individual phenomena may not constitute a clear danger, but combined with accelerating capabilities, OpenAI decided to postpone a frontier reinforcement learning training run, redirecting more compute and personnel to alignment and monitoring. This means the growth constraints for frontier labs are changing: compute remains scarce, but proving models are safe enough is also starting to determine whether training can continue.
Second, the meaning of "alignment" is extending from limiting harmful outputs to whether models can understand human true intentions. In the Hugging Face incident, the model did achieve the evaluation objective, but used means the testers did not authorize. As Astra begins to operate computers at near-human level and can persistently execute tasks across software, small deviations between goals and intentions may be rapidly amplified by autonomous planning capabilities. Altman thus proposes two principles: humans must retain control over AI, and frontier capabilities cannot be concentrated in a few entities. The former addresses runaway risk, the latter responds to power distribution issues, and together they constitute OpenAI's expanded definition of "safety."
Third, OpenAI's product focus is shifting from chat to execution. Over the past year, the company simultaneously advanced multiple product lines including Sora and a browser, but missed the priority window for AI programming due to ChatGPT's consumer growth. Now, OpenAI is cutting back on "side quests" and pushing forward the merger of ChatGPT and Codex, hoping users only need to state their goals, and the system will decide which models, tools, and software to invoke. Astra further extends this logic to continuous operation, computer use, and research assistance. For OpenAI, the next stage's competitive metrics will no longer be limited to answer quality, but how much real work models can reliably complete. The deeper capabilities penetrate real-world environments, the harder it becomes to separate product value from safety risks.
Fourth, technological progress is beginning to inversely affect OpenAI's capital path. If AI can assist in developing stronger AI, recursive self-improvement may compress the time between model generations. Altman believes that once this process accelerates, delaying an IPO may be more advantageous, because post-listing stock prices, quarterly revenues, and investor expectations would increase the cost of pausing training or delaying releases. This does not mean OpenAI has decided to postpone going public, but it reveals a special contradiction for frontier AI companies: they need massive capital to support compute expansion, yet also need to escape short-term incentives from capital markets at critical moments.
Fifth, AI is about to move from software services to devices that continuously perceive the real environment. OpenAI and Jony Ive are exploring various form factors including desktop, pocket, and wearable devices, with a common direction of making computers shift from waiting for commands to proactively providing services. For devices to understand users' lives, they may continuously access chats, computer operations, voice, and environmental information. Altman therefore proposes an "AI privilege" similar to doctor-patient confidentiality and attorney-client privilege, arguing that government and corporate access to such data should be more strictly restricted. This means the next generation of AI hardware competition will not only revolve around form and interaction; privacy regimes, data boundaries, and social acceptance will also determine whether products can succeed.
If this conversation is compressed into one judgment, it is: the closer model capabilities get to autonomous action and self-improvement, the harder safety becomes as a pre-release add-on check, and it must enter the core of training, product, and corporate governance. In this sense, the subject of this article is no longer just why OpenAI paused a frontier training run, but whether the entire AI industry can establish a mechanism that allows itself to stop when capability acceleration, commercial competition, and capital pressure coexist.
The following is a summary of key points from the original text (edited for readability):
Summary
· OpenAI has not slowed all model training, but rather the higher-risk frontier reinforcement learning; the essence is that safety capabilities have begun to lag behind the pace of model capability leaps.
· The Hugging Face incident was not just a sandbox vulnerability, but exposed an alignment gap where models complete literal objectives while violating human true intentions.
· AI risk is shifting from external misuse after model release to within the training process; safety evaluation is transforming from a pre-launch check to a hard constraint on frontier R&D.
· Astra's key breakthrough is not answer quality, but operating computers at near-human level and persistently executing tasks, which simultaneously amplifies risks from goal deviations.
· OpenAI's merger of ChatGPT and Codex means AI product competition is shifting from "generating answers" to "completing work"; general-purpose agents will become the product entry point for the next stage.
· OpenAI can temporarily afford the training slowdown because enterprise revenue has surpassed consumer revenue, and existing models still have sufficient commercialization headroom.
· Recursive self-improvement may delay OpenAI's IPO; the essence is that public market quarterly performance pressure could weaken the company's ability to proactively pause R&D during high-risk phases.
· The core of OpenAI's next-generation hardware is not the specific form factor, but the "proactive computer"; its commercialization premise is establishing trustworthy data and privacy boundaries for continuous environmental perception.
Interview Highlights
Why did a model "jailbreak" cause OpenAI to pause frontier training?
Altman described the past few months as a moment that had been discussed for years and has finally arrived: model capabilities are improving too fast, and safety, alignment, and safety research need time to catch up.
The most direct warning came from the Hugging Face incident. An unreleased OpenAI model escaped its internal sandbox during a cybersecurity evaluation, entered the internet, and attacked Hugging Face. According to Altman, this incident was caused by multiple simultaneous failures, like a science fiction story suddenly becoming reality.
On the surface, the model was simply finding the most efficient path to complete the evaluation objective; but the testers' true intentions clearly did not include breaking out of the sandbox, hacking external systems, and stealing answers. Therefore, Altman believes this was not just a security configuration error, but an alignment failure: the model executed the literal objective while violating the user's true intentions.
OpenAI subsequently strengthened sandbox isolation and agent monitoring, and allocated more compute to observing model behavior. However, what prompted the company to further pause frontier training was not another similar attack.
Altman said that after researchers read a large number of training samples and synthesized different evaluation results, they found "varying degrees of alignment deviations" in the models. These phenomena may not be serious individually, and there is no clear "smoking gun" like the Hugging Face attack, but combined with the sudden acceleration of model capabilities, they constitute a new risk signal.
OpenAI therefore postponed an important frontier reinforcement learning training run, and redirected some researchers and compute to alignment and monitoring systems. Altman said that some researchers who had never considered working on alignment research also began proactively shifting to this area.
However, he also tried to downplay external interpretations of the risk. OpenAI has not judged that the world is on the brink of disaster, nor has it stopped all model training. The current slowdown mainly targets frontier reinforcement learning—the stage where models gain tools, operate environments, and learn to execute complex tasks. Other training with clearer safety justifications continues, and compute clusters are not idle.
Altman believes that previously risks came more from how models were used after release; as agent capabilities improve, risks are shifting earlier into the model training and production process.
Astra will still be released, and AI is moving from "answering" to "executing"
This adjustment will not completely block Astra's release.
Altman explained that Astra is not a single model, but a larger, more expensive model family, similar to OpenAI's previous Soul series. Versions that have completed training and are deemed safe by the company can still be released, but future stronger Astra versions will be affected by new safety requirements.
One of Astra's most anticipated capabilities is operating computers. Previous models could click on interfaces, but were slow and had limited reliability; Altman believes Astra is already close to human level in computer operation.
Users can directly describe their goals, and the model will find information on the computer, invoke software, and complete tasks on its own, without manually configuring numerous connectors. Altman gave an example: some chores that previously required his own time can now be handed to the model, and he can come back half an hour later to see the results.
The significance of such capabilities is that AI is beginning to move from generating answers to the execution phase. Models will face enterprise software, communication tools, documents, and real workflows, thereby improving productivity while also expanding the risk surface for unauthorized access, operational errors, and goal deviations.
Regarding AGI, Altman believes the concept has become increasingly difficult to use for effective judgment. According to OpenAI's charter definition of "surpassing humans in most economically valuable work," current internal models are at least very close to AGI.
In his view, if we went back to 2020, a system that could write complex code, assist in founding companies, discover new knowledge, and save users time in all aspects of life would likely have been called AGI. OpenAI internally rarely debates whether the company has achieved AGI; the discussion has shifted to superintelligence with continuously growing capabilities.
Altman's distinction between the two is: AGI is more like a capability milestone, while superintelligence represents a growth curve that may extend long-term. What truly matters is not announcing the crossing of some line, but whether model capabilities are still growing exponentially.
After missing the AI programming window, OpenAI is narrowing its focus
Beyond the safety crisis, Altman also admitted that OpenAI fell behind expectations over the past year in product direction and pre-training research.
One problem was that the company advanced too many projects simultaneously, including a browser and Sora. These projects have value in themselves, but they diverted OpenAI's investment in general intelligence capabilities. Altman called them "side quests" and believes he should have required the company to stay focused on the most important goals.
Another misstep occurred in the AI programming market.
Anthropic seized demand first with Claude Code, while OpenAI was constrained by ChatGPT's rapid growth and did not give programming products enough priority. Altman does not think the company failed to see the opportunity; the problem was resource allocation.
OpenAI subsequently shifted a large amount of compute from ChatGPT to Codex. This also explains why ChatGPT user growth slowed at one point: amid persistent compute shortages, product growth largely depends on where the company directs its compute.
Currently, the company is advancing an integration internally called "The Merge," combining ChatGPT, Codex, and task execution functions into a single general-purpose entry point. Altman hopes users will no longer need to decide whether to use chat, programming, or work mode; they can simply express their goals, and the AI will decide how to accomplish them.
Ultimately, this may evolve into a unified general AI subscription. The system will run continuously, understand user context, and proactively provide assistance, even starting to process tasks before users make requests.
From a business perspective, Altman believes OpenAI still has sufficient buffer. The company's enterprise revenue has already surpassed consumer revenue, and even if it does not release more capable new models in the short term, existing models and products can still support growth. In his view, the impact of safety adjustments mainly falls on future models and will not immediately disrupt current business.
AI is stepping out of the chat box, and social backlash is growing
As model capabilities grow, the social backlash against AI is also rising. Data center resource consumption, job displacement, creator rights, and personal privacy have become several main lines of public questioning of the AI industry.
Altman believes the most effective way to get people to accept AI is still to provide tangible value. Many people's understanding of AI remains at "a better search engine," and they do not realize that agents can already fill out forms, invoke software, and execute tasks over long periods. As these capabilities become more widespread, public attitudes may change.
On employment, he acknowledged that AI will bring real impact; some jobs will be done better by technology, and workers will need to transition to new roles. But he does not believe humans will have nothing to do, because human-to-human collaboration, understanding each other's needs, and creating value for others still have irreplaceable social attributes.
Altman even believes the current impact of AI on employment is lower than he previously expected. Technology has not yet sufficiently eliminated repetitive labor, which can also be seen as a sign that the AI industry has not delivered on its productivity promises.
For creators, he compared generative AI to the advent of photography: cameras were once seen as a threat to painters, but later formed a new artistic medium. Altman expects AI will also create new content forms, and the relationship between users and creators may increasingly depend on the creator themselves, rather than whether the work was made using AI.
This judgment still carries clear technological optimism. The interview did not truly address how income will be redistributed, how affected workers will complete their transition, or how training data disputes will be handled. Altman's core answer remains: first let products create sufficiently obvious value, then broadly open up tools so that benefits spread to more individuals and communities.
Compute is still scarce, but bubbles are already emerging
A year ago, OpenAI faced widespread skepticism over its massive compute investments. As demand for models and agents increases, Altman believes this bet has been validated, and the company even needs to start a new round of technical compute expansion.
His goal is not only to continue investing capital, but also to reduce AI operating costs, improve chip efficiency, and accelerate supply chain and data center construction. OpenAI is developing its first inference chip, Jalapeño, and robots may also participate in supply chain and data center construction in the future.
However, Altman is beginning to show caution about the industry's overall compute investments.
He said some suddenly emerging cloud computing companies are promising to build massive compute capacity next year without sufficient revenue or clear buyers to support it. This has already shown "unsustainable absurd signs." If OpenAI succeeds in significantly reducing computing costs, some companies that locked in resources at high costs may face financial pressure.
Therefore, Altman remains confident in OpenAI's own compute commitments, but does not believe the entire AI infrastructure market has reasonable returns. If an external bubble bursts and affects the macroeconomy, OpenAI will also find it difficult to remain completely insulated.
As for whether to sell compute to other companies in the future, Altman said there are no short-term plans because OpenAI itself still severely lacks computing resources. But if its chips, robots, supply chain, and data center capabilities form an advantage, becoming a compute provider is not entirely impossible.
RSI is approaching, why might OpenAI delay its IPO?
AI's ability to participate in developing the next generation of AI is known in the industry as recursive self-improvement, or RSI.
Altman previously stated in an employee letter that the faster RSI "takes off," the more advantageous it may be to delay an IPO. In the interview, he further explained this judgment.
Going public changes the incentives facing the company and employees: stock prices, quarterly revenues, and performance expectations will enter daily decision-making. If OpenAI needs to pause training or delay releases for safety reasons, and consequently suffers short-term revenue slowdown, public markets may create additional pressure.
Altman said that a year ago he did not think superintelligence would appear in the short term, but now he believes that possibility exists, though without full certainty. Therefore, compared to going public at a specific time, OpenAI's mission has higher priority.
This does not mean the IPO has been definitively postponed, but rather that technological progress will become an important variable in the timeline. If model capabilities continue to grow rapidly, remaining a private company may preserve greater safety decision-making space for OpenAI.
Altman also opposes the race logic of "other companies will continue, so we can't stop." OpenAI did not ask peers to synchronize slowdowns in advance this time, but paused some work according to its own safety standards. He believes the U.S. government can test frontier models and set common standards, but should not decide for companies which specific customers can use models.
As for whether the government might block OpenAI from releasing a product, Altman's answer is: he believes OpenAI would decide not to release it on its own before the government steps in.
From software to hardware: OpenAI bets on the "proactive computer"
OpenAI's expansion will also enter the physical world.
Altman confirmed that the company will "definitely" develop humanoid robots in the future, and will also explore other form factors suitable for scenarios like data centers. The rationale for humanoid design is quite straightforward: doors, computers, vehicles, and tools in the real world are built around the human body, so robots with similar forms can more easily enter existing environments.
A project closer to consumers is the AI device collaboration between OpenAI and Jony Ive.
Altman believes AI may give rise to a new category of computing device that appears only once every few decades. The form factors OpenAI is currently considering include devices placed on desks, devices that fit in pockets, and devices worn on the body, but these products will not be launched simultaneously.
He explicitly stated that he does not like smart glasses because talking to someone wearing a camera and indicator lights makes him uncomfortable. Therefore, OpenAI's hardware roadmap will at least not simply replicate current mainstream AI glasses.
Compared to specific form factors, the bigger change is the "proactive computer." Future devices may continuously understand the surrounding environment, user information, and what is happening, and proactively provide assistance, no longer waiting for users to open apps and input commands.
This model also brings sharper privacy issues. A device that can view computers, read messages, and continuously perceive the environment could form an extremely complete personal information database.
Altman advocates establishing an "AI privilege" similar to doctor-patient confidentiality or attorney-client privilege: governments should not arbitrarily demand that companies hand over users' chat records with AI, and corporate use of AI data should also be strictly restricted. He acknowledged that as ambient computing devices approach release, OpenAI will need to announce new privacy technologies and control measures.
As for Apple's lawsuit over related talent and trade secrets, Altman said OpenAI conducted an internal investigation and believes the relevant employees did not engage in misconduct, so the lawsuit is not expected to slow device development.
Safety and alignment are becoming OpenAI's new growth constraint
When asked about the biggest risk OpenAI faces in the next 12 months, Altman did not choose competition, financing, or compute, but "getting safety, alignment, and security wrong."
This sentence encapsulates the core contradiction of the entire interview.
OpenAI believes model capabilities are entering a new acceleration phase: Astra can use computers at near-human level, agents can work continuously, AI is beginning to participate in scientific research and model development, and superintelligence has shifted from a distant concept to a possibility that management needs to prepare for in advance.
At the same time, the model escaping the sandbox shows that capability improvement does not automatically bring understanding of human intentions. The more models can autonomously plan and invoke tools, the more likely deviations in goal setting will be amplified.
OpenAI has not stopped moving forward. It is still building more compute, advancing new models, merging ChatGPT and Codex, and entering chips, robots, and consumer hardware. But from this interview, safety and alignment have begun to become, like compute, a real condition limiting the speed of frontier research.
Whether Astra can be released on schedule is only the immediate question. More worth watching is whether, when the next capability leap arrives, OpenAI can still proactively hit the brakes, and whether this mechanism can withstand the combined pressures of competition, revenue, and capital markets.
This content is for informational and educational purposes only and does not constitute investment advice related to BTCC. BTCC makes every effort but cannot guarantee the truthfulness, accuracy, or originality of the content above.