Nat Rubio-Licht
Senior Reporter

Nat Rubio-Licht

Nat Rubio-Licht is a Senior Reporter at The Deep View. Nat previously led CIO Upside, a newsletter dedicated to enterprise tech, for The Daily Upside. They've also worked for Protocol, The LA Business Journal, and Seattle Magazine. Reach out to Nat at [email protected].

Gemini 4 puts Google back in the frontier AI race

As its two biggest rivals continue to one-up each other with new frontier models, Google is making its biggest move of 2026.

On Wednesday, the search giant unveiled Gemini 4 Argon, its latest flagship frontier model. The company said that Argon offers frontier performance in a number of complex workflows, including in software engineering, knowledge work, cybersecurity, and domains such as legal and finance.

Google will begin rolling out the models specifically to cyber defenders in the Fairwind program, its restricted-access AI cyber program that it launched in early September. The company is also currently going through the US government's voluntary pre-release model checks before opening up the model to the general public.

Google noted that Argon surpassed its previous generation, Gemini 3.8 Flash Cyber, in cybersecurity tasks such as real world-vulnerability discovery and penetration testing.

The company said that Argon sets a new state of the art score on DeepSWE v1.1, the benchmark testing performance in real-world long-horizon software engineering tasks, sweeping OpenAI's Astra and Anthropic's Claude Fable 5.1 and Opus 5.5.

  • Google's Argon also outperforms these competitors in benchmarks for knowledge work, long-context tasks, and computer use.
  • Specifically, for knowledge work, the company said that Argon leads on the Vals Index, which measures impact in finance, coding, legal, and tax work, and offers state-of-the-art performance in visual understanding tasks, such as chart, document and video analysis, measuring by the LVBench for multimodal understanding.
  • However, Argon is still beat by Astra in FrontierSWE v2, which evaluates agents on complex, multi-hour technical tasks and Terminal-Bench Science for scientific workflows. Opus 5.5 also beats Argon on PostTrainBench for machine learning engineering, and Terminal-Bench 4.0 for tasks within sandboxed command-line terminal environments
[@portabletext/react] Unknown block type "twitter", specify a component for it in the `components.types` prop

In its announcement, Google noted several ways in which Argon is "fundamentally changing" its own workflows, including helping its quantum researchers optimize algorithms, improving memory efficiency, and handling large-scale codebase migrations.

The model has an output limit of 1 million tokens, up from the previous 64,000 tokens. A Google representative told The Deep View that Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. That price puts it in line with both Anthropic's Claude Sonnet 5.5 and OpenAI's GPT-6.1 Sol, which run at the same cost of $2 per million input tokens and $10 per million output tokens. However, after the introductory period, the price of Argon will double to $4 per million input tokens and $20 per million output tokens.

Google noted that the roll out will begin with paid API customers and Google AI Ultra subscribers.

Our Deeper View

Google could not have picked a more heated time to reenter the high-end frontier model race. In recent months, the company's contributions to the frontier landscape have largely focused on speed and efficiency. For instance, its early September release of Gemini 3.8 Flash cost a fraction of what its frontier competitors were charging at the time. But in recent weeks, with both Anthropic and OpenAI homing in on token efficiency as well, Google's Argon may not be able to compete on just state-of-the-art performance alone, especially as model labs continue to leapfrog each other in capability and efficiency week after week. What a model costs is quickly starting to matter as much as what it can do.

Washington puts AI safety on the honor system

Despite pleas from industry experts to slow down and regulate the pace of AI, the Trump Administration isn't budging on its approach.

Following a Tuesday meeting with AI industry figureheads, President Donald Trump has rejected the notion that the AI industry needs additional regulation, and that existing laws will be enough to prevent the risks the tech poses.

Instead, Trump and six of AI's biggest names, including Anthropic's Dario Amodei, Google's Sundar Pichai, Meta's Mark Zuckerberg, OpenAI's Greg Brockman, SpaceXAI's Elon Musk and Nvidia's Jensen Huang, signed the White House Accord on Super Intelligence, a code that Trump called "morally binding," promising to self-regulate the risks of development.

The accord states that each company will implement four "layers" of controls and audits, including:

  • Implementing robust internal controls to monitor model capability and alignment during training and deployment, with specific areas of risk including cybersecurity, biosecurity and chemical threats.
  • Giving an internal team the authority to ensure that all safeguards, controls and monitoring are operating as intended.
  • Partnership with external auditors to assess whether these safeguards and monitoring systems are operating properly.
  • And creating a designated committee on these companies' boards of directors that oversee the teams running the safeguards and controls and the internal and external teams evaluating them.
[@portabletext/react] Unknown block type "twitter", specify a component for it in the `components.types` prop

The accord states that the signatories would regularly meet to establish standards and improve their safety systems. Additionally, it states that "over time, it may make sense to codify these steps into laws and regulations."

This accord follows Trump rejecting calls for increased regulation and a slowed pace of development, calling the increased AI safety alarmism a "hoax" during an on-stage call with Huang at an event hosted by the "All-In" podcast in mid-September.

Our Deeper View

While these layers of protection are good in theory, the only thing enforcing this accord is the honor system, rather than actual law and regulation that creates consequences for the risks this tech could present. And though the leaders of these companies have largely been urging for safe development, they are all also the ones consistently pushing the frontier. Real safeguards won't result from a good faith system where companies promise they will try their best to keep the risks reined in. Additionally, these companies holding themselves accountable will only become more difficult as Anthropic and OpenAI approach public market debuts. And while regulation has often struggled to keep up with the pace of technology, in the case of tech that these companies claim could cause catastrophic risk, the best way to get substantial safeguards may be from the real consequences and outside accountability that regulation can provide.

OpenAI rolls out deluge of new tools for AI builders

Heavy AI users rejoice. OpenAI has rolled out new options aimed at accelerating the stuff you're already doing and making it safer.

On Tuesday, at the company's fourth annual DevDay event in San Francisco, OpenAI showed off its new "Dots" strategy with AI agents that are always-on and run in their own separate cloud computer. Sound familiar? Yeah, it's not unlike Meta's new Muse agent that went viral this month, but OpenAI's agent is more focused on power users. See Sabrina Ortiz's full breakdown on Dots.

While Dots may have stole the headlines, OpenAI also unveiled other new features that dedicated ChatGPT and Codex users should note:

  • GPT-6.1 Sol: OpenAI's latest model provides "near Astra intelligence" at a fifth of the price. The model sports improved performance in agentic coding, computer use and professional work, and runs $2 per million input tokens, 10 cents per million cached tokens, and $10 per million output tokens.
  • ChatGPT Space: This is a collaborative workspace where you, your teammates and AI can work together. Spaces also allows users to keep everything they share with and create in ChatGPT in one place. It will also let you collaborate in a shared workspace similar to the way you might in a shared Google Doc.
  • New $500/month plan and Ultrafast: OpenAI is creating a high-end tier that offers access to Ultrafast, which delivers up to eight times the speed in Codex using GPT-6 Astra and six times the speed in the API. However, For those on the $200 per month Pro plan, included usage of ChatGPT Work and Codex decreases from 20 times to 10 times that of the Plus allowance, and GPT-6 Pro messages in ChatGPT will be cut from 200 to 100 per week.
  • Codex in the cloud: Codex can now start, review, and continue work from the desktop app, the mobile app, or the web. And since it's not limited to working on just the local files on your machine, you can move between devices much more easily.
  • Plugin extensions: Developers can now build applications and plugins that are native to ChatGPT and Codex, and can reach more users with improved submission and discovery. And now those plugins will appear in the sidebar on the left side of your ChatGPT window, with services like Adobe and Figma being some of the first ones there at launch.
  • Private Intelligence: OpenAI has enabled zero data retention and private safety processing policies for offline, automated safety reviews. Additionally, Private Inference will be available in preview, bringing strict control to confidential computing.
  • OpenAI Marketplace: Eligible US-based enterprise customers can apply their existing OpenAI token spend towards approved third-party software. This marketplace is launching in beta with 32 partners, including Adobe, Figma, Sierra, and Harvey.

Our Deeper View

While it would be easy to look at OpenAI's DevDay as its answer to Meta's newfound momentum in the AI space, the reality is actually quite different. Meta's Muse agent is primarily focused on the consumer space and bringing agents to the masses with its focus on cute interactive avatars and wide publicity across Instagram and Facebook. OpenAI has lessened its focus on consumer AI in 2026 and leaned into enterprises, developers, and power users, and DevDay only doubled down on that strategy. Meta's agent is open to all and has generous token allotments based on Meta's strategy of eventually taking a percentage when you use Muse to sell stuff and make money. OpenAI is only offering its new agent to its highest-tier Pro and Business Premium customers. So OpenAI's focus is on its most-invested AI builders and organizations. And that strategy is supported by all of the other features it released around Dots on Tuesday, from Private Intelligence to GPT-6.1 Sol to its new Ultrafast feature. All of these things are aimed much more at competing with Anthropic for power users than competing with Meta for consumers.

Claude Sonnet 5.5 puts AI efficiency in focus again

Anthropic is releasing a more cost-effective version of its flagship version 5.5 model.

On Monday, the AI lab released Claude Sonnet 5.5, the latest addition to its lineup, a "lower-cost complement" to its recently-released Opus 5.5 model featuring faster performance compared to previous iterations.

The company said that Sonnet 5.5 runs 30% faster than Claude Sonnet 5. Notably, however, Anthropic says that Sonnet 5.5 needs far fewer tokens to do the same amount of work as previous models, costing up to 30% less for most tasks.

On the surface, Sonnet 5.5 costs the same as its previous generation, sitting at $2 per million input tokens and $10 per million output tokens, as well as 20 cents per million cache reads. For reference, OpenAI's GPT-6 Sol runs at the same cost of $2 per million input tokens and $10 per million output tokens, and Anthropic's Opus 5.5 costs double, at $4 per million input tokens and $20 per million output tokens. Sol and Opus are considered the same class of model, while Sonnet is in the same class as OpenAI's mid-tier Terra model.

Sonnet performs strongest in everyday tasks, such as fixing bugs and creating documents, slides and spreadsheets. Compared to the previous generation of Sonnet, Anthropic said Sonnet 5.5 offers a number of improvements:

  • The model sweeps Sonnet 5 across the board in all benchmarks, including agentic coding, knowledge work, computer use and multidisciplinary reasoning. Though it sits just below Opus 5.5's scores in almost every benchmark, Sonnet 5.5 scored higher than both Sonnet 5 and Opus 5.5 on the Terminal-Bench 4.0 benchmark for agentic coding.
  • Anthropic says the model writes more clearly than previous iterations, making it a better partner for collaborative work, and is faster on less complex tasks.
  • Sonnet 5.5 improves or meets Sonnet 5 on most measures for alignment in Anthropic's automated behavioral audit, the company said.

Sonnet 5.5's cyber capabilities are comparable to those of Opus 5, making it the first of the Sonnet series to have the same cyber safeguards as its most capable models. In addition to cybersecurity safeguards, Sonnet 5.5 is built with guardrails against the harmful queries for biology.

Additionally, Sonnet 5.5 is built to prevent model distillation attacks, launching with "safety classifiers" that prevent reasoning extraction and "preserved thinking" that prevents the model's reasoning from being "decoupled from the account that created it."

Claude Sonnet 5.5 is available on all platforms, and with no data retention. Anthropic also reported that it will be releasing a 5.5 version of Claude Haiku, the lightest and lowest-cost model series it offers, in the coming weeks.

Our Deeper View

OpenAI and Anthropic have read the room on enterprise cost cutting. They are now competing heavily on price, pushing the boundaries on just how cheap they can serve up intelligence while leapfrogging each other on performance. The problem that the labs face, however, is that they are relying on conventional model architecture to try and drive down the efficiency curve. Meanwhile, neolabs may present a threat to that strategy, as companies like Jev-maker Typesafe and Pathway gain traction as they look towards a future beyond traditional LLMs that can provide comparable performance at a fraction of the cost.

Why an AI-fueled jobs crisis is not inevitable

One of the foremost AI labs predicted three distinct scenarios for AI's future economic impact, and two involve large swathes of the workforce losing out. But what is the reality of that future?

Earlier this month, Anthropic's economics team released research painting a picture of AI's potential modest, substantial, and extreme impact on the economy by 2030. While all three involve increases in the GDP, ranging from a 1.6% increase to a 32.4% increase, the catch is job displacement. The bigger the impact AI has on the economy, the larger the percentage of knowledge workers who are stripped of their jobs, with a large portion unable to find new work in the most extreme scenarios, according to this research.

Additionally, Anthropic predicts that, as the impact of AI grows more severe, despite the growth in the GDP, the distribution of wealth is uneven, with more money made going back to capital than it does to workers' wages. As it stands, 60 cents of every dollar made goes back to the worker, and 40 cents to capital. AI could flip these figures in the most extreme scenarios, stagnating wages and worsening unemployment. That's the picture Anthropic paints in this report.

Here are the three scenarios:

  • Modest scenario: Anthropic says that AI has roughly the same impact on the economy as the internet, driving significant gains, though within historical norms. The GDP reaches $34.1 trillion, up 1.6%, by 2030, and though 0.3% of knowledge workers are displaced, all of those workers are able to be reallocated into new roles.
  • Substantial scenario: AI has roughly the same impact as the railroad, causing an 8.3% rise in the GDP to $36.3 trillion. Around 2.5% of knowledge workers are displaced, 1.8% of which are able to find new roles by 2030.
  • Extreme scenario: GDP sees a 32.4% increase by 2030 to $44.4 trillion as a result of unprecedented growth, likely led by the development of recursive self-improvement. 13.5% of knowledge workers are displaced, and only 5.2% are able to find new work, leaving 8.3% unemployed.

However, there may be a few hitches in the more extreme scenarios that Anthropic laid out, Julius Probst, senior economist at recruitment marketing firm Appcast, told The Deep View. For one, these scenarios assume that all of the GDP value that AI is generating will be consumed. But in the most extreme scenario, if we are heading towards a labor market facing wide-scale unemployment and a great deal of concentrated wealth, "that is probably a bad assumption," said Probst. Many consumers would not be able to afford to buy things, meaning that the GDP won't actually surge.

"Wealth inequality will soar, but these people will not buy additional cars or additional houses," said Probst. "There's only so many more houses a billionaire can have."

The other hitch is that the model largely lumps together all knowledge workers into one bucket. The reality is that, while knowledge workers will be broadly impacted, many companies are still seeking senior, more skilled workers whose roles can be complemented by AI, said Probst. But even that won't be sustainable for long, he said, as many of those senior workers will retire and companies will realize they need to invest in junior talent again. "I don't think this situation can persist for another three to five years. At some point in time, companies will realize we need to hire junior people again."

And given that a large portion of the labor force is involved in physical work that can't be supplanted by AI, of the three scenarios Anthropic painted, the most likely is the modest one, he said. "AI growth is really showing up in two sectors only, and that is the tech sector and professional business services."

Our Deeper View

There has long been a narrative being pushed by AI stakeholders that the tech is inevitable, and that every worker needs to get on board or be left behind. Anthropic's economic scenarios, especially the most extreme, support that narrative. If you are not able to keep up with AI skills, you could end up one of the 8.3% of knowledge workers that Anthropic predicts will be unable to cross over into a new job by 2030. But there are important caveats to remember in this forecast. The first is that the companies parroting the idea that AI is going to upend our economy and every worker needs to hop on board has a clear incentive to get as many people to adopt its technology as possible. And the second is that the future is not decided. Anthropic, to its credit, notes this in its research. However, in order to prevent the worst outcomes of AI, we need to prepare our economy and workforce for AI now, and stop believing the narrative that AI transformation, and the havoc it can wreak, are inevitable.

OpenAI's mental health test exposes AI's blind spots

People are turning to AI for emotional support more than ever. But these chatbots' ability to provide a shoulder to cry on can vary greatly.

Because of this, OpenAI decided to measure it: On Wednesday, the AI lab unveiled MentalHealthBench, a new open benchmark dedicated to evaluating model capabilities in domains such as safety, seeking user context, preserving user agency, and providing actionable guidance.

To create this benchmark, OpenAI began by developing synthetic conversations that reflected real-world AI use patterns that span multiple topics and run the gamut of severity, ranging from non-acute situations that involve emotional themes, to "high-accuity" situations that indicate more serious concerns or distress, to emergency situations that require immediate support. Then, the company worked with a cohort of 80 mental health professionals across 22 countries, 19 languages and 20 subspecialties to evaluate the responses to the synthetic message.

The benchmark breaks down model performance by conversation severity, as well as a range of 10 dimensions defined by the mental health experts, including whether the model asks the right questions, provides appropriate, clinically accurate guidance, helps the user see reality, avoids harm and recognizes serious risk.

In developing the benchmark, OpenAI also put a number of its own and other models to the test:

  • Astra took the overall best score on the evaluations, scoring 57.3%, with GPT-6 Sol and Luna trailing behind at 53.9% and 50.2% respectively. Claude Opus 5.5 stood third in the line-up at 52.4%.
  • However, performance differs across dimensions of the benchmark. Though Astra still largely outranks other models, all of the models tested tended to perform better in certain areas, such as clinical accuracy, empathy and reality testing, while scoring lower in dimensions such as gathering context and supporting user agency.
  • OpenAI said that MentalHealthBench also points to several opportunities to improve ChatGPT, including asking useful follow-up questions and responding with the right level of urgency, and that it will use this information to guide improvements and track the model's progress.

"This is not a leaderboard," Dr. Declan Grabb, mental health safety research lead at OpenAI, told The Deep View. "What I hope that this benchmark provides is a nuanced view into model behavior, so that people really understand the more complex dynamics of their models."

MentalHealthBench adds to a number of mental wellness-related initiatives that OpenAI has endeavored, including research to combat model sycophancy and improving ChatGPT's responses to sensitive conversations, as well as joining forces with advocacy group Common Sense Media to support the Parents and Kids Safe AI Act. OpenAI said that this is just a piece of its research into mental health benchmarking and alignment in this area, not an end state.

"ChatGPT is not a therapist, and is not here to replace a clinician," said Grabb. "That being said, when I talk to mental health clinicians across the globe, the most responsible and safe thing to do is if people are coming to AI to ask these questions, we absolutely need to have an expert opinion on how you should navigate them."

Our Deeper View

Mental healthcare is a critical area for these models to get right. While OpenAI said that speaking to a chatbot should not supplant actual therapy, the reality is that many people have and will turn to a chatbot for support, seeking both a judgement-free and cost-free alternative to clinical support. The company faces lawsuits involving the deaths of Adam Raine and Joshua Enneking, whose families allege that ChatGPT contributed to their suicides. A benchmark can help identify weaknesses, but a higher score alone does not establish that a model is safe in a real conversation. The gaps in gathering context and supporting user agency are particularly important: an empathetic response is not necessarily an appropriate one. The next test for OpenAI is how it turns those findings into changes that make its models safer for the people relying on them.

CORRECTION: This story has been updated to reflect accurate figures for OpenAI and Anthropic's model scores on this benchmark.

Why OpenAI is giving ChatGPT Voice a bigger job

OpenAI has created an assistant to help you get work done anywhere.

On Wednesday, the lab released ChatGPT Voice "on-the-go," bringing its voice models to the web and mobile experiences of its agentic tools. The product builds on GPT-Live, its upgraded voice model, by bringing ChatGPT Voice to its Work and Codex experiences.

OpenAI said that this release allows users to get work done whenever they speak to the model. For instance:

  • For those with Pro and Plus plans, users can use Voice in ChatGPT Work on mobile to create docs, decks and spreadsheets, work with connected plugins, draft emails, create sites, and talk to Slack.
  • Free and Go users can also use Voice in the chatbot itself to interact with plugins that they've connected to the text model, including tools and connected apps.
  • The voice model is also connected to the text model, meaning that text conversations flow as ChatGPT answers vocally, allowing you to switch seamlessly back and forth between voice and text conversations. The company noted that ChatGPT Voice will become more powerful as ChatGPT itself grows more capable.

This update comes a day after the company unveiled the latest versions of its flagship product, GPT-6 Sol and Luna, both coming in cheaper than the models of its biggest competitor, Anthropic.

This expansion to ChatGPT Voice's agentic capabilities also comes as a new competitor emerges: Meta's Muse personal agent is rapidly rising in popularity. According to The Information, Muse has surpassed 500,000 users in the first week, including 250,000 daily active users. Muse is outpacing ChatGPT's early mobile launch.

Our Deeper View

OpenAI is feeling pressure on all sides. The company faces pressure from Anthropic to compete both on frontier model development and on enterprise adoption, while Meta has popped up as substantial competition on the consumer side. This update to ChatGPT Voice may actually address both of those fronts. While bringing ChatGPT Voice to platforms like Codex and ChatGPT Work certainly helps with enterprise adoption, bringing the experience to mobile could provide those same agentic capabilities to a far wider audience. That's where Meta's Muse agent is making such rapid progress by making agents easier for the average user to access. Essentially, this substantial update to ChatGPT Voice is trying to bridge the gap between enterprise and consumer adoption.

Opus 5.5 makes frontier AI cheaper and safer

Days after calling for a slowdown on frontier model development, Anthropic is back with another model release. The catch is that the company claims this one is safer.

On Tuesday, the company unveiled Opus 5.5, the latest of its flagship Claude models and what the company calls its "strongest-performing model" on behavioral alignment yet. Additionally, the model features safeguards developed specifically for its most capable models.

One of those safeguards is falling back on previous generations of Opus. For instance, in cybersecurity use cases, most tasks will be rerouted to Opus 4.8, and requests flagged for biology classifiers will be routed to Opus 5. Only vetted organizations through Anthropic's Life Sciences and Cyber verification programs will be able to use Opus 5.5 for these tasks.

[@portabletext/react] Unknown block type "twitter", specify a component for it in the `components.types` prop

Alignment and safety aside, Anthropic laid out a few improvements featured Opus 5.5, including:

  • Improved performance, representing a major step up compared to Opus 5 when it comes to complex work, beating previous generations of Opus and Fable 5.1 in benchmarks for agentic coding, knowledge work, computer use and visual chart recognition.
  • Improved natural communication, offering clearer writing that's easier to follow, putting the most important information at the top of the outputs.
  • Better speed, generating outputs more than 30% faster than Opus 5.

And of course, Anthropic addressed the elephant in the room: costs. Opus 5.5 now requires less compute to serve than its predecessor with pricing to match. Opus 5.5 costs 40% less than Opus 5 on typical workloads. Input tokens cost $4 per million and output tokens cost $20 per million, representing a 20% decrease from Opus 5, though still more than OpenAI's most recent comparable release, GPT-6 Sol, which costs $2 per million input and $10 per million output.

Additionally, cache reads, which the company says make up the majority of agentic and coding work, sit at $0.20 per million tokens, 60% less than Opus 5. Anthropic is also increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans. The company also said 5.5 versions of Sonnet and Haiku will be made available in the coming weeks, and those will likely cost even less.

Our Deeper View

There's one thing that Anthropic CEO Dario Amodei noted in his latest essay urging for pacing the frontier that has stuck with me since it was published. He emphasized that "progress will still seem fast." Releasing yet another powerful iteration of its models is seemingly an example of this. However, what's clear with this release is that, while Anthropic is eager to keep up with the stiff competition, not every task is right for frontier AI, and frontier AI is not ready for every task. It's why the fall back plan for safeguards isn't simply an outright refusal to do certain tasks, but rerouting those tasks to less capable versions of its models. And though Amodei said in his essay that we "must make wise use of the time we gain" that we get as a result of pacing, the question we have to ask is how much time do measures like this buy us before these models are being leveraged for riskier tasks?

OpenAI’s cheaper GPT-6 models change the math

OpenAI is releasing a more budget-friendly version of its most powerful model.

On Tuesday, the company unveiled GPT-6 Sol and GPT-6 Luna, the latest additions to its lineup following the release of Astra, which has quickly become one of the world's top performing models but is also neck-and-neck with Anthropic's Fable 5.1 as one of the most expensive models. OpenAI said that GPT-6 Sol and Luna were trained with similar methods to Astra, touting advancements in factuality, coding, computer use, and alignment.

The bigger highlight, however, is the cost: These models are priced around 50% cheaper per million tokens than previous iterations of Sol and Luna.

  • GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, compared to 5.6 Sol's cost of $4 per million input tokens and $20 per million output tokens.
  • Meanwhile, GPT-6 Luna runs at 10 cents per million input tokens and 50 cents per million output tokens, compared to its previous generation's costs of 20 cents per million input tokens and $1.20 per million output tokens.

Though OpenAI still says Astra is its "best model across the board," GPT-6 Sol outperforms previous OpenAI models, as well as Anthropic's Claude Opus 5 on a number of benchmarks, including AutomationBench, which tests business workflows across apps, Agents' Last Exam for complex agentic workflows, and DeepSWE v1.1 for complex software-engineering tasks in real codebases.

Additionally, OpenAI said GPT-6 Sol and Luna both feature Astra's improved communication style, featuring more clarity, less jargon, slightly shorter answers and fewer "low-value details." Also as a result of Astra, these models feature improved alignment compared to previous iterations, showing significantly lower rates of circumventing warning messages, coding deception, and unauthorized agent interactions.

These models are currently available in ChatGPT Work and Codex for Plus, Pro,

Business, and Enterprise users. Free and Go users can access GPT-6 Luna in the desktop app. The models are not yet available in traditional chat.

Our Deeper View

It's clear that OpenAI is reading the tea leaves on cost. For many everyday tasks, enterprises don't want to pay for the most expensive models, no matter how powerful and capable they are. This is especially true as agentic deployments start to consume a greater amount of tokens. OpenAI is showing it is capable of bending with the trend, not only by retrofitting its state-of-the-art model for efficiency, but by cutting the cost of that model from its previous generations. Attracting users with low prices and solid performance may be OpenAI's best shot at keeping its lead and fending off innovations that threaten its bet on conventional scaling laws, such as the recent innovations from Jev, AlohaJet and Pathway that The Deep View has reported on.

or get it straight to your inbox, for free!

Get our free, daily newsletter that makes you smarter about AI. Read by 750,000+ from Google, Meta, Microsoft, a16z and more.