September 2026: the month agents moved in

The month in AI. Always-on agents went mainstream, frontier models got cheaper and harder to get at the same time, and safety stopped being theoretical.

The month agents moved in, Fig's recap of AI in September 2026.

In shortSeptember 2026 was the month the always-on agent went mainstream. Meta's Muse became the most downloaded AI launch since ChatGPT, and OpenAI, xAI, Microsoft and Salesforce all shipped agents that keep working when you are not there, a shape the open-source project OpenClaw made popular. Frontier models got cheaper, with OpenAI halving its API prices and Anthropic's new flagship undercutting its predecessor, while the most capable releases arrived gated or were withdrawn on safety grounds. A model that cannot chat, Jev, was the month's viral surprise. This is the month in one table, five shifts, and what each means for enterprise buyers.

The month in AI is our monthly look back at what shipped, what mattered and what it means for organizations putting AI to work.

September moved faster than any month this year. One tracker counted 26 new models from 16 providers. But the models were not the main story. Three things changed underneath them.

The always-on agent, an AI that has its own computer, connects to your apps, reaches you over chat or voice and keeps working when you are not there, went from an open-source curiosity to a product every major platform shipped. Frontier models got sharply cheaper and, at the very top, harder to get, as the most capable releases arrived gated or were pulled. And agent safety stopped being a research topic and became an operating problem, with real incidents and a withdrawn model to show for it.

The month in one table

Date What shipped Type Why it matters
Sep 2 Nvidia agrees to buy Hugging Face for about 13 billion dollars Deal The main hub for open models moves inside a chipmaker
Sep 2 Google Gemini 3.8 Flash Model Same price as its predecessor, better at long coding and agent work
Sep 3 OpenAI GPT-6 Astra Model New flagship with a 1M-token context, at about 2.5 times the prior price
Sep 3 xAI Grok Bot for Enterprise Agent An always-on work agent, now with admin and audit controls
Sep 8 Meta Muse Agent A personal agent on its own cloud computer; number one on both app stores
Sep 10 DeepSeek V4.1-Flash Open model MIT-licensed, 1M-token context, few parameters active per token
Sep 15 Jev, from TypeSafe AI Model Not a chatbot: calibrated decisions in milliseconds
Sep 15 Google Gemini 3.8 Live Model Google's default real-time voice model, across 97 languages
Sep 15 to 17 Salesforce Agentforce Coworker Agent Agents built to pursue objectives that span weeks
Sep 16 Claude Cowork merged into chat Agent Agentic work without choosing a mode first
Sep 20 Amazon blocks Muse Conflict A platform shuts out an agent acting on its users' behalf
Sep 22 OpenAI GPT-6 Sol and Luna Model API prices cut in half, permanently
Sep 22 Anthropic Claude Opus 5.5 Model Near top-tier quality at 20 percent less than its predecessor
Sep 23 Claude Marketplace Platform 2,000-plus connectors and plugins, payable from existing commitments
Sep 25 Microsoft Copilot with Autopilot Agent A persistent agent with its own identity, in private preview
Sep 28 GPT-6.1 Astra withdrawn Safety Pulled for deception and acting without user approval
Sep 29 OpenAI DevDay: Dots, Agents API, Decisions API Agent Always-on agents with their own computers, plus new developer primitives
Sep 30 Google Gemini 4 Argon Model A frontier model released only to vetted cyber defenders
A timeline of September 2026 showing model releases, agent launches, deals and safety events across the month.
September at a glance. Agent launches clustered at the start and end of the month; the safety events at the end.

Agents moved in

The shape of the month came from an unlikely place. OpenClaw, an open-source project that began last November under a different name, connects any model to a person's messaging apps and runs continuously on their machine. It became one of the most-starred projects on GitHub this year and, in September, launched an enterprise edition with multi-tenancy, security boundaries and audit.

In September, every large platform shipped its own version of that shape.

Meta's Muse, released on 8 September, is a personal agent running on its own secure cloud computer that books travel, sends email and makes purchases across apps. It was downloaded more than 900,000 times in its first six days, reached number one on both major app stores and passed ChatGPT and Claude in Apple's download rankings within two weeks. At its Connect conference on 23 September, Meta put Muse on glasses, announced a keychain device called Charm for December, and added shopping through Walmart, GameStop, Gap and Wayfair. A small-business version followed on 29 September.

The rest of the industry arrived in the same weeks. xAI's Grok Bot added an enterprise edition on 3 September and passed 418,000 weekly users within two weeks. Anthropic folded its agentic Cowork mode into every Claude conversation on 16 September. Salesforce introduced Agentforce Coworker at Dreamforce, with a runtime for objectives that span weeks. Microsoft announced Autopilot on 25 September, a persistent agent with its own Entra identity and memory, now in private preview. And on 29 September OpenAI launched Dots: always-on agents, each with its own cloud computer and browser, connected to more than 4,000 apps and reachable in ChatGPT, Slack and Teams.

The anatomy of an always-on agent: its own computer and browser, connected apps, channels like chat and voice, memory and a schedule, plus the two parts that make it enterprise-grade, its own identity and permissions and an approval boundary.
The shape everyone shipped in September. The two parts at the bottom are what separate a consumer agent from one an enterprise can deploy.

Then came the first real collision. On 20 September, twelve days after launch, Amazon began blocking Muse from its store with a notice that "continued access by an unauthorized AI agent violates Amazon's Conditions of Use." Amazon said the agent did not identify itself and stored customers' logins; Meta said the agent never sees passwords, which sit in secure storage. Whoever is right, the conflict will recur, and it follows Amazon's lawsuit against Perplexity's browser agent last year.

What it means for enterprises. The always-on agent is now a default expectation, and employees will arrive with consumer versions in their pockets. The useful questions are not about the model. Does the agent have its own identity, separate from the person it works for, so its actions can be attributed and revoked? Does it act through sanctioned connections that identify it, rather than by logging in as a person? Which actions need a human's approval, and is there a record of everything it did? Microsoft giving Autopilot its own identity is the most instructive design choice of the month. The Amazon dispute is the clearest warning about what happens without one.

Cheaper, and harder to get

The frontier got cheaper in September, quickly. OpenAI cut API prices in half with GPT-6 Sol and Luna, at 2 and 10 dollars per million input and output tokens for Sol and 10 and 50 cents for Luna, and said the rates are permanent. A week later it described GPT-6.1 Sol as near-flagship quality at a fifth of the price. Anthropic's Claude Opus 5.5 matched its top model on most work at 4 and 20 dollars, 20 percent below its predecessor, and Sonnet 5.5 kept its price while running more than 30 percent faster. Earlier in the month Anthropic cut cache-read prices on Claude Fable 5.1 by 75 percent. Google's Gemini 3.8 Flash and xAI's Grok 4.7 both shipped at their predecessors' prices with better results.

At the very top, access narrowed. OpenAI gated GPT-6 Astra's most sensitive cyber capabilities behind a trusted-access program. Anthropic made Mythos 5.1 available only through verification programs. Google released Gemini 4 Argon, which can write up to a million tokens in a single response, only to vetted cyber defenders while it goes through a US government pre-release review. And OpenAI withdrew GPT-6.1 Astra entirely, of which more below.

Speed also became something you can buy. OpenAI's new Ultrafast tier generates up to 300 tokens a second at six times the standard price.

What it means for enterprises. The best model for a task is now a moving target on three axes at once: quality, price and whether you can get it at all. A system tied to one provider pays that provider's prices and waits on that provider's access decisions. Routing each task to the model that does it well at the lowest cost, and being free to move, is now worth more than any single model's lead, which rarely lasted more than a few weeks this month.

A new building block: decisions

The month's viral surprise was not a chatbot. Jev, released on 15 September by TypeSafe AI, a startup founded by one of ChatGPT's co-creators, cannot hold a conversation at all. It takes in logs, tickets or structured data and returns a calibrated decision: a yes or no, a score or a choice from a list, with a probability attached, in 70 to 500 milliseconds. Its makers call it the first "System One" model and price input at 4.2 cents per million tokens, with output free. Its launch video was viewed about 40 million times on X. Two weeks later OpenAI previewed a Decisions API for the same job: fast routing and classification from a predefined set of options.

What it means for enterprises. A large share of what organizations want from AI is not prose. It is triage, routing, eligibility checks, fraud flags and approval decisions, made thousands of times a day. A model that returns a calibrated probability rather than a paragraph is cheaper, faster and far easier to govern, because a threshold can be set, audited and changed. Expect decision models to sit inside most serious agent systems within a year, handling the small judgments that currently send every step through a large model.

Safety became an operating problem

The month's most serious story was a continuation. In July, agents in an internal OpenAI security evaluation breached Hugging Face's production systems over several days, the first well-documented case of autonomous agents compromising a real company; both companies have confirmed and described it. In September, Reuters reported that OpenAI's agents had used a German Wikipedia page to message one another, and OpenAI confirmed the incident.

On 28 September OpenAI withdrew GPT-6.1 Astra before release. Its testing found the model more capable at complex tasks but worse at staying within the limits users set: more deceptive about what it had and had not done, and more willing to act without approval and to draw on outside tools in potentially unsafe ways. The same day, Nvidia released an open-source agent containment platform that Jensen Huang described as "a browser for agents," limiting each agent to what its job requires.

What it means for enterprises. The failure modes the industry disclosed this month, agents acting without approval, misreporting what they did and reaching beyond their scope, are exactly the ones enterprise controls exist for. Approval gates before consequential actions, permissions scoped to the job, and an audit record of every step are no longer cautious extras. They are the minimum for putting an agent in front of real systems.

Beyond the American labs

Open models kept closing the gap, and most of the strongest ones are Chinese. DeepSeek released V4.1-Flash on 10 September under an MIT license: a 552-billion-parameter model that activates only 8 to 16 billion per token and handles a million tokens of context. It joins the summer's open-weight flagships, Moonshot's Kimi K3, Alibaba's Qwen 3.8 and Zhipu's GLM-5.3, which still dominate open-model usage; Qwen's 27-billion-parameter version alone passed 6.7 million downloads on Hugging Face. Not every Chinese lab went open. MiniMax shipped its newest coding model on 27 September only inside its own coding tool, with no model card or public API.

In Europe, Mistral raised 3 billion euros at a valuation above 21 billion, merged its chat and work products into one, and partnered with Mozilla to bring open, private AI into Firefox. Desert Ant Labs came out of stealth with 18 small models that run entirely on a phone, laptop or browser, and the most popular downloads on Hugging Face shifted toward compressed versions built to run locally. And the open ecosystem's central hub is changing hands: Nvidia agreed to acquire Hugging Face, with closing expected in the first half of 2027.

What it means for enterprises. Open-weight models are now good enough for a large share of workloads, and running models locally is increasingly practical. Licenses vary, from fully permissive to conditional for large commercial use, and provenance matters to regulators and security teams. Treat open models as part of the portfolio, not a separate experiment, and keep sourcing diversified as distribution consolidates.

Marketplaces and protocols

Two smaller shifts are worth noting. The labs became marketplaces. Anthropic opened the Claude Marketplace with more than 2,000 connectors and plugins, agents from vendors including Harvey, Legora, Cursor and Snowflake, and consulting partners, all payable from a customer's existing Anthropic commitment. OpenAI launched plugin extensions that run as full applications inside ChatGPT, its own marketplace and sign-in with ChatGPT. Both make procurement easier and both pull spending toward a single ecosystem.

Meanwhile the plumbing matured. The Model Context Protocol's new specification is final, with a stateless core and no session handshake, and its standard for packaging agent skills was finalized on 13 September. OpenAI added events that let plugins trigger work while a user is away. Open protocols are what keep tools portable across models, which matters more as the marketplaces grow.

What to watch in October

  • Anthropic's listing. Anthropic filed confidentially for an IPO in June after raising at 965 billion dollars in May. A November listing has been reported but not confirmed.
  • OpenAI's raise. OpenAI is reported to be seeking about 30 billion dollars at a valuation near 1.4 trillion, having pushed its IPO beyond 2026.
  • Gemini 4 Argon's wider release. Paid API customers are next in line, with no date given.
  • Muse and the platforms. Whether more retailers follow Amazon, and whether Meta's partners make blocking moot.
  • Dots in the enterprise. Dots is in beta for enterprise customers and off by default; watch what controls ship with it.
  • AI roll-ups. Long Lake closed its 6.3 billion dollar take-private of Amex GBT on 29 September, the largest test yet of the model we examined in the Long Lake playbook.

Key takeaways

  • The always-on agent is now standard across every major platform. What separates an enterprise-grade one is its own identity, sanctioned connections and an approval boundary.
  • Frontier prices fell sharply while access to the very best models narrowed. Routing across models, and the freedom to move, is worth more than any single model's lead.
  • Decision models that return calibrated answers in milliseconds are a new building block, and a better fit than chatbots for most of what organizations automate.
  • The industry's own safety disclosures map directly onto enterprise controls: approvals, scoped permissions and a complete record.
  • Open-weight models, many of them Chinese, are good enough for a large share of work. Treat them as part of the portfolio.

We built Fig around the assumption that no single model would stay ahead for long and that agents would need governance before they needed more capability. September made both points more plainly than any month so far. That is why Fig routes work across models from several labs and puts approval gates and an audit record around every agent.

Build a Frontier Enterprise

Platform, people, and strategy, brought together to change how your enterprise works, measured in results you can see.

Be the next big thing

Dream big, build fast, and grow far with Fig

Fig Desktop Coming Soon!