OpenAI published a customer story on September 11, 2026, titled "Cognition helps Devin test its own work with GPT-6 Astra," describing how the AI coding startup is using OpenAI's newest model to make its autonomous engineer, Devin, test its own code and show that the work actually functions. The rollout covers Devin's cloud agent, CLI, and desktop products, and Cognition frames it as a way to shrink the amount of code engineers have to manually check before it ships.
The news is small in scope but well timed. It arrives three days after Cognition closed a $2 billion Series E that valued the company at $48 billion, roughly double its price tag from four months earlier. Put together, the two stories say something specific: Cognition is betting that better proof of work, not just faster code generation, is what keeps enterprise customers paying at this scale.
What OpenAI and Cognition Actually Announced
According to OpenAI's post, Cognition is applying GPT-6 Astra across its product lineup to strengthen Devin's testing step, the part of the workflow where the agent has to demonstrate that code it wrote actually works. Walden Yan, Cognition's co-founder, told OpenAI the model's main contribution is its ability to test and prove that its work actually functions the way you expect. In one example OpenAI cites, Devin uses Astra to test an iPhone game called Otter Run and returns a recording of the game running in a simulator, alongside a report listing which checks passed and which areas remain untested.
Cognition says the same setup is helping close the loop on customer bug reports faster. When a customer sends a screenshot of a broken feature, the team passes it to Devin, which uses Astra to fix the issue and return a screenshot showing the result. None of this changes what Devin writes. It changes what Devin has to show for it, which is the part Cognition says has become the bottleneck as its engineers generate more code than they can individually review.
The Benchmark Numbers Are Cognition's, Not an Outside Lab's
On its own blog and on X, Cognition published specific figures: on its FrontierCode 1.1 benchmark, GPT-6 Astra scores within 0.4 points of Anthropic's Claude Fable 5 while costing 64 percent less, and it sets a new high on Cognition's internal testing benchmark for generating comprehensive tests and clearer reports. A second Astra launch post from OpenAI carries a similar claim from Silas Alberti, Cognition's SVP of research, who said Astra's writing and codebase understanding meant videos are noticeably easier to follow, and reports are clearer and more concise.
Those numbers are worth taking at face value only as far as they go. FrontierCode 1.1 is Cognition's own benchmark, graded by Cognition, on tasks Cognition selected. Neither figure has been replicated by an independent evaluator, and OpenAI's post doesn't include one either. That's a familiar pattern: a foundation model launches, and the partner company that gets early access publishes numbers that look great and can't easily be checked from outside. It doesn't make the claim false. It does mean it's a company's marketing number until someone outside Cognition runs the comparison.
It's also worth putting Astra's broader reception next to Cognition's praise of it. Independent testers who got early access to Astra after OpenAI launched the model on September 3 found it excelled at spatial, mechanical, and agentic tasks, the same category testing code falls into, while several rated its general writing below its predecessor. That split lines up with what Cognition is describing: a model that's genuinely stronger at producing and verifying structured technical output, deployed for exactly that kind of task.
Why It Matters
I've watched a lot of "agent tests itself" announcements over the years, and most of them are really about trust, not capability. Devin has always been able to write code fast. The harder problem, the one that actually determines whether an enterprise lets an autonomous agent merge to production, is whether a human can trust the agent's own account of whether that code works. A video of a simulator run and a pass/fail report are cheap for Cognition to generate and genuinely useful for an engineer deciding how closely to check something. That's a real product improvement, even if the specific benchmark numbers behind it are self-reported.
The timing also tells you where Cognition's incentives sit right now. The company's run-rate revenue reportedly grew from $492 million in May to nearly $900 million by its September funding round, and testing depth is exactly the kind of feature that keeps large customers, the ones running Devin against real codebases at big banks, comfortable scaling usage rather than pulling back. A better testing story is cheaper to ship than it is to walk away from once you've promised investors this kind of growth curve.
What to Watch
The next real signal here isn't another Cognition blog post, it's whether an independent group, something like Artificial Analysis or a third-party SWE-bench style eval, runs Astra-powered Devin against a testing benchmark Cognition didn't design. Also worth tracking: whether Cognition's own safety researchers raise the same concerns that OpenAI's Astra launch drew elsewhere, since a model that's trusted to verify its own work carries more downstream risk if its judgment about what counts as "passing" is wrong. If Devin's testing reports start showing up in customer case studies with third-party validation attached, that's the confirmation. If they stay purely self-reported, treat the efficiency numbers as directional, not verified.
FAQ
How does Cognition use Devin?
Cognition uses Devin both as a product it sells to customers and, by its own account, internally: in a May 2026 interview with TechCrunch, CEO Scott Wu said 89 percent of code committed by Cognition's own engineers was committed by Devin, with the remainder coming from Windsurf, the AI coding IDE Cognition acquired. The September 11 update focuses on the testing layer, using GPT-6 Astra to have Devin verify and document its own output across Devin's cloud agent, CLI, and desktop products.
Is Devin better than ChatGPT?
That comparison doesn't quite hold up, because the two aren't built for the same job. ChatGPT is OpenAI's general-purpose assistant, while Devin is Cognition's autonomous coding agent, and as of this announcement it runs partly on OpenAI's own GPT-6 Astra model for testing. Cognition positions Devin as competing with other autonomous and semi-autonomous coding tools, including GitHub Copilot, Cursor, Anthropic's Claude Code, and OpenAI's own Codex, rather than with ChatGPT directly.
Who is behind Devin AI?
Devin is made by Cognition AI, Inc. (also known as Cognition Labs), a San Francisco company founded in November 2023 by Scott Wu, Steven Hao, and Walden Yan. Cognition has since acquired Windsurf, the AI coding IDE, and closed a Series E in September 2026 that valued the company at $48 billion.
Who is the CEO of Cognition AI?
Scott Wu is Cognition's co-founder and CEO. Before starting Cognition, he founded Lunchclub, an AI-powered networking platform, and competed at the International Olympiad in Informatics, where he won three gold medals.
Is Devin AI free?
Yes, in a limited form. Devin's own pricing page lists a Free plan that includes limited Devin usage, Devin Review, and DeepWiki. Running Devin's full autonomous cloud agent at meaningful scale requires a paid plan, starting at $20 a month for Pro, with Teams at $80 a month and Max at $200 a month, plus custom Enterprise pricing.
Key Takeaways
- OpenAI confirmed on September 11, 2026 that Cognition is using GPT-6 Astra to strengthen how Devin tests and documents its own code, across Devin's cloud agent, CLI, and desktop products.
- Cognition says Astra can generate video evidence, such as a simulator recording of an iPhone game, alongside reports of what passed and what remains untested.
- Cognition's efficiency claims, including a FrontierCode 1.1 score within 0.4 points of Claude Fable 5 at 64 percent lower cost, come from the company's own benchmark and haven't been independently verified.
- The announcement follows Cognition's September 8 Series E, which valued the company at $48 billion and put its run-rate revenue near $900 million, up from $492 million in May.
FAQ
How does Cognition use Devin?
Cognition sells Devin as an autonomous coding agent to enterprise customers and, per CEO Scott Wu's account to TechCrunch in May 2026, uses it internally as well, with 89 percent of Cognition's own committed code coming from Devin. This latest update adds GPT-6 Astra to Devin's testing step, so it can verify and document its own code across the cloud agent, CLI, and desktop products.
Is Devin better than ChatGPT?
They're not really comparable products. ChatGPT is OpenAI's general-purpose assistant, while Devin is Cognition's autonomous coding agent, and it now runs partly on OpenAI's GPT-6 Astra for testing. Cognition frames Devin as competing with tools like GitHub Copilot, Cursor, Claude Code, and Codex, not with ChatGPT.
Who is behind Devin AI?
Devin comes from Cognition AI, Inc., founded in November 2023 by Scott Wu, Steven Hao, and Walden Yan, and headquartered in San Francisco. Cognition also owns Windsurf, an AI coding IDE it acquired in 2025.
Who is the CEO of Cognition AI?
Scott Wu, one of Cognition's three co-founders, serves as CEO. He previously founded Lunchclub and is a three-time gold medalist at the International Olympiad in Informatics.