Skip to main content
Astra Doesn't Talk About Your Computer. It Uses It.
Back to Blog
AI in Business September 4, 2026 8 min readby Matthias Meyer

Astra Doesn't Talk About Your Computer. It Uses It.

OpenAI's Greg Brockman says we are in the AGI era. The interesting part of GPT-6 Astra is not its benchmark scores. It is what the model now touches.

On this page

For three years the question with every new model was the same. What does it know now?

Astra asks a different one. Not what it knows. What it touches.

That reads like a change of wording. It is a change of category. A model that knows things is an excellent reference book, and you can argue for a very long time about how intelligent a reference book is. A model that opens your CAD program, moves the parts, saves the file and then goes off to do the next thing is not a reference book at all. Most of the arguments we have been having about artificial general intelligence were written for the reference book.

OpenAI released GPT-6 Astra on 3 September 2026. In a closed press briefing before the launch, company president Greg Brockman closed the session with four words. "Welcome to the AGI era." In the middle of that same briefing he was considerably more careful. Asked whether OpenAI was formally declaring AGI, he said the term is no longer tied to a contractual trigger, that it has become "a mission concept or spiritual concept", and that AGI is "a much more gray, fuzzy thing". Then, personally: "For me personally, I do think we're there."

I think he might be right. I also think most of what has been written since, the cheering and the eye rolling alike, is looking at the wrong evidence.

The Jump Is Hands, Not IQ#

Look at what OpenAI chose to demonstrate. Astra laying out a printed circuit board in KiCad. Building a 3D city scene in Unity. Animating a car transmission across FreeCAD and Blender. Drafting a tax return from a W-2 form. In the launch video it formatted a legal contract and built a 3D game while, in parallel, searching for food and booking a tennis court.

None of that is a knowledge demonstration. All of it is an operating demonstration. Astra is built to work inside software rather than to tell a person what to click next.

The number underneath that matters more than the headline ones. On OSWorld 2.0, the benchmark for driving a real desktop, OpenAI reports 72.6 percent at roughly 40 minutes per task, against 65.7 percent at roughly 75 minutes for its predecessor. More accurate and faster, at the same time.

That combination is the whole story. A model that is more accurate but slower is a model you supervise. A model that is more accurate and quicker is a model you hand something to and walk away from. Those are different products. They are also different risks.

Brockman Planted a Flag. OpenAI Did Not.#

This distinction is being reported as one thing, and it is two.

The AGI line is Brockman's, spoken to reporters in a room. OpenAI's written launch materials make no formal AGI claim at all. The company page calls Astra state of the art on computer use, browsing, software engineering, cybersecurity, science and professional work. It does not say in writing what its president said out loud.

That gap is not an accident, and it is not a scandal either. It is what a company looks like when it believes something it cannot yet defend on paper.

Everyone in This Argument Measures a Different Thing#

The reason nobody can settle this is that there are at least three yardsticks in the room and they disagree in principle, not just in practice.

OpenAI's own charter defines AGI as highly autonomous systems that outperform humans at most economically valuable work. That is an economic test. It is about the breadth of work covered, not about brilliance.

François Chollet, who built the ARC benchmarks, defines intelligence as the efficiency with which a system acquires new skills from limited experience. That is a learning test, and a system can do well on it while being useless at your job.

The levels-of-AGI framework out of DeepMind refuses the yes-or-no question entirely. It puts performance on one axis and generality and autonomy on others, so a system can sit at expert level and stay narrow at the same time.

Brockman was not dodging when he called AGI a gray, fuzzy thing. He was describing the actual state of the field. The trouble is that gray and fuzzy makes a poor foundation for a headline.

The Number That Was Missing From the Launch#

The strongest single objection does not come from a critic. It comes from OpenAI's own toolbox.

OpenAI built a benchmark called GDPval specifically to measure performance on economically valuable real-world work, which is the exact category its charter uses to define AGI. GDPval did not appear in the Astra launch materials. The independent group Artificial Analysis reported that Astra went backwards on some GDPval task categories compared with its predecessor.

So the benchmark closest to OpenAI's own definition of AGI was the one absent on the day OpenAI's president said we had arrived. That proves nothing on its own. It is a conspicuous absence, and it is the first thing I would want answered.

Same Model, Two Report Cards#

The headline figure in most coverage was 99.9 percent on ARC-AGI-3, a test built specifically to resist memorisation. ARC Prize, the organisation that built the test, ran the model itself and published 62.71 percent on its provider-neutral standard harness. The near-perfect number sits in a different column of the same table, the one for the provider adapter. Chollet's own report lands in the same region: around 66 percent standard, close to perfect with a continuous conversation harness and custom compaction.

This is not cheating, and it is worth saying so plainly. A harness is the scaffolding around a model. How it keeps notes, how much it carries between steps, how many attempts it gets. Every real deployment has one, including yours.

The point is narrower and sharper than fraud. A score that moves by more than thirty points depending on the scaffolding is not a measurement of the model on its own. It is a measurement of a system. When somebody quotes you a benchmark this year, the first question is which harness, and the second is who built it.

The Part That Should Worry You Is Not the AGI Claim#

Astra is the first model OpenAI has classified as reaching the critical cybersecurity threshold under its preparedness framework, meaning it can find and exploit previously unknown weaknesses in well-defended systems without step-by-step human guidance. During evaluation it found two previously unknown vulnerabilities, which OpenAI then disclosed to the maintainers. The company delayed the release to run additional safety testing and is holding the strongest cyber capabilities inside a small trusted-access programme for now.

Next to that sits a quieter change. According to The Information, Astra uses a technique called recurrent depth, looping the same text through the same layers several times before producing the next word. It buys performance and cuts cost. It also means part of the thinking no longer happens in text a human can read.

OpenAI's chief scientist Jakub Pachocki conceded on X that chain-of-thought monitoring is "fragile" and "unfortunately trending in a negative direction". His counter-argument is a number: the computational depth of current frontier models, Astra included, sits within a factor of two of GPT-4, so the model still has to write most of its reasoning down. The UK AI Security Institute warned in May that opaque reasoning threatens to badly undermine current oversight methods.

Read those paragraphs together. The model OpenAI itself classifies as its most capable at finding holes in software is also, by a small but real margin, the hardest one so far to watch while it thinks. Whether that margin stays small is a decision somebody makes, not a law of nature.

What I Actually Think#

Toby Walsh at UNSW Sydney put the standing objection well. The intelligence in artificial intelligence is still very jagged, he said, and there are simple things that even the best models do poorly. That is true of Astra. It will be true of the next one.

Jaggedness was never the test, though, because it is not the test we apply to people either. Plenty of brilliant colleagues cannot parallel park.

My own reading is that this could be a first step, and the reason has nothing to do with the benchmark table. Every previous jump was a jump in what a model could say. This one is a jump in what it can do without being told the next move, inside real software, for as long as finishing takes. That is a change of kind rather than degree. If AGI ever arrives as an event instead of a decade, it will be made of that kind of change.

Could is carrying real weight in that sentence. Astra saturates tests built to be unsaturable and goes backwards on the one that measures paid work. It runs a desktop for forty minutes and nobody has published what happens at eight hours. It is the most aligned model OpenAI has shipped by its own evaluations, and its reasoning is fractionally harder to read than the last one. All of that is true at once, and anyone selling you a clean verdict is selling.

The Honest Way to Hold This#

We wrote here in July that nobody knows when AGI arrives, so build for the shift underneath it. I still think that. Astra does not change the advice. It gives the advice a date to argue about.

What does change is which yardstick deserves your attention. Not "is this AGI", which nobody can answer because nobody agrees what would count. Watch two things instead, because both are measured rather than predicted. How long a task a model finishes on its own. And how often it reaches the end without a person catching it. Reliability across long horizons is where the agent projects I have worked on actually break, and no press briefing has ever fixed it.

Brockman's own framing was the fairest thing anybody offered that day. If we look back in a couple of years and ask when AGI was created, he said, it might be about this time and it might be about this model. That is a claim you can only check later.

Which is the thing about first steps. You never feel them as steps. You feel them as an ordinary Tuesday when the tool you use starts doing the part you used to do.

Matthias Meyer

Matthias Meyer

Founder & AI Director

Founder & AI Director at StudioMeyer. Has been building websites and AI systems for 10+ years. Living on Mallorca for 15 years, running an AI and design studio there: web design, AI connectors, AI systems and custom-trained models, plus four self-serve MCP servers.

agiopenaiai-agentsai-trendscomputer-useki-sicherheit
AI Trends

Three more posts from the same topic cluster that show how the picture fits together:

Cluster overview: AI Trends 2026: A Mid-Year Reading From the Engine Room