GPT-6 Astra: Brilliant Inside the Right Box
Astra is token efficient and strong on long-running work, if you build it a contained environment and, genuinely, motivate it. Outside that setup it is sloppy and takes liberties. 6.7 out of 10.
I have had GPT-6 Astra for less than a day. Call this a first impression rather than a verdict, and expect me to revisit it, because a few hours is not enough to be confident about a model either way.
OpenAI hyped this one as the next level. The thing that would change how I work.
My honest position is that it depends almost entirely on how you set it up, which is not the answer anyone wants about a flagship model.
Give it a box and it is excellent
Put Astra in a contained environment. Give it the skills it needs, tell it how you want code written, and, I am aware of how stupid this sounds, motivate it. We have genuinely arrived at the point where encouraging a model changes its output.
Do all of that and it excels. It is very good at long-running tasks, the kind that would exhaust most models halfway through, and it is remarkably token efficient while doing them. It gets to the point and stays there. There is a genuinely good model somewhere inside this thing.
Outside the box it gets lazy
The problem is that it does not bring that capability on its own. Without the setup, it is simply lazy. The attention to detail is shocking for something this capable, and the gap between what it clearly can do and what it bothers to do is the most frustrating part of using it.
It also tests things that do not need testing. It will produce work, then go and verify some part of it that was never in question, while the actual problem sits untouched.
It takes liberties
There is a streak in Astra that I do not like. If I have a bug in my HTML, it will rewrite the HTML. Not point at the bug, not ask. Rewrite it, and hand me a much larger change than the one I wanted.
Give it a simple task and there is a real chance it will misread you in a way that costs you. Not a small misunderstanding. A confident wrong turn on something that was not ambiguous, followed by a solution far larger than the problem.
Against Fable 5.1
Both of these are strong models and they are strong at opposite things.
Fable 5.1 is the one you do not have to watch. It writes the minimum, it finds a bug and asks whether you want it fixed, and I will merge its work without reading it. It needs no scaffolding to behave that way.
Astra needs the pen built for it, and inside that pen it out-works Fable on long grinding tasks and costs fewer tokens getting there. Outside it, Astra is the one rewriting your HTML while Fable is the one asking permission. Pick accordingly.
They did fix Sol
Credit where it is owed. Sol was awful to talk to and awful to deal with, and Astra is not. The interaction itself is fine now. They solved that problem.
What replaced it is sloppiness, which is not what you expect from a model this strong. I would rather have a model that is unpleasant and careful than pleasant and careless.
The verdict: 6.7/10
I rate GPT-6 Astra 6.7 out of 10.
That number is for the model as it behaves by default. Set it up properly and it is worth considerably more than 6.7, which is the most interesting and most annoying thing about it. Less than a day of use is not much, so treat this as provisional. I will write a second verdict once I have run it against real work.