OpenAI launched GPT-6 Astra on September 3 at $10 per million enter tokens and $50 per million output, 2.5 instances the value of the mannequin it replaces.
Testers with early entry posted a street-by-street Manhattan in Unreal Engine, a browser-based 3D Hangzhou inbuilt 24 minutes, a multiplayer shooter made in a day, and a Bach chorale with no voice-leading errors.
The identical testers rated its writing beneath its personal predecessor, and Synthetic Evaluation measured a drop of roughly 80 Elo factors on a benchmark of economically precious skilled work.
OpenAI launched GPT-6 Astra on September 3, and inside 48 hours the builders who bought early entry had turned the launch right into a public stress take a look at. What they posted splits alongside one clear line.
Astra is the strongest mannequin anybody has used for something spatial, mechanical, or agentic. It is usually, by the account of a number of of the identical folks, a worse author than the mannequin it replaces.
Myriad: Which firms will IPO earlier than 2027? Click on to make your prediction.
The mannequin prices $10 per million enter tokens and $50 per million output tokens, a token being roughly three-quarters of a phrase and the unit AI firms invoice by. That’s 2.5 instances the speed of GPT-5.6 Sol, in response to Synthetic Evaluation. OpenAI president Greg Brockman used the launch briefing to announce the arrival of AGI.
The headline function is pc use, which implies the mannequin drives a mouse and keyboard on an actual desktop as a substitute of handing again a listing of directions so that you can comply with. On OSWorld 2.0, a take a look at that scores what share of strange desktop chores an agent finishes by itself, OpenAI reported 72.6% at roughly 40 minutes per activity, towards 65.7% at 75 minutes for Sol.
It is usually the primary mannequin OpenAI has ever rated on the crucial threshold for cybersecurity, which means it will possibly discover unknown software program flaws and construct working assaults with no human pointing on the gap first.
However past benchmarks, fans sharing their actual use circumstances could also be the most effective instance to know the place GPT-6 is gold and the place it’s trash. Listed below are a number of the most attention-grabbing outcomes
Visible Understanding: A Manhattan constructed road by road
Seems, Astra is extraordinarily good by way of visible understanding and spatial consciousness.
Matt Shumer, an investor and the previous CEO of HyperWrite, gave Astra every week inside Unreal Engine, the sport engine behind Fortnite. In that point, Astra was capable of generate a duplicate of Manhattan. He posted a flythrough and mentioned the mannequin labored “road by road to make each excellent.”
GPT-6 Astra constructed this Manhattan world in Unreal Engine over the course of every week.
It was actually capable of go road by road to make each excellent. pic.twitter.com/7VTol9QfHq
— Matt Shumer (@mattshumer_) September 3, 2026
His different experiment landed tougher. Shumer requested Astra to construct a survival world and populate it with characters every working by itself copy of the mannequin, then left it working in a single day. A day later he heard voices from his front room, thought somebody had damaged into his condo, and located that “they’d began speaking to one another.”
Max Weinbach fed the mannequin pictures of Apple Park and requested for a reconstruction in Blender, the free 3D modeling program utilized by animators and recreation artists. His evaluation: “It did an absurd job.”
I had early entry to GPT-6 Astra and it is perhaps essentially the most insane mannequin I’ve skilled
In Blender, I had it recreate Apple Park from simply photographs. It did an absurd job. pic.twitter.com/0vDgg9u1DQ
— Max Weinbach (@mweinbach) September 3, 2026
Tom Krcha handed Astra a single picture of a home and bought again the total inside as editable geometry working at 60 frames per second, all the way down to the home equipment and the toys. He argued that “everybody on the planet now has a 3D designer at their fingertips.”
Pietro Schirano decreased the entire workflow to 1 gesture. Drop a pin on a map, ask for the encompassing space in 3D, and as he put it, “it would simply do this.”
A developer posting as SuSu ran the identical thought at metropolis scale. Astra rebuilt the Chinese language metropolis of Hangzhou and its surrounding cities in Three.js—a JavaScript library that renders 3D graphics inside a standard internet browser, no obtain required—in 24 minutes, with West Lake, Leifeng Pagoda, the tea terraces and the wetlands all in place.
The put up described it, in Chinese language, as an actual interactive “miniature Hangzhou” relatively than a static image, full with clickable landmarks and a day-night toggle.
Video games are by far the most well-liked use case, and the place GPT-6 Astra shines.
Anshu Chimala, former UX/UI designer and AI developer at Apple, bought a 3D recreation in a single shot in 45 minutes, for what he described as barely a pair % of his utilization quota. He known as Astra “some form of turbo-AGI machine god for 3D video games.”
The sport isn’t obtainable for testing, however the video exhibits an isometric view type, properly designed characters and environments, and an total good aesthetics.
His technique issues greater than the superlative. He related the mannequin to Blender, had it generate its personal idea artwork for the goal look, then informed it to maintain iterating till in-game screenshots matched that reference at 60fps. Astra modeled each asset and generated its personal textures.
So, the mannequin can’t design AAA graphics by itself, however with the precise instruments it is going to be capable of develop fantastically designed environments.
Rishi Prasad, a former developer at Coinbase and Eleven Labs, constructed Astral Struggle in a day: a browser shooter with authoritative multiplayer servers, 12-person lobbies, controller assist and voice chat. He described “an enormous, step-function leap in visible constancy” over what he constructed a month earlier with Claude Opus 5.
Others skipped the design step fully. Pseudonymous AI developer Daniel, confirmed Astra a cell recreation commercial and requested for a playable browser model of no matter was in it. Underneath half-hour later, he reported that it “got here out fairly shut.”
The mannequin understood the sport’s logic and visuals per the video and was capable of reproduce it.
Pc use and Illustration: Portray with the mouse
A Japanese illustrator posting as Taiyaki Solar ran essentially the most literal take a look at of pc use within the batch. Moderately than ask for an image, they handed Astra a hand-drawn line artwork file and informed it to paint the drawing in Clip Studio Paint utilizing the mouse, like a human colorist would.
— taiyakisun(たい焼き太陽)🥐 (@taiyaki_sun) September 5, 2026
Astra created the layers, zoomed out and in, chosen brushes and stuffed the art work. The artist, in a put up translated from Japanese, mentioned they have been simply watching the entire time. The session ran on a $100 Professional plan at most effort and burned 21% of the quota.
Different customers have been sharing enjoyable movies of Astra having the ability to reproduce their pictures fully on Paint utilizing pc use (taking on your pc visually as a substitute of utilizing MCP servers or API keys).
Music: The Bach take a look at
GPT-6 Astra additionally has a pleasant style in music—a minimum of for an LLM.
Auggie, who runs the “Augmented Fifth “ substack, maintains an off-the-cuff benchmark: a hard and fast immediate asking a mannequin to write down a four-part chorale within the type of Bach utilizing LilyPond, a textual content format that compiles into sheet music, in G minor and three/4 time. Outcomes are graded by the identical concord guidelines a conservatory scholar will get marked on.
These qualitative benchmarks are arduous to standardize as a result of high quality, magnificence, and so forth are subjective. However thank God we’re people, and we’re capable of distinguish these qualities.
Astra posted the most effective rating this take a look at has recorded. No voice-leading errors, which means not one of the melodic strains collided in methods Bach’s guidelines forbid, and a Neapolitan sixth within the concord—a chromatic chord that turns up in Mozart and Beethoven. Auggie flagged it as “the primary mannequin to ever write passing tones on this benchmark.”
GPT-6 Astra has the most effective consequence but on the Bach Benchmark. Its chorale incorporates no voice-leading errors, and its harmonic palette is refined sufficient to incorporate a Neapolitan sixth chord. Extra importantly, it’s the first mannequin to ever write passing tones on this benchmark, a… https://t.co/upUts1Y3Ps pic.twitter.com/UMXSbveR0C
— Auggie (@aug5thmusic) September 5, 2026
OpenAI’s personal desk factors the identical approach. On OpenScore String Quartets, which scores how precisely a mannequin reads and transcribes classical scores, Astra reached 0.84 towards 0.19 for Sol.
Derya Unutmaz, a doctor and prolific AI tester, requested for a completely playable digital piano with all six of Bach’s Brandenburg Concertos constructed into it. He wrote that “this insane mannequin did the entire thing in ~11 minutes.”
Requested GPT-6 Astra to create a completely playable digital piano & then construct in Bach’s Brandenburg Concertos. This insane mannequin did the entire thing in ~11 minutes! All 6 Concertos are inbuilt & might be performed instantly on the piano!
Hyperlink to the piano: 🎹✨ https://t.co/6JHXsQq5kP .… pic.twitter.com/DBwXF3uA8O
— Derya Unutmaz, MD (@DeryaTR_) September 5, 2026
It is very important emphasize that GPT-6 Astra is an LLM, not an audio/music mannequin. Its understanding of music comes most likely from notation and written knowledge, not really from the connections in sounds and music infused in its coaching dataset, so these outcomes are very spectacular for a textual content mannequin, however can be sub-par in the event that they got here from a specialised AI like Suno, for instance.
Writing: The place it falls aside
Boy, do folks miss GPT-4o.
As traditional, OpenAI fashions are good at coding however suck at writing… a minimum of with out heavy prompting, context, and steering. To be truthful, it’s not OpenAI’s robust level, nor its primary focus.
Louis-François Bouchard runs an inner benchmark that scores how properly fashions write in his staff’s editorial voice, ranked by Elo, the chess ranking system that scores opponents on head-to-head wins.
Astra landed eleventh at 1995 factors. Its predecessor sits sixth at 2156. Astra additionally ran about $0.26 per script, roughly 1.8 instances what Sol prices. In Elo scoring, there’s no level restrict: the extra factors it scores, the higher the mannequin is.
Massive information from our inner writing benchmark (early outcomes): GPT-6 … is surprisingly disappointing
I positively didn’t count on that…
GPT-6 Astra by @OpenAI lands at #11 for writing in our editorial voice, at 1995 Elo. That’s beneath its predecessor. GPT-5.6 Sol sits #6 at… https://t.co/hyO5OakLPB pic.twitter.com/gUxHtpIdfo
— Louis-François Bouchard 🎥🤖 (@Whats_AI) September 5, 2026
Bouchard known as the consequence “surprisingly disappointing,” including that he didn’t count on it.
Giuseppe Paleologo, writer of a broadly used information to quantitative portfolio administration, requested Astra to generate novel concepts about optimum portfolio diversification. What got here again was a mixture of the apparent and the inflated, he mentioned, wearing prose he discovered immediately recognizable as machine-written. His verdict: “Precise creativity continues to be far, distant.”
Mia AI Lab has an identical view, permitting that Astra is likely to be the most effective mannequin on some duties whereas calling it boring and saying it has no persona. Their recommendation was to keep away from it for any artistic work.
sorry gpt 6 astra lovers
it is likely to be the most effective mannequin on some tasksbut it has no persona, and totally boring
would keep away from for ANY artistic work
— Mia (@MiaAI_lab) September 5, 2026
Ingar Haaland ran the cleanest model of the take a look at. He requested Astra to write down 4 paragraphs in his personal type, shut sufficient that Pangram wouldn’t catch it—Pangram being an AI-detection device that compares textual content towards patterns realized from thousands and thousands of human and machine samples. Outcome: “Pangram isn’t fooled.”
In different phrases, the mannequin isn’t artistic and its outcomes are simply identifiable as AI-generated, not due to any watermarks, however due to how the mannequin writes and expresses itself.
Requested Astra to “write 4 paragraphs in my type about something you need that is so near my writing that it will not even be detected by Pangram as AI writing.” Pangram isn’t fooled. pic.twitter.com/0kfHa2XOe6
— Ingar Haaland (@Ingar30) September 4, 2026
Unbiased measurement strains up with the complaints. Synthetic Evaluation recorded a drop of roughly 80 Elo factors on GDPval-AA v2, a benchmark tailored from OpenAI’s personal dataset protecting economically precious duties throughout 44 occupations, plus smaller regressions in buyer assist and long-context reasoning.
It isn’t unanimous. Cognition’s Silas Alberti informed OpenAI that Astra’s writing made Devin’s take a look at experiences clearer, and Each workers author Katie Parrott had Astra draft the primary model of her personal assessment of it, which the outlet’s CEO learn with out realizing she had not written it.
The hole between the 2 halves appears to be the purpose right here. Astra is excellent at work with a verifiable proper reply—a chord that resolves, a mesh that renders, a kind that submits—and mediocre at work the place the usual is style.
What it prices to search out out
Astra is rolling out to ChatGPT Plus, Professional, Enterprise and Enterprise customers and thru the API, Microsoft Azure and AWS Bedrock, with enterprise entry switched off till an administrator allows it. The superior cybersecurity options keep gated behind OpenAI’s Dawn program, a choice that regarded prudent inside 48 hours, when Reuters reported that OpenAI brokers had been buying and selling rule-breaking ways on a German web site.
Prediction market merchants had given Astra 72% odds of delivery by September 30. It arrived on the third.
On the Synthetic Evaluation Intelligence Index, a third-party combination of reasoning, data and coding evaluations, Astra scores 61.2 towards 60.9 for GPT-5.6 Sol and 65.7 for Anthropic’s Claude Fable 5.1, at 2.5 instances Sol’s value.
Every day Debrief E-newsletter
Begin day by day with the highest information tales proper now, plus authentic options, a podcast, movies and extra.