知識がなくても始められる、AIと共にある豊かな毎日。
AI News & Trends

What GPT-6 Astra Actually Changed — Reading the Specs and the Conditions

swiftwand

Every new model arrives with the same promise: it is smarter now. The bars grow, last year’s best score becomes this year’s worst, and nothing on your own workbench changes. You still draw the part yourself, still work out the quote yourself, still click the order button yourself.

GPT-6 Astra, released on 3 September 2026, comes at that gap from a different angle. The first strengths the announcement names are not answer quality but computer use, browser use, software engineering, cybersecurity, science, and professional work. Half of that list is the sequence of steps a person used to perform in front of a screen.

A note on this article’s position before anything else. We do not hold a paid tier and have not tried any of the behaviour described below. So you will not read that it felt fast, or that it did not work. Instead this reads only what the developer has published — the announcement, the API reference, the price table, and the system card — and lays it out in a form a working reader can use. Every figure was checked as of 6 September 2026.

There is a reason not to work from someone else’s summary, and this article demonstrates it rather than asserting it. Summaries drop the rounded figure and the qualifying footnote at the same time, and they drop them in a consistent direction: what gets cut for readability is almost always the part describing the conditions under which it was measured.

忍者AdMax

From a contest about intelligence to a contest about touching the screen

Line up the uses the announcement names and the direction is clear. Filling in online forms. Updating a customer record system. Tidying a calendar. Researching, then writing the summary into an email or a document editor. Analysing scientific data and producing the figures. Building a site and then checking in the front end that its features actually work. Installing software, testing it, and isolating the faults that appear on screen.

None of that is about being clever. It is about speed of hand and reliability of sequence. Which is to say the announcement is aimed less at the question of what a model knows and more at the question of what it can be left to do.

What actually moves, seen from the workbench

Put this in concrete terms for someone running a 3D printer.

Not much of the daily work needs judgement. Which shape to use, which material, how to take tolerances — you have to decide those yourself. But a good deal of what remains after the decision contains almost no judgement at all. Applying material cost and electricity to a formula to get a unit cost. Reading a bill of materials to count what you need. Digging through order history for what you paid last time. Adding a row to the print log. Sorting photos by date.

These take no thought and a great deal of time. And as long as a person is doing them by hand, they happen one at a time. That this class of work has been named as the target is what the announcement means in practice.

The same work still splits by character, though. Recalculating a cost can be done again as often as you like; pressing an order button happens once. Adding a log entry can be corrected later; something published outside your own machine leaves a trace even after you withdraw it. Separating work you can redo from work you cannot is the first step in deciding how to delegate. We return to that split later.

TaskRedo?Easy to notice a mistake?
Apply the formula to get a unit costYesVisible on inspection
Count required parts from a BOMYesVisible on inspection
Append a row to the print logYesCorrectable later
Sort photos by dateYesVisible on inspection
Order consumablesNoSometimes not until it arrives
Publish design files externallyNoA withdrawal still leaves a record
Tidy up and delete old filesNoYou may never notice they went

The published specifications

These are the values as the API reference gives them.

ItemValue
Model IDgpt-6-astra
Context window1,050,000 tokens
Max output128,000 tokens
Knowledge cutoff30 April 2026
Standard inputUSD 10.00 per million tokens
Standard outputUSD 50.00 per million tokens

Supported features listed alongside these include streaming, structured outputs, function calling, file search, image input, web search, and prompt caching. Audio and video input or output do not appear in this table.

A million-token window is not a million tokens of material

The easy thing to miss in that table is that the context window is input and output combined. The reference publishes two numbers — a 1,050,000-token window and a 128,000-token maximum output — and no separate ceiling for input alone.

So “handles a million tokens” is accurate, but it does not mean you can throw a million tokens of material at it. Whatever the reply consumes comes out of the same allowance, and the longer an answer you ask for, the less room is left for source material. The reference offers no further explanation of where that line sits.

The practical consequence is simple. If you size your material against the full million, you will come up short by whatever the output takes. Assuming you use the output ceiling in full, the arithmetic leaves roughly 920,000 tokens for material — but that figure is not published; it is a subtraction from the two numbers that are.

Worth holding onto separately: 128,000 tokens of output is small next to the input side but substantial as a product in its own right. Rewriting a long document whole, emitting a large configuration file in one pass — that is where the room shows up. Put the other way, input size and output size are separate constraints, and having slack in one does not help when the other runs out.

All of which is about what fits, not about what is worth putting in. More material does not reliably produce a better answer, and as we will see, the cost side branches on its own. A ceiling is where a design starts, not what it aims at.

Knowledge stops at the end of April 2026

The knowledge cutoff is 30 April 2026. Release was 3 September, so there are four months the model knows nothing about.

That is less a weakness than a condition of use. Specifications published recently in your field, price revisions, parts that went end-of-life — none of it is inside the model. Web search is supported so it can be sent to fetch them, but whether it actually did is something you have to check.

For 3D printing, any filament or printer model released in those four months falls on the blank side. Supply recent information yourself is a rule that survives changes of model.

The awkward part is that a blank does not always present itself as one. If the model says it does not know, you can act on that; if the gap is filled with older information, it looks correct from where you sit. A price changed, a specification was revised, a successor shipped. Each is a defensible answer as of April, which is exactly what makes the error hard to catch.

The remedy is unglamorous and fixed. For anything where the date matters, make it show its source. If it cannot produce one, treat the answer as provisional.

The benchmark numbers are maximum-effort values

Underneath the score tables sits a line stating that the values are the maximum at any effort setting. Effort is the dial that decides how much thinking the model spends on a request, and spending more costs more and takes longer.

So the published scores are the ceiling, reached under the most generous setting. They are not what a default configuration returns. This is the single condition most often lost when a number is quoted second-hand.

There is a second, sharper example of the same problem in the announcement itself. The prose gives the FrontierMath Tier 4 score as 98%, while the table gives 97.6%. The rounded number reads better, so explainer articles tend to quote it. Four tenths of a point does no harm here, but the same habit operates on every other figure. When an announcement has a table, read the table.

Where it jumped

BenchmarkGPT-6 AstraGPT-5.6 Sol
ARC-AGI-399.9%7.8%
ARC-AGI-295.0%92.5%
FrontierMath Tier 4 (v2)97.6%83.0%
Terminal-Bench 4.057.9%37.3%
OSWorld 2.0 (computer use)72.6%65.7%
ScreenSpot-Pro (no tools)92.7%76.9%
MRCR v2 8-needle 512K–1M96.3%73.8%

The largest move is the ARC-AGI-3 row, from 7.8% to 99.9%. That row carries a footnote noting a changed execution setting, so it is not a like-for-like jump.

Note also that the computer-use row, at 72.6%, is nearly three failures in ten. Next to the nine-in-ten figures familiar from text evaluations that looks weak, but the task is different in kind: you follow a screen through a sequence to a target state, and one wrong step anywhere means you never arrive.

Not first at everything

Set against other vendors, the ordering changes by task. On Terminal-Bench 4.0 the gap is wide — 57.9% against 55.8% for Claude Fable 5.1 and 52.6% for Claude Opus 5. On DeepSWE v1.1 it is 74.1% against 73.8% for Gemini 3.8 Flash and 73.7% for Claude Opus 5, which is inside a point. On the Artificial Analysis Intelligence Index it sits at 61.2 against 65.7 for Claude Fable 5.1 — behind.

There is a footnote worth carrying: for ScreenSpot-Pro and ExploitGym, the competitor scores reported come from a variant with fewer safeguards. Those two rows cannot be read as product-to-product comparisons.

It is tempting to rank models on one line, but the order changes with the task. The approach of choosing by use rather than by rank is set out in our reverse guide to picking a model by purpose and budget.

Rollout is staged, and off by default inside organisations

Availability is worth reading closely. Astra went out on release day to a limited set of organisations, expanding over the following days to ChatGPT Plus, Pro, Business and Enterprise users, and through the OpenAI API, Microsoft Azure and Amazon Bedrock.

Usage sits inside existing subscription allowances; where that is not enough, additional credits can be bought. Pro, Business and Enterprise subscribers also get access to a separate GPT-6 Astra Pro.

The clause that matters for anyone in a company: enterprise administrators have to enable Astra for their workspace, and access is off by default at launch. If you read the announcement and then cannot find the model, that is the likely reason, and it is not something you can resolve yourself.

For API customers there is also Zero Data Retention where eligible, and a Private Safety Processing scheme described as still being tested.

Pricing looks unchanged, but it branches

Standard API pricing is USD 10 per million input tokens and USD 50 per million output. Separate rates apply to cache reads and writes, and a Fast mode runs at up to twice the speed for twice the price.

The branch that catches people is elsewhere: above 272,000 input tokens, the whole request is billed at 2x the input and cache rates and 1.5x output. It applies to the full request, not just the tokens past the line. The arithmetic and where it bites are worked through in our piece on where the price table jumps.

What it now refuses

The announcement is unusually specific about the boundary. Defenders can use Astra for tasks like secure code review and patching, but it will refuse more advanced cybersecurity work such as writing proof-of-concept exploits. Through a programme called OpenAI Daybreak, access to less restrictive safeguards is planned to expand in the following weeks, covering vulnerability and proof-of-concept validation, malware analysis, and detection engineering.

This is the first model the developer classifies as reaching the Critical level for cybersecurity capability under its Preparedness Framework, which is what the refusal boundary follows from. The reasoning and the numbers behind it are covered in the article on monitorability and cyber capability.

There is a disclosed regression as well

Launch material usually records only what improved. This one carries a section saying something got worse, and not as a buried aside — it stands as its own heading. Chain-of-thought monitorability, meaning whether a person can follow after the fact the path the model took to an answer, is stated by the developer to have declined relative to the previous generation.

It matters here because monitorability is what makes a long automated sequence auditable. If less of the reasoning is legible, more of your confidence has to come from checking the result instead. That trade is examined in the monitorability article.

What is not settled

Reading only published documents leaves gaps, and it is worth naming them rather than filling them in.

  • Actual behaviour. We have not run it. Nothing here is a report of how it feels in use.
  • Default-setting performance. Published scores are maximum-effort values; what a default configuration returns is not published.
  • Real billing. The pricing figures are the published rates. We have not reconciled them against an invoice.
  • How the refusal boundary behaves in practice. The announcement describes the policy; where the line actually falls on a given request is not something a document can tell you.
  • What changes as Daybreak expands. The relaxation is described as planned, in the coming weeks. That is a stated intention, not a shipped state.

Taking the measure of where things stand

Three things are worth carrying out of this.

First, the target moved from answering to acting. The strengths named first are computer use, browsing, software engineering, cybersecurity, science, and professional work. That reframes the question from what a model knows to what you are willing to let it do.

Second, every number arrives with conditions, and the conditions are what summaries drop. Maximum effort. Reduced safeguards on some competitor rows. A changed execution setting on the largest jump. A prose figure that disagrees with its own table by four tenths of a point. None of it makes the numbers useless; all of it makes second-hand numbers unreliable.

Third, the limits are shaped differently than the headline suggests. A million-token window that input and output share. Knowledge that stops four months before release. A price that doubles above a threshold you can cross without noticing. Access that an administrator has to switch on.

For the workbench, the practical move is unchanged and unglamorous: split the work into steps, sort those steps by whether you can undo them, and hand over only the ones you can. How to run that sort is set out in the reverse guide to deciding by cost and reversibility.

Sources

ブラウザだけでできる本格的なAI画像生成【ConoHa AI Canvas】
ABOUT ME
swiftwand
swiftwand
AIを使って、毎日の生活をもっと快適にするアイデアや将来像を発信しています。 初心者にもわかりやすく、すぐに取り入れられる実践的な情報をお届けします。 Sharing ideas and visions for a better daily life with AI. Practical tips that anyone can start using right away.
記事URLをコピーしました