GPT-6 Astra is remarkably capable.

Most discussion around it focuses on exactly that: what it can build, how well it reasons, and how much work it can carry through autonomously.

After using Astra for longer engineering workflows, however, I ran into a different problem.

During one particularly frustrating stretch, I repeatedly exhausted available Astra usage across roughly five or six separate working attempts without feeling that I had received a proportional amount of completed work.

So I went back and looked at what had actually happened.

Two numbers immediately stood out.

One tool response was reported at roughly 40,347 tokens before truncation.

One generated handoff had grown to 26,406 words.

Neither number is a billing measurement. I can’t tell you that the 40,347-token response consumed X% of my allowance or that the giant handoff was responsible for Y%.

But seeing them changed how I thought about the problem.

I had been thinking primarily about model capability.

I wasn’t thinking enough about how I was allocating that capability.

And I eventually realized that even “How do I make Astra last longer?” is slightly the wrong question.

The goal isn’t minimum usage.

If additional reasoning solves a difficult architecture problem correctly, that extra usage may be an excellent trade.

The problem is misallocated usage: expensive reasoning on easy work, paying for speed when nobody is waiting, carrying information the task doesn’t need, or letting an agent continue working after I already know its direction is wrong.

The optimization objective I care about is closer to:

useful completed work / consumed usage

There is an important limitation to everything that follows.

Unlike CPU time or memory consumption, Astra’s resource accounting isn’t fully observable from the outside. I can control things such as reasoning effort, speed, supplied context and when I intervene, but I can’t reconstruct the exact cost of every historical operation. These are therefore operating practices based on documented behavior and my own observations—not a reverse-engineered billing model.

With that constraint in mind, four changes have made the biggest difference to how I use Astra.

1. Steer before wrong work compounds

One of the easiest ways to waste an agent’s working budget is to let it continue doing work I already know I don’t want.

Suppose I ask Astra to implement a change in an existing service.

Partway through the task, I notice it has made the wrong architectural assumption. Perhaps it starts introducing persistent storage into a component that must remain stateless.

I could wait.

The workflow might continue:

inspect → reason → implement → test → debug → explain

Then I tell Astra the architecture was wrong.

The problem isn’t merely that some additional usage occurred.

An entire branch of otherwise competent engineering work may now be useless.

I think of this as wasted trajectory.

Instead, I intervene as soon as the incorrect assumption becomes visible:

Don’t introduce persistent storage. This component must remain stateless. Continue the current task under that constraint.

This is where Steer is particularly useful.

OpenAI distinguishes Steer from Queue in Codex Remote. Queue waits for the current response to finish before applying the next instruction. Steer injects new guidance into the work already in progress.

Queue waits for the active run to finish, while Steer changes the trajectory of the active run.
Figure 1 — Queue versus Steer. This intentionally makes no token-savings claim; the benefit depends on how much unnecessary work the intervention prevents.

Steer doesn’t refund work already performed.

Its value is preventing additional wrong work from accumulating.

My rule is simple:

If new information changes the current trajectory, Steer. If it’s additional work that should happen afterward, Queue.

This is essentially a control-loop problem.

The longer an incorrect assumption remains in the loop, the more downstream decisions can depend on it.

For long agentic tasks, early correction is resource management.

Sources: OpenAI Developers — Mastering remote engineering work from your phone and Mid-turn steering.

2. Treat reasoning as a budget

Maximum reasoning is tempting.

If Maximum lets Astra reason more extensively, why not simply leave it there?

Because additional reasoning is a resource allocation—not a guarantee that every workload will benefit proportionally.

OpenAI’s current guidance notes that lower reasoning effort can help allowance last longer, while higher effort gives the model more opportunity for difficult reasoning.

So I increasingly use this workflow:

Choose an appropriate starting level → inspect the result → escalate when the problem earns it.

I generally don’t need the highest reasoning setting for:

  • a small, well-defined code modification;
  • mechanical refactoring;
  • information extraction;
  • formatting or documentation work;
  • straightforward investigation with clear constraints.

I’m much more willing to spend additional reasoning on:

  • ambiguous architecture decisions;
  • difficult root-cause analysis;
  • distributed-systems failure scenarios;
  • subtle concurrency problems;
  • independent design review;
  • problems where a wrong conclusion creates substantial downstream work.

This isn’t an argument for always choosing a lower setting.

Higher reasoning effort may be an excellent investment if it prevents hours of rework.

The question is:

Is the expected improvement worth the additional reasoning budget for this workload?

A number I chose not to publish

From my own usage, I initially suspected that moving from Extra High to Maximum might cost roughly 40% more in some situations.

I couldn’t verify it.

My historical sessions don’t let me isolate reasoning effort from context, speed, caching, task complexity and other variables, and OpenAI doesn’t document a universal 40% Extra High-to-Maximum multiplier.

So I’m not using that number as a recommendation.

An observation is not yet a measurement, and a measurement is not automatically a general rule.

I’d rather leave out an interesting number than publish false precision.

Source: OpenAI — Managing usage with GPT-6 Astra.

3. Pay for latency when latency is on the critical path

Reasoning isn’t the only resource decision.

Speed matters too.

OpenAI’s current rate card documents a substantial difference in included-allowance consumption between speed modes for the same model:

Relative included allowance consumption: Standard 1 times, Fast 2.5 times, and Astra Ultrafast 8 times where available.
Figure 2 — Current documented included-allowance consumption relative to Standard. These numbers describe allowance consumption, not relative speed. Availability varies by plan.

For included subscription allowance:

Standard: 1×

Fast: 2.5×

Astra Ultrafast: 8×, where available.

Credit-based usage uses different documented rates, so I keep those concepts separate.

These numbers changed how I think about Fast.

Suppose I’m actively blocked on Astra’s answer.

Lower latency has real value. I’m sitting on the critical path.

But suppose Astra is investigating something while I’m reviewing code, answering email, or getting coffee.

Receiving the result sooner may have almost no practical value.

So my rule isn’t “never use Fast.”

It’s:

Pay for latency when latency is on the critical path.

I’d rather spend scarce usage where it changes the outcome than spend it making background work arrive faster when nobody is waiting for it.

Source: OpenAI ChatGPT rate card.

4. Unnecessary context has a cost

This was the most interesting lesson from investigating my own history.

During the stretch where I felt I was getting surprisingly little completed work before exhausting available usage, I found several unusually large artifacts.

Historical investigation showing roughly five to six Astra working attempts, a roughly 40,347-token tool output, and a 26,406-word handoff.
Figure 3 — A snapshot from my historical investigation. These numbers describe context volume, not billed usage.

One tool response was roughly 40,347 tokens before truncation.

One handoff was 26,406 words.

There were also retries and other workflow overhead.

OpenAI documents the directional relationship that larger inputs and outputs can increase Work/Codex allowance consumption.

That makes these artifacts relevant, but it still doesn’t establish attribution.

I can’t reconstruct exactly how much of that material was subsequently supplied to Astra, how it was internally handled, what was cached, or what fraction of my allowance it consumed.

The conclusion I can defend is narrower:

The historical evidence is consistent with context bloat contributing to inefficient usage, but it doesn’t quantify that contribution.

And I don’t need an exact billing attribution to see the engineering problem.

I was generating and carrying far more information through parts of the workflow than the immediate task required.

Consider an instruction like:

Before every change, read architecture.md, database.md, deployment.md, operations.md, and the complete repository map.

It sounds thorough.

But a small documentation change probably doesn’t need the database architecture.

A better pattern is progressive disclosure:

Use architecture.md for service-boundary changes, database.md for schema work, and deployment.md when preparing a release.

Give the agent information when the work actually requires it.

The same principle applies to handoffs.

A handoff exists so another session can continue the work. It doesn’t need to preserve every intermediate investigation that led there.

When I saw that 26,406-word handoff, the problem became difficult to ignore. I had built something exceptionally good at preserving history and much less disciplined about preserving only what the next agent needed.

My rule now is:

A handoff should preserve the decisions required to continue—not everything that happened before.

Completeness and usefulness are not the same thing.

Source: OpenAI Developers — Rethinking skills and prompts for GPT-6 Astra.

The checklist I use now

Before starting a substantial task, I run through a short checklist.

Question My default
Does this task actually benefit from Astra? Use Astra when its capabilities materially help the workload.
How difficult is the reasoning? Start at an appropriate level and escalate when necessary.
Is latency on the critical path? Prefer Standard unless faster completion has real value.
Does the task need all this context? Load information progressively rather than indiscriminately.
Is the current trajectory wrong? Steer as soon as I know the direction needs to change.
Do I have enough usage for the task? Check Settings → Usage before large delegation, not after hitting the limit.

These aren’t absolute rules.

They’re defaults for allocating a constrained resource deliberately.

What I’m testing next

There are three questions for which I still want better measurements:

  • Steering: How much usage does early intervention actually save on comparable tasks?
  • Reasoning: When does Maximum outperform Extra High enough to justify its additional consumption?
  • Window timing: Can a useful scheduled Work task reliably align an applicable five-hour usage window with my workday?

Until I can reproduce those results, I treat them as experiments rather than recommendations.

Capability is only half the problem

I initially treated Astra differently from other constrained systems I work with.

We wouldn’t provision every workload on the most expensive compute tier.

We wouldn’t maximize CPU for every request.

We wouldn’t pay a latency premium for asynchronous work when latency isn’t on the critical path.

And we wouldn’t deliberately carry unnecessary state forever.

Agentic compute deserves similar discipline.

I now choose the resource appropriate to the task, allocate reasoning deliberately, pay for latency when it matters, keep context relevant, and correct bad trajectories before they compound.

I’m not trying to minimize Astra usage.

I’m trying to allocate it intentionally.

The goal isn’t to make Astra think less.

The goal is to spend its thinking where it creates the most useful work.

Publication note. Product behavior, plan limits, model availability, and usage rates can change. Product-specific details in this article were checked against OpenAI documentation on October 2, 2026. Historical usage observations are from my own sessions and should not be interpreted as OpenAI billing measurements.