- Everyday Work: Matt Shumer, a CEO of OthersideAI, found the GPT-6 Astra model strong at coding and computer work, while ambitious builds needed human direction.
- Manager Loop: His two-session pattern separated planning from implementation, helping long projects continue after single agents became stuck on details.
- Human Role: Shumer still set goals, supplied assets, reviewed mockups, changed delegation limits, and intervened; several showcase projects remained unfinished.
- Cost Boundary: Five Macs and a cloud machine supported the tests, but missing usage records leave the tests’ token charges and total project cost unknown.
In a detailed account published September 3, 2026, Matt Shumer, founder and CEO of AI software company OthersideAI, described using OpenAI’s GPT-6 Astra model for backend repairs, browser and desktop work, and sprawling multi-agent builds. He found Astra unusually practical for everyday engineering and long conversations. On ambitious projects, however, he found it slow and often absorbed in low-value detail. His largest demonstrations depended on substantial human setup, existing assets, and a resource-heavy coordination pattern.
Everyday Work Required Less Reorientation
Shumer’s clearest positive result came from ordinary software work. In one example, a friend sent a phone request about a broken backend service and reported it working again about an hour later. Shumer said he did not intervene. The account does not identify the service, diagnosis, changes, or verification record, so the episode remains one reported outcome rather than a reliability measure.
His other examples were less dramatic but more representative of daily computer use. Astra worked through browser and desktop interfaces to help prepare newsletters, manage an inbox, and handle advertising tasks. A colleague or agency reviewed some of that work, according to Shumer. He also describes a disk-space tool built from two top-level prompts that moved old message threads to cloud storage and restored them on demand. The number counts only Shumer’s top-level requests; the exchanges, tool calls, and retries underneath them are unknown.
Shumer says Astra gives concise, plain-language progress reports and retained useful context across long conversations. When the system condenses earlier exchanges to make room for more work, it sometimes looses details, but he finds the remaining continuity good enough to avoid repeatedly explaining the project. His own writing style is a notable exception: Shumer says the model absorbes preferences less reliably there than in other work.
For his testing, Shumer ran Astra inside Codex, OpenAI’s environment for software agents to work with code and tools. He said he usually chose Medium as reasoning effort for daily tasks and Ultra for larger experiments.
Manager Loop Splits Direction From Execution
Shumer tried long single-agent runs, isolated lanes, periodic supervision, and several multi-agent arrangements. The systems often kept working but became absorbed in low-value details, leaving large parts of a project untouched. Persistence did not guarantee useful direction.
His Manager Loop experiments separated those jobs. Shumer first established the vision and answered questions about the goal. A coordinator session then divided the agreed project into phases and sent one phase at a time to a separate implementer session. The implementer executed the phase, checked its work, reported back, and could call scoped sub-agents. The coordinator reviewed each completion report before advancing the plan.
The coordinator and implementer were peer sessions, not a manager agent with the implementer hidden beneath it. Optional helpers sat below the implementer. That distinction kept strategic direction in a context separate from day-to-day execution, reducing the chance that implementation detail would consume the entire plan.
The arrangement automated an intervention Shumer had been making himself: telling the working agent to continue to the next phase. He reports that projects moved farther after the split, with less ongoing steering. The pattern builds on Codex’s existing support for parallel agent work. Its important change assigned direction and implementation to different conversations.
Shumer still chose the goals, supplied inspiration, reviewed interface mockups, configured permissions and delegation, encouraged greater ambition, and judged the results. He also intervened in the largest city project. Manager Loop reduced one kind of repetitive supervision; it did not transfer product judgment or accountability to the model.
Big Demos Needed Assets, Hardware, and Patience
The most striking demonstration was a simulated built of the game civilization in Unreal Engine. Shumer had already experimented with the coordination pattern and supplied tools and existing assets, including MetaHuman characters. Within that environment, he reports that Astra produced a world with animals, model-driven inhabitants, and audible dialogue.
My first “holy shit” moment with GPT-6 Astra:
I asked it to create a world in Unreal Engine, and fill it with humans (each an Astra-powered agent) who all have to work together to survive.
A day later, I was in my bedroom and heard voices coming from the living room… I… pic.twitter.com/INzt3t2bp1
— Matt Shumer (@mattshumer_) September 3, 2026
Those conditions help explain why the demonstration looks stronger than his direct attempts in Three.js or Blender. Existing game-engine assets reduced how much visual material Astra had to create from scratch. Shumer says he preferred Claude Fable for visual taste and asset creation, and Astra also failed to produce a website redesign he liked.
His larger New York project made the limit more visible. Astra refined an initial street in detail, but much of the city remained to be built and the project was still in progress when Shumer wrote. A browser project had earlier reached a recognizable first version and then plateaued with substantial work left. Both were colorful but partial outputs.
GPT-6 Astra built this Manhattan world in Unreal Engine over the course of a week.
It was literally able to go street by street to make each one perfect. pic.twitter.com/7VTol9QfHq
— Matt Shumer (@mattshumer_) September 3, 2026
The human operator also carried the physical and organizational burden. Shumer described five Macs and a cloud machine in use across his experiments, along with memory and disk pressure. Existing assets, careful setup, permissions, process monitoring, and acceptance decisions were part of the result, even when the top-level request looked short.
Hardware and Pricing Leave Project Cost Unknown
Shumer increased one computer’s configured sub-agent cap from four to 16 and set a cap as high as 96 on another machine.
Shumer’s machine inventory and reported memory and disk pressure make the setup operationally heavy. He describes his token consumption as enormous.
OpenAI’s pricing schedule reviewed on September 5 separates input, cached input, cache writes, and output, while Astra’s model documentation applies higher rates to requests that exceed 272,000 input tokens. Without Shumer’s token mix, project boundaries, tool charges, or invoice, those rates cannot reconstruct what the experiments cost. They do show why longer conversations and more concurrent agents turn coordination design into a budget decision.
Human Direction Remains Central
Artificial Analysis’ independent model comparison gives Astra High and Fable 5 the same aggregate intelligence index score of 53, while Astra produces output faster but takes longer to return its first token. The mixed result supports parity on that aggregate index alongside different latency profiles, rather than a single winner across tasks. Its settings and tasks differ from Shumer’s work, and it covers Fable 5 rather than every Fable 5.1 claim in his evolving account.
The comparison cannot supply a matched check of Shumer’s Manager Loop, backend repair, civilization, or city project. Astra appeared useful when he gave it familiar software work, usable tools, clear permissions, and a long-running context. Separating direction from execution may help such projects move beyond a detail-fixation plateau.


