Skip to content
Menu

Comparison

GPT-6 Astra vs Claude Fable 5.1: What the Demos Actually Show

Compare GPT-6 Astra and Claude Fable 5.1 through eight creator builds, from playable shooters to a Final Cut workflow, and two lab-reported benchmarks. The demos reveal useful ways to work with each model, but their different build conditions do not establish an overall winner.

Watch the comparison

Riley Brown’s Claude Fable 5.1 and GPT-6 Astra shooter demos shown side by side in Reindent’s video.

Frame from Reindent’s video · 9 min 18 sec
Watch the comparison on YouTube

Riley Brown’s two shooters give this comparison a concrete starting point. Both can be played. Players move through environments, aim, take damage and return for another round. But sharing a builder and a genre does not make the two projects a controlled test.

Reindent’s September 6 video examines those games alongside six other creator demonstrations. Together, they show several ways to build with GPT-6 Astra and Claude Fable 5.1: generate a new experience, improve an existing project, turn repository information into an interface, or prepare work inside another application. They also show why a polished clip is only part of the evidence.

The video contains the gameplay, app interactions and creator footage discussed below. This article expands the comparison into a guide to reading those demonstrations. Build times, prompt counts and iteration claims are the creators’ reports; Reindent has not independently reproduced the builds.

Game development: two shooters, different build conditions

In Brown’s Fable 5.1 shooter, the ammunition, health and respawn sequence help establish a playable loop. The important evidence is that the state of the game changes as someone plays it. Brown reported using three prompts, but that count does not supply the complete instructions, settings or development history.

His Astra shooter adds a useful view of the process. A coding conversation appears alongside the running game. In a follow-up post, Brown said he played for around two hours and requested changes between matches. The recording shows the setup and gameplay, rather than every change being implemented.

That distinction changes the comparison. Three prompts and two hours of playtesting describe different aspects of two different runs. Neither is a complete measure of effort or cost. We can compare the environments, interface readability and visible play. We cannot turn those preferences into a reliable ranking of the models’ engineering ability.

The practical lesson is to include the feedback loop in the story. Playing, noticing a problem and requesting a revision are part of making a game, even when an AI model writes the code.

Apps, simulations and editing: six more creator demos

SWARM FLOOR: does the scene respond to its controls?

Ryan Sael’s SWARM FLOOR, built with Fable 5.1, presents a warehouse floor with moving robots, task information and changing counters. Sael describes pathfinding, collision avoidance, heatmaps and adjustments to fleet size.

His prompt excerpt asks for a Three.js fulfillment floor based on a reference image, with specific facilities such as aisles, packing stations and charging pads. He reports a build time just under 24 minutes.

The interesting connection is between the visual scene and the simulation state. A beautiful warehouse is one result; a warehouse whose controls change its behavior is another. The demonstration offers something to inspect, while leaving real logistics reliability untested.

The iPod app: does the object serve a purpose?

Pietro Schirano asked Astra for a Mac app containing a 3D iPod built in Blender, using its familiar interface to browse Codex threads. The demonstration shows menu and thread navigation. Schirano reports 15 minutes and says the click sounds were made in code.

The iPod is an interface with a job to do. Its value as a demo comes from the relationship between the modeled object, its interactions and the information it exposes. The reported time describes this creator’s run, not a delivery estimate for another project.

The Long Silence: what existed before the model started?

Anshu Chimala’s space-game demonstration is an upgrade to an existing game. The earlier project was built with Claude Opus 5; Fable 5.1 is credited with improving its appearance. The creator’s before-and-after edit leads into gameplay showing changes to surfaces, foliage and lighting.

The public repository for The Long Silence gives readers another source to inspect, including its use of Blender and image-material tooling. That does not make the project a controlled benchmark, but it provides more context than a clip alone.

Credit belongs to both stages: Opus 5 for the earlier game and Fable 5.1 for the reported upgrade. For someone with a working application, improving an existing project may also be the more relevant task to evaluate.

Microduck: can an interface reflect an existing project?

Dilum Sanjaya’s Microduck app shows a robot changing poses and colors beside its controls. Sanjaya says Fable 5.1 explored available open-source repositories and used what it learned to select interface features.

This is a demonstration of building on existing material. Fable did not invent the underlying robotics project, and the app footage does not establish that those controls drive a physical robot. The other-model demonstrations in the surrounding thread should not be attributed to this Fable build.

Van Gogh town: do the spaces connect?

Peter Gostev’s Van Gogh town turns paintings into a continuous Three.js environment. The walkthrough moves from a bedroom toward a cafe, making navigation between spaces part of the result.

Gostev identifies the run as GPT-6 Astra at “Max,” describes six connected paintings and reports zero iterations. Those are his descriptions of the run. They do not establish its total cost, elapsed time or likelihood of success on another attempt.

The useful question is spatial: do the scenes fit together into a place someone can explore? Attractive individual views would answer a smaller question than the one posed by his continuous-world prompt.

Final Cut: which steps can an assistant take on?

Ben Davis describes asking Astra to import recorded clips, prepare their color and synchronize them in Final Cut. The footage shows Davis and his editor reviewing the result. He also credits Astra with editing the reaction clip and its captions.

We do not have a complete recording of the agent session, so we cannot verify each preparation step. The narrower opportunity is still useful: an assistant handles parts of the setup, while a person reviews the result and decides how the finished video should work. The demonstration does not establish that an editor can be replaced.

GPT-6 Astra vs Fable 5.1 benchmarks

The model names in the video match the primary announcements: OpenAI’s GPT-6 Astra and Anthropic’s Claude Fable 5.1. The video includes two selected lab-reported results:

Selected lab-reported benchmark results
EvaluationGPT-6 AstraClaude Fable 5.1
Terminal-Bench 4.057.9%55.8%
Humanity’s Last Exam, with tools57.2%65.0%

These values appear in OpenAI’s comparison table, checked again for this article on September 7, 2026. The lead changes between the two evaluations. The test conditions also matter: OpenAI describes results at the maximum of any effort setting in research or API environments, while Anthropic describes evaluation with production safeguards.

Those qualifications belong beside the numbers. The setups are not established as matched, and Reindent did not rerun the evaluations. Neither this small selection of scores nor the creator demos establishes an overall engineering winner.

How to compare the models for your own project

The examples suggest four questions to ask before treating an impressive clip as evidence for a model choice:

  1. What was the starting point? A blank project, an existing game and an existing editor require different kinds of work.
  2. What can someone actually do with the result? Look for interaction, changing state and a complete task, alongside visual quality.
  3. What did the person contribute? References, repositories, playtesting and revision requests help explain the result.
  4. What remains outside the record? Missing prompts, budgets, failed attempts and long-term behavior limit what can be concluded.

In the video, Diego says GPT-6 was working better in his own workflow at that point. That is a personal observation, not the result of a controlled engineering comparison.

A useful next experiment would put both models on a task you understand, with the same starting files and acceptance criteria, and record the help each one needs. These demos can suggest what to try. Your own results will tell you more about the work you need to deliver.

For more of Reindent’s work, explore our tools and products. If you have an AI-built project that needs help reaching production, see how Reindent works with builders.

Sources and scope

This article adapts Reindent’s September 6, 2026 video and the original creator posts linked throughout. Creator claims were retained from the video’s archived source research. Some X posts could not be reopened during this article’s review, so their build-time and process claims remain attributed reports. The official model pages and The Long Silence repository were accessible during the September 7 review.

The projects were not independently rebuilt, the hosted demos were not freshly playtested for this article, and their results should not be read as current product guarantees.