GPT-6 Astra vs Claude, Gemini, Qwen and DeepSeek: Who Wins the Pelican Bicycle Test?

Lauren Pan is the founder of ZimaSpace and the architect behind the acclaimed ZimaBoard series. Blending industrial design with embedded engineering, Lauren launched ZimaSpace with a clear mission: to democratize personal cloud computing. He operates on the belief that hardware should be both "hackable" and beautiful—closing the divide between industrial-grade servers and consumer gadgets. Today, he leads the engineering team in building tools that give creators full control over their digital lives.

GPT-6 Astra produces one of the most convincing pelicans yet, but the Pelican Bicycle Test becomes more interesting when it is compared with Claude Fable 5.1, Gemini 3.8 Flash, Qwen 3.8, DeepSeek V4, Kimi K3, and GLM-5.3-Flash. A one-line request to generate an SVG of a pelican riding a bicycle forces a language model to translate words into code, geometry, anatomy, object relationships, and a scene whose mistakes are immediately visible.

The latest results show that there is no simple winner. GPT-6 Astra raises the quality floor, Claude can spend dramatically more reasoning to repair geometry, Gemini adds exceptional visual polish, Qwen shows how close local models can get, and DeepSeek demonstrates how much reasoning settings can change one model's output. Tool-assisted GLM and DeepSeek animations reveal something different again: once an agent can inspect and iterate on its own artifact, the test stops measuring only the model.

Pelican Bicycle Test Results at a Glance

Model Test Type What Stood Out Main Caveat
GPT-6 Astra Standard SVG, Low to Max Very strong baseline; Low already outperformed the previous GPT-5.6 family visually Below Max, leg placement around the bicycle frame was still inconsistent
Claude Fable 5.1 Standard SVG + separate animation pass Max produced a highly deliberate composition with strong rider-to-bike interaction Max took almost 14 minutes and cost dramatically more than Low
Gemini 3.8 Flash Standard SVG, Low to High Exceptional visual polish and decorative coherence Visual flair is not the same thing as physical correctness
Qwen 3.8 27B Local standard SVG One of the strongest Pelicans generated by a roughly 17GB local model Default xhigh reasoning consumed more than 22,000 reasoning tokens
Qwen 3.8 Flash-Next Local standard SVG Strong scene from a 125B MoE with roughly 6B active parameters Local result depends heavily on quantization and available hardware
Qwen 3.8 Max Standard SVG / larger model Clean rider-to-bike interaction and polished composition Much larger model, so it is not a fair local-hardware comparison
Kimi K3 Standard SVG Coherent result from an extremely short prompt Used more than 13,000 reasoning tokens
DeepSeek V4 Pro 0813 Standard SVG, Low to High Reasoning levels produced radically different visual strategies High variance makes one screenshot a poor model summary
DeepSeek V4 Flash 0731 Standard SVG High reasoning repaired a badly broken default bicycle Reasoning improvement was dramatic but not guaranteed
GLM-5.3-Flash Agent-assisted animated HTML/SVG Rich animation, pedals, scene controls and iterative visual work Chrome MCP and an agent environment make it non-comparable with one-shot tests

This is not a scientific leaderboard. The public runs were performed at different times, with different APIs, reasoning settings, quantizations, and—in some cases—different tool environments.

The more defensible comparison is behavioral. Does the bicycle have a coherent frame? Do the feet reach the pedals? Does the bird actually interact with the handlebars? Does extra reasoning correct those relationships, or merely add more decoration?

What Does the Pelican Bicycle Test Actually Measure?

The original prompt is deliberately minimal:

Generate an SVG of a pelican riding a bicycle.

Producing a convincing answer requires several capabilities to work at once. The model must write valid SVG, construct a recognizable bicycle, approximate a pelican's anatomy, place the bird in the correct part of the scene, and make separate objects interact in a visually plausible way.

That makes the test useful for observing:

  • SVG and frontend code generation,
  • object geometry,
  • spatial composition,
  • anatomical consistency,
  • object-to-object interaction,
  • instruction following,
  • and reasoning-effort efficiency.

It does not directly measure factual accuracy, scientific reasoning, repository-scale coding, agent reliability, professional knowledge, or general intelligence.

Simon Willison, who has repeatedly used the prompt across model generations, has also become more cautious about it. He now treats it primarily as a quick behavioral test and a useful way to compare releases inside the same model family rather than a universal intelligence benchmark.

His Kimi K3 Pelican analysis explicitly warns against treating individual Pelican outputs as a serious model leaderboard.

That limitation is important because modern models are now good enough that visual taste increasingly influences the result. Once every output contains a recognizable bird and bicycle, a beautiful sunset or fish basket can make one image feel smarter even when another model has more accurate geometry.

How Did GPT-6 Astra Perform on the Pelican Bicycle Test?

GPT-6 Astra's most important result is not simply that Max looks impressive. It is that Low already produces a much stronger Pelican than the previous GPT-5.6 family.

GPT-6 Astra and GPT-5.6 Pelican Bicycle Test results across multiple reasoning levels
GPT-6 Astra compared with GPT-5.6 Sol, Terra and Luna across multiple reasoning levels. Source: Simon Willison's Astra comparison.

Willison tested Astra at Low, Medium, High, XHigh and Max reasoning. His recorded output token counts increased from 1,906 at Low to 12,638 at Max.

Astra Reasoning Output Tokens Recorded Cost
Low 1,906 $0.0955
Medium 2,560 $0.1282
High 3,671 $0.1837
XHigh 6,766 $0.3385
Max 12,638 $0.6321

Willison's strongest observation was that Astra Low looked better than any GPT-5.6 Sol Pelican he had generated at any reasoning level. That makes this less about Max reasoning and more about a generational increase in baseline visual coding ability.

The bicycle becomes more coherent, the pelican becomes recognizably integrated with the vehicle, and the whole scene requires less interpretation from the viewer.

But the task is still not solved perfectly. Below Max, Astra did not reliably position the two legs on opposite sides of the bicycle frame.

That is exactly why this silly benchmark remains useful. A polished output can give an immediate impression of competence, while a small relationship such as leg placement exposes whether the scene is structurally consistent.

Astra appears to understand the composition better, but the Pelican Test cannot establish that it maintains a human-like physical model of bicycle riding.

GPT-6 Astra Also Appeared in OpenAI's Developer Demo

OpenAI's GPT-6 Astra developer presentation contains another Pelican reference while demonstrating richer visual and 3D generation. This is not the same controlled SVG run, so it should be treated as supplemental evidence rather than part of the comparison.

Supplemental GPT-6 Astra visual example from OpenAI's developer presentation. It should not be compared directly with the controlled SVG grid.

Claude Fable 5.1: Does More Thinking Produce a Better Pelican?

Claude Fable 5.1 demonstrates that extra reasoning can improve geometric relationships—but it also shows how quickly the cost of a simple visual task can explode.

Claude Fable 5.1 Max Pelican Bicycle Test result
Claude Fable 5.1 Max result. Source: Simon Willison's Fable 5.1 test.

Low and Medium completed in roughly 23 seconds and cost about ten cents each. High took about 29.6 seconds.

Then the curve changed dramatically.

Reasoning Time Recorded Cost
Low 23.8 seconds ~$0.10
Medium 23 seconds ~$0.10
High 29.6 seconds ~$0.13
XHigh 7 minutes 51 seconds $1.83
Max 13 minutes 54 seconds $3.30

The Max result does improve details that matter. The two legs are visibly placed around the bicycle frame, the feet reach the pedals, a wing reaches the handlebars, and the bicycle geometry receives far more deliberate attention.

The reasoning trace is particularly revealing because Fable actively reconsidered parts of the drawing. It noticed problems in the front fork geometry, reconsidered helmet placement around the beak, and repeatedly checked whether decorative elements would collide with the rest of the scene.

That gives us a rare visible connection between more reasoning and a more internally consistent artifact.

But the cost difference is enormous.

The useful question is not whether Max is better. It is whether a nearly 14-minute Pelican is valuable enough to justify more than 30 times the cost of Low.

Claude's Pelican Was Then Animated in a Second Pass

The public animation was not created by the original Pelican prompt. Willison took the Max SVG and sent it back to Fable 5.1 with a separate instruction to animate the existing artifact.

That second pass used 6,121 input tokens and 26,201 output tokens and added another recorded cost of about $1.37.

Because the animation was a second task, it should not be ranked directly against static one-shot results from other models. It is still useful for showing what happens when the model must preserve geometry while introducing motion.

Watch the original Claude Fable 5.1 animated Pelican.

Gemini 3.8 Flash: Does Visual Flair Equal Better Understanding?

Gemini 3.8 Flash stands out for a different reason: it makes the benchmark look like illustration work instead of a geometry exercise.

Gemini 3.8 Flash High reasoning Pelican Bicycle Test beside the sea
Gemini 3.8 Flash at High reasoning. Source: Simon Willison's September 2026 Gemini test.

The High result includes a teal cruiser bicycle, a red polka-dot scarf, a fish basket, a beach boardwalk and a glowing seaside background.

It is immediately attractive.

That makes Gemini a useful example of a problem with visual benchmarks: humans naturally reward style.

A more polished composition can feel more intelligent before we inspect whether the bicycle frame closes correctly, whether the foot truly reaches the pedal, or whether the pelican's body is positioned consistently relative to the saddle.

Visual polish is a real capability, but visual polish and physical consistency are different capabilities.

As models improve, the Pelican Test increasingly measures both. That makes simple winner rankings less meaningful than inspecting the specific ways each model succeeds or fails.

Qwen 3.8 27B: Can a 17GB Local Model Compete With Frontier Pelicans?

Qwen 3.8 27B may be the most important result for local AI users because this was not a massive cloud-only model.

Willison ran a roughly 17GB Q4_K_M quantized version locally through LM Studio, including experiments on high-memory local hardware.

Qwen 3.8 27B locally generated Pelican Bicycle SVG at xhigh reasoning
Qwen 3.8 27B running locally with xhigh reasoning. Source: Simon Willison's local Qwen 3.8 27B test.

The result is unusually coherent for a model that can be packaged into a roughly 17GB quantized file.

The bicycle has the expected frame shape. The two legs appear on opposite sides of the bicycle, which is a surprisingly difficult relationship for many models. The pelican has a clear pouch, its wing reaches the handlebars, and the motion lines are placed behind rather than cutting through the scene.

Willison described it as the best Pelican SVG he had generated locally at that point.

But there is a major catch.

Qwen 3.8 27B massively overthought the task.

At its default xhigh reasoning level, the model took around 21 minutes. It used 22,276 reasoning tokens before producing 3,223 final output tokens.

When reasoning was disabled, the generation dropped to a little over two minutes—but the output became noticeably worse.

Qwen 3.8 27B Pelican Bicycle result with reasoning disabled
The same Qwen 3.8 27B prompt with reasoning disabled. Bicycle geometry and rider interaction deteriorated substantially.

The frame became weaker, the feet missed the pedals, and the model no longer made a serious attempt to connect the wing to the handlebars.

This creates almost the opposite story from GPT-6 Astra.

Astra's most impressive characteristic is that Low is already strong. Qwen 3.8 27B demonstrates how much visual quality can be extracted from comparatively compact local hardware if the model is allowed to spend a large amount of time reasoning.

It is not evidence that a 17GB model has caught GPT-6 overall. It is evidence that “small enough to run locally” and “capable of producing sophisticated structured output” are no longer mutually exclusive.

Qwen 3.8 Flash-Next: What Can a 6B-Active MoE Draw?

Qwen 3.8 Flash-Next provides a second local-AI data point from a very different architecture.

The model has 125 billion total parameters but activates roughly 6 billion parameters for each token. That distinction can reduce compute requirements even though the full weight set remains much larger than 6B.

Qwen 3.8 Flash-Next SVG showing a pelican riding a bicycle in a landscape
Qwen 3.8 Flash-Next using a quantized UD-Q2_K_XL build on DGX Spark. Source: Simon Willison's Qwen 3.8 Flash-Next testing.

The xhigh result is a polished scene. The pelican sits over a red bicycle, a fish occupies the front basket, and the bicycle itself maintains a recognizable structure rather than collapsing into disconnected arcs and tubes.

What makes this result interesting is not whether it is prettier than Gemini or Astra.

It is that sparse MoE models are making simple parameter-count comparisons less useful.

A 125B model with roughly 6B active parameters does not behave like a conventional dense 6B model, nor does it have the same storage and memory characteristics as one.

The Pelican demonstrates the capability side of that equation: relatively low active compute can still coordinate a surprisingly sophisticated structured scene.

Qwen 3.8 Max: What Changes When the Model Gets Much Larger?

Qwen 3.8 Max provides a useful upper-end comparison within the same broader model family.

Qwen 3.8 Max Pelican Bicycle Test result
Qwen 3.8 Max Pelican result. Source: Simon Willison's Qwen 3.8 Max example.

The result is clean and easy to parse. The legs reach the pedal region, the bicycle has a conventional form, and the overall scene stays simple enough that the geometry remains readable.

But Qwen Max reinforces another lesson from the test:

A larger model should not automatically receive a higher score because it is larger.

For practical local AI, Qwen 3.8 27B may actually be the more interesting Pelican because it tells us something about what can run on hardware a sophisticated home user could realistically own.

Max demonstrates a model ceiling. The local 27B result demonstrates deployment progress.

Kimi K3: How Much Reasoning Does One Pelican Need?

Kimi K3 produced another coherent Pelican, but its hidden compute use may be more interesting than the picture.

Kimi K3 Pelican Bicycle Test result showing a pelican riding a red bicycle
Kimi K3 result from the standard Pelican prompt. Source: Simon Willison's Kimi K3 test.

The prompt was only a few words long, yet the recorded run used 95 input tokens and 16,658 output tokens.

Of those output tokens, 13,241 were reasoning tokens.

The recorded cost was about $0.25 for that run.

The final scene is competent: recognizable bicycle, recognizable pelican, reasonable rider placement, background details, road and motion elements.

But the numbers reveal something that the image cannot.

A visually simple result can hide an enormous amount of model-side work.

That matters for real applications. A model can look inexpensive when judged by a single successful result but become costly or slow when the same reasoning pattern is repeated hundreds of times inside an agent workflow.

For Kimi, the Pelican becomes a reasoning-efficiency test almost as much as a drawing test.

DeepSeek V4 Pro: Why Do Different Reasoning Levels Look Like Different Models?

DeepSeek V4 Pro 0813 produced one of the strangest reasoning comparisons in the archive.

Low, Medium and High did not simply look like increasingly polished versions of the same design. They appeared to pursue different visual strategies.

DeepSeek V4 Pro Low reasoning Pelican Bicycle Test
Low
DeepSeek V4 Pro Medium reasoning Pelican Bicycle Test
Medium
DeepSeek V4 Pro High reasoning Pelican Bicycle Test
High

DeepSeek V4 Pro 0813 across Low, Medium and High reasoning. Source: Simon Willison's DeepSeek V4 Pro test.

Low is relatively clean and minimalist. Medium becomes dramatically more abstract: broken wheel arcs, loose line work and a strange elongated shape dominate the composition. High changes direction again, returning to a more conventional red bicycle and adding a basket, pennant and musical notes.

The lesson is not simply “High is better.”

Reasoning effort appears to influence which visual strategy the model chooses, not just how carefully it executes one fixed composition.

This makes DeepSeek V4 Pro an excellent reminder that one screenshot is a very weak basis for claiming that a model is good or bad at visual coding.

Sampling variance matters. Reasoning settings matter. The API or harness may matter. The Pelican makes those differences visible because we can inspect the result immediately.

DeepSeek V4 Flash: Can More Reasoning Repair a Broken Bicycle?

DeepSeek V4 Flash 0731 provides an even cleaner example of reasoning affecting geometry.

The default run produced a severely malformed bicycle. Wheels appeared as incomplete arcs, frame tubes floated without connecting, and the pelican hovered over the scene rather than convincingly riding it.

DeepSeek V4 Flash default reasoning Pelican Bicycle result with malformed bicycle geometry
DeepSeek V4 Flash 0731 at its default reasoning setting.

Willison then increased reasoning effort to High using the same basic task.

The result changed dramatically.

DeepSeek V4 Flash High reasoning Pelican Bicycle result with coherent bicycle geometry
DeepSeek V4 Flash 0731 at High reasoning. Source: Simon Willison's DeepSeek V4 Flash reasoning test.

The High output has complete wheels, a recognizable frame, a crank area, a pelican gripping the handlebars, and a foot positioned near the pedal.

This does not prove that more reasoning always fixes visual generation.

It does suggest that some bad SVG results are coordination failures rather than evidence that the model has no representation of bicycle geometry at all.

The pieces may already exist inside the model; extra reasoning can sometimes help assemble them correctly.

Why DeepSeek V4 Flash Once Beat V4 Pro on the Same Pelican Prompt

An earlier DeepSeek V4 comparison produced another counterintuitive result.

DeepSeek V4 Flash Pelican Bicycle result with coherent bicycle geometry
DeepSeek V4 Flash
DeepSeek V4 Pro Pelican Bicycle result with unusual pelican anatomy
DeepSeek V4 Pro

Source: Simon Willison's DeepSeek V4 comparison.

The Flash version had an excellent bicycle by Pelican-Test standards: a recognizable frame, visible chain, a reflector, wings reaching the handlebars, and feet on the pedals.

The Pro model retained a reasonable bicycle but generated a much stranger pelican with an oversized body and inconsistent anatomy.

This should not be interpreted as evidence that Flash is generally smarter than Pro.

It demonstrates something narrower and more useful:

Model size and benchmark reputation do not guarantee the best output from one stochastic visual coding prompt.

GLM-5.3-Flash vs DeepSeek V4 Flash: What Happens When the Models Get Tools?

GLM-5.3-Flash introduces a different kind of Pelican experiment.

A community comparison used an agent environment rather than the simple one-line test. Both GLM-5.3-Flash and an experimental DeepSeek V4 Flash vision setup were run with OpenCode, Trellis, Chrome MCP and Max reasoning.

The instruction asked the models to create a complete HTML page containing a 2D SVG animation and explicitly told them to treat the task as a competition.

That changes what is being measured.

The model is no longer only producing one SVG from one prompt. It can spend more time, use tools, inspect a browser environment and behave more like a coding agent.

GLM-5.3-Flash Animated Pelican

GLM-5.3-Flash animation produced using OpenCode, Trellis, Chrome MCP and Max reasoning. Source: original community comparison.

The GLM result is substantially more ambitious than the static Pelicans above. It contains a complete scene, explicit bicycle components, motion, pedals and a richer page-level presentation.

Community commenters generally preferred the GLM result, with several pointing to the more complete pedal and animation work.

There are still visible mistakes. Some foot motion appears mechanically unnatural, and adding more animation introduces more opportunities for relationships to break.

This creates a useful distinction:

A more feature-rich result is not automatically a more physically correct result.

DeepSeek V4 Flash Vision Animated Pelican

DeepSeek V4 Flash vision experiment generated under the same tool-assisted environment. Source: community test methodology and discussion.

The DeepSeek version takes a different approach, using a sunset composition and a cleaner responsive presentation.

Community feedback generally favored GLM's mechanical completeness while giving DeepSeek credit for visual composition and mobile presentation.

More importantly, both results demonstrate how quickly the meaning of an AI benchmark changes once tools are introduced.

The system being tested is now:

  • the base model,
  • reasoning configuration,
  • agent harness,
  • browser tooling,
  • visual feedback,
  • execution time,
  • and iteration strategy.

Calling the result simply “GLM vs DeepSeek” hides much of what actually produced the artifact.

Why One-Shot Pelicans and Agent-Assisted Animations Should Not Share One Ranking

This distinction is important enough to make explicit.

Test What It Mostly Measures
One-line SVG prompt Model reasoning, SVG generation and spatial coordination in one response
Higher reasoning effort Whether additional inference helps repair geometry and relationships
Second-pass animation Whether a model can preserve an existing composition while introducing motion
Agent + Chrome MCP Model, harness, tools, feedback loop and iterative coding together

A strong agent-assisted animation is arguably more relevant to modern coding workflows than a one-shot SVG.

But it is a different test.

If GPT-6 Astra gets one response while GLM receives twenty minutes, a browser and visual feedback, we cannot responsibly claim that the final animation proves GLM is better at the original benchmark.

What it proves is that models become substantially more capable when they are placed inside systems that can inspect, execute and revise their own work.

Which AI Model Actually Won the Pelican Bicycle Test?

There is no scientifically defensible single winner, but the current results do reveal several category standouts.

Category Standout Why
Strongest baseline improvement GPT-6 Astra Low reasoning already produces a major improvement over the GPT-5.6 family
Most deliberate high-effort geometry Claude Fable 5.1 Max Explicit reasoning repaired rider, fork and object relationships
Strongest visual flair Gemini 3.8 Flash Turns the minimal prompt into a highly designed illustration
Most impressive local result Qwen 3.8 27B A roughly 17GB quantized local model produces unusually coherent geometry
Most interesting sparse-MoE result Qwen 3.8 Flash-Next Strong composition despite roughly 6B active parameters per token
Most extreme reasoning case Kimi K3 More than 13,000 reasoning tokens for a tiny prompt
Best reasoning-repair example DeepSeek V4 Flash High reasoning dramatically improves a broken default bicycle
Highest output variance DeepSeek V4 Pro Low, Medium and High produce radically different visual strategies
Most ambitious tool-assisted animation GLM-5.3-Flash Builds a richer animated HTML artifact when given an agent environment and browser tools

If the question is simply which current model produces the strongest one-shot Pelican with minimal reasoning, GPT-6 Astra is difficult to ignore.

If the question is which result is most surprising for hardware that can actually sit on a desk and run locally, Qwen 3.8 27B becomes much more important.

If the test shifts from one-shot generation toward a coding-agent workflow with tools and iteration, GLM-5.3-Flash demonstrates how different the ceiling can become.

Those are three different questions, which is exactly why one numerical winner would hide more than it explains.

Do These AI Models Actually Understand What They Draw?

The Pelican Bicycle Test cannot prove genuine physical understanding.

A model may generate a correct-looking bicycle because its training has given it powerful representations of bicycles, birds, SVG code and common visual compositions.

Successfully coordinating those representations is evidence of useful spatial competence.

It does not establish that the model understands balance, gravity, mechanical load, pedal motion or locomotion in the same way a human does.

The remaining failures make this particularly clear.

Astra can create an attractive bicycle while still placing both legs incorrectly around the frame. Claude's animated wheels can reveal motion issues that were invisible in the static SVG. Qwen can spend 22,000 reasoning tokens building a scene that another model produces much faster. DeepSeek can produce a broken bicycle at one reasoning level and a coherent one at another.

These models are becoming excellent at constructing visual structures from language.

Whether that should be described as “understanding” depends on what level of internal representation we require from the word.

Why the Chinese Open Models Are the Most Interesting Part of This Test

The most important change since earlier versions of the Pelican Test may not be that GPT-6 produces a better bird.

It is that open and locally deployable models are now producing results that would have looked frontier-class surprisingly recently.

Qwen 3.8 27B can generate a coherent Pelican locally from a roughly 17GB quantized model. Qwen 3.8 Flash-Next combines a large total MoE capacity with much lower active compute. DeepSeek V4 Flash can recover from a broken visual result when additional reasoning is enabled. GLM-5.3-Flash can operate inside a browser-equipped coding-agent loop and produce a sophisticated animated artifact.

None of that means they have universally caught GPT-6 Astra or Claude Fable 5.1.

It means the gap is becoming workload-specific.

For the hardest reasoning problems, frontier cloud models can still have a clear advantage.

For structured generation, coding, private workflows and increasingly sophisticated local agents, the question is becoming less obvious.

The Pelican Test accidentally visualizes the same trend happening across local AI: the frontier keeps moving, but the level of capability available outside the frontier is moving almost as fast.

What Should We Learn From the Pelicans?

The clearest result is not that one AI has finally learned how a pelican should ride a bicycle.

It is that the baseline quality of code-generated visual artifacts is rising quickly while the differences between models are becoming more subtle.

GPT-6 Astra shows that strong structured visual generation can now appear even at Low reasoning. Claude Fable 5.1 demonstrates how much extra compute can be spent repairing geometry. Gemini 3.8 Flash shows how strong aesthetic composition can influence human judgment. Qwen 3.8 27B shows what is possible from a compact local deployment. Kimi K3 exposes the hidden cost of heavy reasoning, while DeepSeek demonstrates how unstable individual visual generations can remain.

GLM-5.3-Flash adds the final piece: once a model receives browser tools, visual feedback and permission to iterate, the benchmark starts measuring an AI system rather than an isolated model.

That makes the Pelican Bicycle Test less useful as a leaderboard and more useful as a microscope.

It exposes the difference between:

  • a model that can write valid SVG,
  • a model that can maintain spatial relationships,
  • a model that can spend more reasoning to repair mistakes,
  • a local model that can approach frontier-looking output,
  • and an agent system that can inspect and improve its own artifact.

GPT-6 Astra may currently produce one of the strongest Pelicans, but the bigger story is that Claude, Gemini, Qwen, DeepSeek, Kimi and GLM now fail—and succeed—in increasingly sophisticated ways.

The bird is still ridiculous. The benchmark is imperfect. But as a visual snapshot of how quickly model behavior is changing, the Pelican Bicycle Test remains surprisingly revealing.

Tech & AI HUB

More to Read

Get More Builds Like This

Stay in the Loop

Get updates from Zima - new products, exclusive deals, and real builds from the community.

Stay in the Loop preferences

We respect your inbox. Unsubscribe anytime.