The author used GPT 6.1 Soul and Claude Sonnet 5.5 to create interactive games and websites from prompts. In these runs, Claude, in the author's assessment, more often matched the task mechanics accurately, while GPT made a more interactive 3D vending machine.
How the comparison of GPT 6.1 Soul and Claude Sonnet 5.5 was conducted
The author tested the models on five tasks: a 3D drink vending machine, an insectarium, a 2D shooter, a pizzeria game, and a website. For generation, the author chose High reasoning mode rather than the maximum level. The author explained that this made the test closer to ordinary use: they assumed that the maximum mode might use up the limits before all the tasks were completed.
This comparison shows the results of specific runs, not fixed properties of the models. The author assessed the finished pages and games personally, and some token and limit readings could not be interpreted unambiguously. The test did not cover backend development.
What matters more when choosing a model: speed or accuracy in following a prompt? In this set of tasks, the criteria diverged: faster generation did not always produce a game that better matched the request.
3D vending machine: GPT won thanks to working buttons
In the GPT version, you could enter a notional amount, choose a drink, and start the purchase: the available bottle rolled out, and the vending machine's little door opened. If the drink could not be purchased, the action did not work.
In the Claude version, the author could not find any working interaction: the buttons did nothing when clicked. The appearance of the machine and bottles seemed roughly comparable to GPT's result. The author declared GPT the winner in this category for its interactivity, not for visual superiority.
Token usage readings for this task were contradictory: the interface showed 133 thousand tokens for GPT, another figure was 2,87 million, and the author cited 250 thousand for Claude. The author could not determine which figure to trust. Both interfaces showed roughly similar portions of the five-hour and weekly limits. Without knowing exactly what each figure measured, these numbers cannot be treated as an accurate comparison of usage.
This difference matters for prototyping: a polished image does not necessarily mean that the user can complete the main action. In this kind of test, it is worth checking the key workflow separately rather than relying on a first visual impression.
Insectarium: visual assessment and disputed token figures
In the insectarium, GPT placed butterflies, a stick insect, a praying mantis, and additional details around the scene. The author noted that some elements looked odd, particularly the praying mantis, but there were noticeably many objects. The butterflies and insect images in Claude's version seemed more interesting to the author; the author also noticed something strange about the praying mantis image there.
According to the interface readings, GPT used about 1,3 million tokens, while Claude used roughly 640 thousand. The author considered Claude more economical on this measure, but noted that the portions of the limits used did not differ all that noticeably. This is an observation based on the interface, not an independently verified measurement.
2D shooter: Claude seemed more dynamic
In GPT's game, the author noted changes in weapons and locations, as well as a tank and various enemies. The pace seemed slow, the enemies quickly became repetitive, and the boss appeared too early. To reach it, the player had to replay the basic sections.
In Claude's game, the author saw more intense combat: enemies descended by parachute, drones exploded when they fell, and a little vehicle and enemies throwing barrels appeared. The game included medkits and checkpoints. In the author's impression, it was more challenging and dynamic, but the author did not reach the boss: completing the game seemed to take too long.
This assessment depends on the player's preferences. A more challenging shooter with more events may seem more impressive, but a simple, short game can be more convenient for quickly demonstrating a mechanic. In the author's assessment, Claude had the edge in the test for intensity and variety.
Pizzeria game: Claude followed the prompt more closely
In the GPT version, the gameplay was set up as a clicker: the player chose ingredients and watched them cook. The author expected a more hands-on mechanic, with a larger table and the ability to add ingredients independently, so considered the result a poorer match for the prompt.
Claude's version let the player roll out the dough, add specified amounts of ingredients, and put the pizza in the oven. The player had to monitor the baking: if the pizza was taken out too soon, it remained raw, and if left too long, it burned. The author tested this scenario and noted that the pizza really did burn. Sound effects and animation complemented the mechanics.
For this task, the author rated Claude's result higher for matching the described game. In the author's assessment, GPT made a website instead of the expected game. The author considered Claude's version a good starting point for further development, but this does not confirm that the product is ready for publication or commercial use.
Websites: functionality versus clarity of presentation
On GPT's website, the author found an infographic and a color setting for the 3D object. In the author's assessment, the page adapted well to different screen sizes, but contained little information. The author allowed that they might not have found all the sections.
Claude's website included sound, a menu, a contact form, and an interactive presentation of materials. Its sections contained information about websites, pricing, and projects. However, the author did not fully understand the purpose of the page or how some interactive elements worked. Visual richness alone does not guarantee a clear user journey.
For a practical assessment of a website, it is useful to separate three questions: Does it display correctly on different screens? Do the interface elements work? Does the visitor understand what is being offered? GPT received a positive assessment from the author for its responsiveness and visual details, while Claude was praised for the number of interactive features. These observations do not provide an unambiguous conclusion that one website is superior by every criterion.
Time and resource usage: why the figures should be treated with caution
The author reported separate time measurements: in one case, GPT took 11 minutes to complete a task, while Claude took about 14β15 minutes. In comments about this comparison, the author also cited GPT's usage as about 1,8 million tokens. However, the author cautioned that they were confused about the limit percentages and were giving approximate figures. These numbers cannot be combined into a precise efficiency ranking.
How many tokens will a similar project require next time? This test cannot provide an exact answer: the interface readings were ambiguous, and the result depends on the task and the stages of generation. To compare the models yourself, it is better to run both with the same prompt and settings, record the time and usage from the same interface fields, and then manually check that the features work.
Conclusion of the GPT 6.1 Soul and Claude Sonnet 5.5 comparison
In the author's overall assessment, Claude won almost every category except the 3D vending machine. It followed the prompt better in the pizzeria game, and its shooter seemed more dynamic and varied to the author. Both models showed useful strengths when creating websites, but part of Claude's interface was not entirely clear, while GPT's website seemed sparse.
This conclusion reflects the author's opinion based on a limited set of tasks, not a universal ranking of the models. For quickly prototyping a specific mechanic, the test gives grounds to try Claude Sonnet 5.5. If working interactivity for a simple object is the priority, the vending machine test shows a strength of GPT 6.1 Soul. Before choosing a model for a real project, it is worth separately checking the code, feature availability, responsiveness, and how each workflow behaves.
Brief takeaway: in this comparison, Claude more often matched the game task accurately, GPT performed more convincingly on the vending machine, and the token and limit figures remained approximate.
Compare models before you start
The service sets its plans, limits and model catalog. If they differ from this article, contact us so we can update it and record a new review date.
Browse modelsAffiliate link: your price stays the same and the project earns a commission.