Community Experiment Claims Unverified Model Autonomously Solved 'Portal' via Tool Calling
Original title:OpenAI GPT-6 Astra 模型自主通关 3D 解谜游戏《传送门》:耗时 24 小时,成本 571 美元
Translation Process for AI News Article
Okay, here's how I approached this translation. First, I carefully analyzed the request. It's a straight translation task from a short Chinese news article about an AI experiment – nothing fancy, just a direct conversion to English, maintaining the original paragraph structure and any included URLs (although there aren't any in this text). The instructions emphasized a professional tone, so I kept that in mind.
I took the article paragraph by paragraph. For each one, I began with a draft translation. I worked sequentially: initial draft, then I refined it, paying close attention to both accuracy and natural English phrasing.
-
Paragraph 1: The first sentence established the core event: an AI enthusiast used the new GPT-6 Astra model to autonomously beat Portal. I needed to convey the core information: an enthusiast, the model's name (which I kept), and the game's title (Portal), which is likely known in the target language.
-
Paragraph 2: This paragraph highlighted the significance of the achievement. I focused on making the contrast between past AI limitations (e.g., losing at chess) and the current capability (completing a 3D puzzle game). I used "groundbreaking" and "compelling" to underline the novelty.
-
Paragraph 3: This one dealt with the cost of the operation. I made sure to clarify the units (dollars, RMB) and the exchange rate. I also ensured the reader understood that the experimenter had a subscription that covered the costs.
-
Paragraph 4: This explained the technical details of the game control. I translated the technical terms related to the model (MCP, SPT). I needed to get the explanation of how the model interacted with the game right. The "thinking" state vs. "executing" state had to be clear.
-
Paragraph 5: It described how information was provided to the AI. I emphasized the visual input (screenshots) and the positional data. I made sure the description of the AI action flow was accurate.
-
Paragraph 6: This paragraph explains the length discrepancy between the livestream and the edited video. I wanted it to be concise.
-
Paragraph 7: Here, the experimenter's pragmatic view of the experiment was being described. I focused on conveying that the playthrough was not seen as a benchmark, which is key to avoid over-exaggeration.
-
Paragraph 8: It focused on what the achievement meant in relation to the initial project goals. I needed to convey the sense of progress, with "original vision" (OpenAI's goal) being realized.
-
Paragraph 9: It provided context to this "vision".
-
Paragraph 10: This contained information about the new GPT-6 Astra model. I wanted to clarify what this meant in terms of AI capabilities (computer use, browsing, software, and so on).
-
Paragraph 11: The last paragraph was about future expectations. I kept it short.
Finally, I did a quick review: No commentary included, the paragraph breaks were preserved and the tone was journalistic. I was ready to deliver the translation.
IT Home reported on September 7 that an AI and large language model enthusiast recently conducted an experiment in which OpenAI's new model, GPT-6 Astra, autonomously beat Valve's 3D puzzle game Portal.
From the results, this is undoubtedly a groundbreaking leap forward for large language models and a compelling demonstration of multimodal "AI intelligence." Not long ago, AI was still losing matches in Atari 2600 chess games, but now it can autonomously complete an entire 3D puzzle game.
This Portal playthrough involved a total of 3,336 tool calls, amounting to a cost of $571.18 based on API pricing (IT Home note: approximately 3,838 RMB at current exchange rates). CozyBlaze, the creator of the experiment, later clarified that these costs were actually covered by his $200/month (approximately 1,344 RMB at current exchange rates) Codex Pro subscription.
CozyBlaze explained: "The model controls Portal via MCP, which stands for Model Context Protocol, and a modified SourcePauseTool. While the model is thinking, the game remains paused; once the model sends a set of input commands, SPT unpauses and executes those actions."
While the AI is thinking, the system provides it with in-game screenshots and information such as the player character's position. The game then resumes running, and GPT-6 Astra carries out movements and other actions according to its pre-planned steps.
This also explains why the edited highlights video runs for about 2 hours, whereas the full livestream recordings of GPT-6 Astra playing Portal add up to roughly 24 hours.
Regarding the outcome of this experiment, CozyBlaze maintains a fairly pragmatic stance. He acknowledged that many issues still need to be addressed and that this playthrough should not be treated as a standardized benchmark for AI capabilities.
However, he remarked: "Still, seeing a general-purpose agent autonomously explore a game world and eventually complete the entire game makes it feel like the original vision is gradually becoming a reality."
The "original vision" CozyBlaze referred to is a long-term goal OpenAI proposed back in 2016, when the company envisioned creating a single AI agent capable of tackling diverse types of gaming tasks.
GPT-6 Astra became OpenAI's new flagship model earlier this month. OpenAI stated that the new model represents a "new generation of intelligence" and achieves state-of-the-art performance in areas such as computer use, web browsing, software engineering, cybersecurity, scientific research, and professional knowledge work.
In the coming days, Astra is expected to further showcase its powerful capabilities.
Why it's worth reading
Regardless of whether the cited model is genuine or speculative, the methodology of coupling the Model Context Protocol with engine-pause mechanics illustrates a tangible paradigm for long-horizon spatial reasoning.