GPT-6 Astra on robot arms
Robocurve ran OpenAI’s GPT-6 Astra against Claude Fable 5 and 5.1 on identical YAM robot-arm tasks. Its bowl-placement result was 19 out of 20 for Astra, versus 8 out of 20 and 1 out of 20. The reported trial time was 2.5 minutes against 6.8, with estimated cost per run of $0.94 against $2.12. Physical tasks are a useful test because the environment can punish a plausible plan that does not translate into action. These are third-party reported results, not a general robotics verdict, but the task design makes the operational signal worth watching.
Our takeWe want evidence from real environments, not just a model saying it understands one. This is a strong result set, and it needs replication beyond one task family.
Source