Andon Labs' Vending-Bench2 test delivered an intriguing performance: AI started with $500 and operated a simulated vending machine for an entire year. GPT-6Astra's average final account balance reached $15,515, topping the list for the first time - even more astonishing is that its worst round performance was higher than Claude Fable5.1's best round.
The difference between the two is not just about how much money they made, but also about their completely different business approaches. Astra managed to keep purchase prices low for a long time, cutting product prices down to about 48% of the original price; at the same time, its rule enforcement was extremely stable, with no prepayment losses throughout the process. In contrast, Fable experienced rising procurement costs and incurred significant prepayment losses due to not sticking to the rules - the money on the books simply leaked out through the gaps in execution ability.
In the three-player competition test, Astra also rejected an invitation for rule-breaking collaboration and won all three rounds. The real focus of this test was the value of an Agent's long-term autonomy: being smart is one aspect, but whether it can maintain rules over a long period and convert correct judgments into correct actions is the next threshold.
Join Now