OpenAI GPT-6 Astra Model Breaks Records in ARC-AGI-3 Benchmark Examination

Serdar HocamAuthor & Editor

GPT-6 Astra achieved outstanding success in the ARC-AGI-3 benchmark test with standard and provider adapter hardware, surpassing human efficiency.

◉ 0 views
ARC-AGI-3 leaderboard showing GPT-6 Astra Standard and Provider Adapter results
GPT-6 Astra achieves state-of-the-art scores on ARC-AGI-3 with both the Standard and Provider Adapter harnesses. Higher reasoning levels generally cost less because Astra solves games in fewer actions, reducing the total number of model calls and tokens. View the full results.

Developed by OpenAI, GPT-6 Astra achieved record-breaking scores in the ARC-AGI-3 Semi-Private benchmark test, setting a new threshold in its field.

Successes Achieved in the Benchmark Test

The Astra model achieved a score of 62.7 percent at a cost of $26,000 with Standard hardware in the ARC-AGI-3 test, while reaching 99.9 percent success at a cost of $19,000 with Provider Adapter hardware.

Surpassing Human Efficiency

In evaluations, GPT-6 Astra outperformed the human baseline by performing fewer actions than the median human player tested in 96 percent of the levels.

Symbolic World Models and Action Planning

It was observed that Astra can convert unknown environments into compact symbolic world models, represent game mechanics as logical rules, and create its own domain for state tracking and action planning.

ARC-AGI-3 Measurement Criteria

A third-generation benchmark, ARC-AGI-3 tests four core components of agentic intelligence: exploration, modeling, goal setting and planning, and execution.

Research Findings and Scope Limitations

While replay analyses revealed custom algebraic notation, high action efficiency, and specialized tools built on advanced agent hardware, it is emphasized that the test does not fully reflect the open-endedness of the real world.