Roboflow has named GPT-6 Astra the strongest vision model it has tested after evaluating the OpenAI model across object detection, counting, visual reasoning and several tasks outside its standard benchmark. The results, published on September 29, are Roboflow’s own measurements rather than an independent industry-wide ranking.

At low reasoning effort, Astra recorded 82.1% mAP@50 in Roboflow’s object-detection evaluation. Roboflow reported that score as 5.4 points above Qwen3.8 Max and 13.7 points above GPT-5.6 Sol. The company said none of the competing models it tested reached Astra’s result even when configured for higher reasoning effort.

The evaluation emphasized the model’s ability to connect small visual details with category meaning. In one example, Astra separated several sizes of LEGO bricks despite differences in color, orientation and overlap, reaching 99.8% mAP@50. In another, it distinguished full and empty gas cylinders by interpreting a small valve detail without being explicitly told what physical feature marked the difference. Roboflow has made Astra the default model in its Auto Annotate product, while cautioning that unusual classes requiring specialized organizational knowledge may still need more human correction.

Roboflow also tested visual examples as prompts. Giving Astra positive and negative boxes helped narrow detections among similar dice, raising the reported score from 21.3% to 90.1%. A football example rose from 46.9% to 100%. The company said examples could transfer between images and videos, including tests involving tablets, capsules and bottle caps.

Counting results were 80.2% at low effort and 81.1% at high effort, compared with 74.3% and 76.1% for Sol. Astra also led Roboflow’s visual-reasoning evaluation with 87.2% accuracy at low effort and 91.2% at high effort. Tasks included reading dimensions from a technical drawing and determining whether a proposed container load exceeded a printed weight limit.

The strongest caveat appeared in segmentation. Astra could use language to identify the intended regions and return polygon vertices, but its generated outlines sometimes shifted or simplified object boundaries. Roboflow found SAM 3 more precise when exact mask shape mattered, and described Astra’s polygon generation as slower and more expensive. Its proposed workflow is therefore hybrid: use Astra to understand a requested class and provide boxes, then use SAM 3 to create tighter masks.

The tests indicate that Astra’s language understanding can be useful for visual annotation and reasoning, but the figures should be read within Roboflow’s specific evaluation setup. Performance, cost and accuracy can vary with prompts, images and domain-specific classes.