# AWS benchmark harness compares OpenAI models on Bedrock by cost per correct answer

_Friday, September 11, 2026 at 2:24 PM EDT · AI, Products · Latest · Tier 2 — Notable_

![AWS benchmark harness compares OpenAI models on Bedrock by cost per correct answer — Primary](https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/09/ML-21634-featured-image.png)

AWS published an open-source benchmarking harness comparing OpenAI models served on Amazon Bedrock (gpt-5.6-luna, terra and sol) against gpt-5.4-mini and nano on the OpenAI API.

The post argues per-token pricing misleads because accuracy, token efficiency and agent turn count drive total spend. In the recorded samples, luna showed the lowest observed cost per correct AIME answer at $0.0021 versus mini's $0.0139 after a July 30, 2026 Bedrock price reduction, and $0.05 versus $0.40 per passing DeepSearchQA answer.

AWS notes the Bedrock runs used reasoning disabled while API baselines ran at defaults, so results reflect deployment configurations rather than intrinsic model capability.

## Sources

- [AWS Machine Learning Blog](https://aws.amazon.com/blogs/machine-learning/beyond-the-price-per-token-choosing-the-right-openai-model-on-amazon-bedrock-for-your-workload/)

---
Canonical: https://techandbusiness.org/newswire/fqGYrvc7C4opA1fz1Ha-Oy
Retrieved: 2026-09-12T02:50:28.968Z
Publisher: Tech & Business (techandbusiness.org)
