Cheapest GPT API
Compare Cost per Accepted Output
最便宜的 GPT API:比较每个有效输出的成本
As of July 21, 2026, GPT-5.6 Luna has the lowest OpenAI list price at $1 input and $6 output per million tokens. The cheapest production route still depends on accepted-result quality, cache use, output length, retries, tools, and service tier.
截至 2026 年 7 月 21 日,GPT-5.6 Luna 的 OpenAI 目录价最低:每百万 Token 输入 1 美元、输出 6 美元。但生产环境中的最低有效成本仍取决于合格结果质量、缓存利用率、输出长度、重试、工具调用与服务层级。
TL;DR
Luna is the list-price floor in the GPT-5.6 family; Sol and Terra buy higher capability at higher token rates.
Compare exact models on identical prompts, context, reasoning effort, tools, output limits, quality gates, and region.
Include input, cached input, cache writes, output, tools, retries, batch discounts, regional premiums, and gateway fees.
Divide every attempt and non-token charge by outputs that pass the same quality, schema, safety, and latency gate.
Luna 是 GPT-5.6 系列的目录价底线;Sol 与 Terra 以更高 Token 费率换取更强能力。
用完全相同的提示词、上下文、推理强度、工具、输出限制、质量门槛和区域比较准确模型。
计入输入、缓存输入、缓存写入、输出、工具、重试、批处理折扣、区域溢价和网关费用。
把全部尝试与非 Token 费用除以通过同一质量、结构、安全和延迟门槛的结果数。
What “cheapest” should mean “最便宜”应如何定义
Compare direct provider access, cloud-platform access and managed gateways only when they deliver the same model version, region, features and service level. Otherwise label the difference rather than hiding it.
只有在模型版本、区域、功能与服务级别相同时,才能比较直连、Cloud Platform 与 Managed 网关;否则应明确标记差异,而不是隐藏。
Calculate total charges for a representative batch, then divide by outputs that pass your acceptance gate. Add engineering, support and compliance costs separately when they materially affect the decision.
计算代表性批次的总费用,再除以通过验收门禁的输出数。当工程、支持与合规成本显著影响决策时,应单独列出。
GPT access path comparison GPT 接入路径对比
| Access path 接入路径 | Best fit 最适合 | Verify before choosing 选择前验证 |
|---|---|---|
| Direct provider 直连供应商 | Official models and features with a direct commercial relationship. 官方模型与功能,以及直接商业关系。 | Current model price, tiers, regions, support, data terms and limits. 核对当前模型价格、层级、区域、支持、数据条款与限制。 |
| Cloud platform Cloud Platform | Existing cloud identity, networking, procurement or residency matters. 现有云身份、网络、采购或数据驻留很重要。 | Deployment names, regional availability, feature lag and platform pricing. 核对 Deployment、区域可用性、功能差异与平台价格。 |
| Managed gateway Managed 网关 | Multi-provider policy, observability and one integration are valuable. 多供应商策略、可观测性与统一集成有价值。 | Gateway fee, model markup, routing policy, data path and evidence quality. 核对网关费用、模型加价、路由策略、数据路径与证据质量。 |
| Batch or async 批处理或异步 | Work can wait and the official route supports asynchronous processing. 工作可等待且官方路由支持异步处理。 | Eligibility, completion window, failure handling and actual acceptance rate. 核对资格、完成窗口、失败处理与真实通过率。 |
True-cost worksheet 真实成本工作表
Separate prompt, cached input, reasoning and visible output dimensions.
Add search, code, files, images, audio, video and storage charges when used.
Count timeouts, invalid outputs, safety failures and fallback attempts.
Use workload tests with explicit quality, schema and latency gates.
区分提示词、缓存输入、推理与可见输出维度。
使用时加入搜索、代码、文件、图片、音频、视频与存储费用。
计算超时、无效输出、安全失败与回退尝试。
用明确质量、结构定义与延迟门禁进行工作负载测试。
Run a defensible cost test 运行可辩护的成本测试
- Snapshot current official pricing and terms for each exact route.
- Replay the same representative workload with fixed acceptance criteria.
- Reconcile native usage, retries, tools and route charges to the bill.
- Repeat after model, price, prompt or workload distribution changes.
- 为每条准确路由保存当前官方价格与条款快照。
- 使用固定验收标准回放同一代表性负载。
- 把原生用量、重试、工具与路由费用对账到账单。
- 模型、价格、提示词或负载分布变化后重新测试。
Measure cost after the quality gate 在质量门禁后计算成本
A versioned workload suite calls each eligible route. The ledger records every attempt and billing dimension. A quality gate marks accepted results. The comparison engine calculates total bill divided by accepted outputs and retains source URLs and timestamps.
版本化负载套件调用每条合格路由;账本记录每次尝试与计费维度;质量门禁标记有效结果;比较引擎用总账单除以有效输出,并保留来源 URL 与时间戳。
Production rule: never publish a cheapest claim without an exact model, route, workload, date and acceptance method.
生产规则:未说明准确模型、路由、负载、日期与验收方法时,绝不发布“最便宜”结论。
Separate model cost from capability cost 分离模型成本与能力成本
Keep inference and capability spend as separate ledger lines. Record GPT model, reasoning effort, cached and uncached tokens, and retries on the model span; record QVeris search, inspection, external API execution, billed amount, and result status on the capability span. Join both with one workflow ID only when calculating cost per accepted business outcome.
推理费用与能力费用应分别记账。模型链路记录 GPT 模型、推理强度、缓存与非缓存 Token 及重试;能力链路记录 QVeris 搜索、检查、外部 API 执行、计费金额与结果状态。只有在计算每个合格业务结果的成本时,才通过同一工作流 ID 汇总两类费用。
Current GPT-5.6 list-price ladder 当前 GPT-5.6 目录价阶梯
Official OpenAI prices verified July 21, 2026, in USD per 1 million tokens. Luna is the lowest list-price GPT-5.6 option, but the cheapest accepted route depends on task quality, reasoning effort, cached input, output length, retries, and tool fees.
以下 OpenAI 官方价格于 2026 年 7 月 21 日核验,单位为美元/百万 Token。Luna 的 GPT-5.6 目录价最低,但真实最低合格成本仍取决于任务质量、Reasoning Effort、缓存输入、输出长度、重试与工具费用。
| Model or route 模型或路由 | Input / cost 输入/成本 | Output 输出 | Scope 适用范围 |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $30.00 | Cached input $0.50 缓存输入 $0.50 |
| GPT-5.6 Terra | $2.50 | $15.00 | Cached input $0.25 缓存输入 $0.25 |
| GPT-5.6 Luna | $1.00 | $6.00 | Cached input $0.10 缓存输入 $0.10 |
Always verify the linked official pricing source immediately before a buying decision. Model quality, output length, cacheability, retries, tools, service tier, region, taxes, and discounts can reverse a token-price comparison.
采购决策前务必立即核验页面所链接的官方定价来源。模型质量、输出长度、可缓存性、重试、工具、服务层、区域、税费与折扣都可能逆转 Token 单价比较。
Build an Effective Cost Model for the cheapest GPT API route为成本最低的 GPT API 路径建立有效成本模型
For GPT-5.6, list price is only the first row of the model. The comparison must also account for Sol, Terra, or Luna selection, reasoning effort, explicit cache writes, discounted cache reads, long-context pricing, tool charges, retries, and outputs that fail the acceptance gate.
对 GPT-5.6 而言,目录价只是成本模型的第一行。比较还必须计入 Sol、Terra 或 Luna 的选择、推理强度、显式缓存写入、折扣缓存读取、长上下文计价、工具费用、重试以及未通过验收门槛的输出。
Build separate samples for short extraction, long-context research, structured output, tool use, and code. Run each sample on Luna, Terra, and Sol at the reasoning efforts the workload actually needs; report p50, p95, and the share of requests that cross the long-context pricing threshold instead of hiding them in one average.
分别建立短文本抽取、长上下文研究、结构化输出、工具调用和代码任务样本。按工作负载真实需要的推理强度在 Luna、Terra 与 Sol 上运行,报告 p50、p95 以及跨过长上下文计价阈值的请求占比,不要用单一平均值掩盖差异。
Give every task a pass/fail gate before running it: required facts, schema validity, tool correctness, safety constraints, maximum latency, and permitted review time. Calculate Luna, Terra, and Sol cost per passing result with all failed attempts included; a model is cheaper only when it clears the same gate.
运行前为每项任务定义通过/失败门槛:必需事实、结构有效性、工具正确性、安全约束、最大延迟与允许的复核时间。把所有失败尝试计入 Luna、Terra 与 Sol 的每个合格结果成本;只有通过同一门槛时,模型才真正更便宜。
Add GPT-5.6 cache-write charges, cache-read discounts, Batch savings where latency permits, tool-call fees, regional processing premiums, gateway markup, evaluation traffic, observability storage, reserved capacity, unused credits, taxes, and the engineering cost of failures or manual review.
加入 GPT-5.6 缓存写入费用、缓存读取折扣、延迟允许时的 Batch 节省、工具调用费、区域处理溢价、网关加价、评估流量、可观测存储、预留容量、未用额度、税费,以及失败或人工复核带来的工程成本。
Version the OpenAI pricing URL, retrieval timestamp, Sol/Terra/Luna model IDs, currency, region, processing mode, service tier, cache policy, context threshold, and discounts with the estimate. Recompute immediately after any model or pricing update and alert when invoice cost per accepted result diverges from the forecast.
把 OpenAI 定价 URL、获取时间、Sol/Terra/Luna 模型 ID、币种、区域、处理模式、服务层级、缓存策略、上下文阈值和折扣与估算一起版本化。模型或价格更新后立即重算,并在每个合格结果的发票成本偏离预测时告警。
FAQ
Within OpenAI's GPT-5.6 family on July 21, 2026, Luna has the lowest list price. Verify the live pricing page and your accepted-result cost before buying.
No. Include output mix, retries, tools, quality acceptance and support terms.
After material model, price, prompt, feature or workload-distribution changes.
在 2026 年 7 月 21 日的 OpenAI GPT-5.6 系列中,Luna 的目录价最低。采购前仍应核对实时定价页和自身的合格结果成本。
不应。应纳入输出结构、重试、工具、质量通过率与支持条款。
模型、价格、提示词、功能或负载分布发生重大变化后。