QVeris
COST DECISION GUIDE成本决策指南

Cheapest GPT API
Compare Cost per Accepted Output
最便宜的 GPT API:比较每个有效输出的成本

There is no permanently cheapest GPT route. Model versions, access paths, processing modes and prices change; the right comparison uses the same workload and current official terms.

不存在永久最便宜的 GPT 路由。模型版本、接入路径、处理模式与价格都会变化;正确比较必须使用相同负载与当前官方条款。

Cheapest GPT API true workload cost comparison worksheet

TL;DR

Freeze the workload

Use the same prompts, context, tools, output limits, quality bar and region.

Verify current prices

Record retrieval time, model version, access path and official source URL.

Count the full bill

Include input, output, cache, batch, tools, media, retries and route fees.

Divide by accepted outputs

A low unit rate loses when quality failures and retries consume extra calls.

固定工作负载

使用相同提示词、上下文、工具、输出限制、质量门槛与区域。

核对当前价格

记录获取时间、模型版本、接入路径与官方来源 URL。

计算完整账单

纳入输入、输出、缓存、批处理、工具、媒体、重试与路由费用。

除以有效输出

单位价低但质量失败和重试更多时,真实成本可能更高。

What “cheapest” should mean“最便宜”应如何定义

Compare direct provider access, cloud-platform access and managed gateways only when they deliver the same model version, region, features and service level. Otherwise label the difference rather than hiding it.

只有在模型版本、区域、功能与服务级别相同时,才能比较直连、云平台与托管网关;否则应明确标记差异,而不是隐藏。

OpenAI's current table separates short- and long-context prices and distinguishes Standard, Batch, Flex, and Priority service. It also lists cached input and, for supported models, cache-write rates. Regional processing can add an uplift, while web search, file search, containers, image generation, audio, and other tools have their own billing units. A headline input rate therefore describes only one part of one route.

OpenAI 当前定价表区分短上下文与长上下文价格,也区分 Standard、Batch、Flex 与 Priority 服务;同时列出缓存输入,并为支持的模型列出缓存写入价格。区域处理可能产生溢价,Web Search、File Search、Container、图像生成、音频及其他工具也有独立计费单位。因此,标题中的输入单价只代表某条路由的一部分。

Calculate total charges for a representative batch, then divide by outputs that pass your acceptance gate. Include input, cached input, cache writes, visible output, reasoning tokens where billed, retries, failed attempts, and tool charges. Add engineering, support and compliance costs separately when they materially affect the decision.

计算代表性批次的总费用,再除以通过验收门禁的输出数。应纳入输入、缓存输入、缓存写入、可见输出、按规则计费的推理 Token、重试、失败尝试与工具费用。当工程、支持与合规成本显著影响决策时,再单独列出。

GPT access path comparisonGPT 接入路径对比

Access path接入路径Best fit最适合Verify before choosing选择前验证
Direct provider直连供应商Official models and features with a direct commercial relationship.官方模型与功能,以及直接商业关系。Current model price, tiers, regions, support, data terms and limits.核对当前模型价格、层级、区域、支持、数据条款与限制。
Cloud platformCloud PlatformExisting cloud identity, networking, procurement or residency matters.现有云身份、网络、采购或数据驻留很重要。Deployment names, regional availability, feature lag and platform pricing.核对 Deployment、区域可用性、功能差异与平台价格。
Managed gatewayManaged 网关Multi-provider policy, observability and one integration are valuable.多供应商策略、可观测性与统一集成有价值。Gateway fee, model markup, routing policy, data path and evidence quality.核对网关费用、模型加价、路由策略、数据路径与证据质量。
Batch or async批处理或异步Work can wait and the official route supports asynchronous processing.工作可等待且官方路由支持异步处理。Eligibility, completion window, failure handling and actual acceptance rate.核对资格、完成窗口、失败处理与真实通过率。

True-cost worksheet真实成本工作表

Input and output

Separate prompt, cached input, reasoning and visible output dimensions.

Tools and media

Add search, code, files, images, audio, video and storage charges when used.

Retry waste

Count timeouts, invalid outputs, safety failures and fallback attempts.

Acceptance rate

Use workload tests with explicit quality, schema and latency gates.

输入与输出

区分提示词、缓存输入、推理与可见输出维度。

工具与媒体

使用时加入搜索、代码、文件、图片、音频、视频与存储费用。

重试浪费

计算超时、无效输出、安全失败与回退尝试。

通过率

用明确质量、结构定义与延迟门禁进行工作负载测试。

Run a defensible cost test运行可辩护的成本测试

Use production-shaped prompts rather than a synthetic one-line benchmark. The replay set should include the real context-length distribution, cacheable and unique prefixes, short and long answers, structured outputs, tool calls, rate-limit pressure, and requests that fail the quality gate. Compare service tiers separately because a lower token rate is not equivalent if completion time or availability changes.

不要只用一行合成 Prompt 做基准,应使用接近生产的请求。回放集要覆盖真实上下文长度分布、可缓存与独特前缀、长短答案、结构化输出、Tool Call、限流压力以及未通过质量门禁的请求。不同 Service Tier 应分开比较,因为完成时间或可用性变化后,较低 Token 单价并不代表等价服务。

  • Snapshot current official pricing, model ID, context band, service tier, region, and terms for each exact route.
  • Replay the same representative workload with fixed schema, quality, latency, and safety criteria.
  • Capture native request IDs and usage fields instead of estimating tokens from text length.
  • Reconcile cached input, cache writes, reasoning or output, retries, failed attempts, tools, and route charges to the bill.
  • Calculate cost per accepted result and per completed user workflow, not only cost per call.
  • Repeat after model, price, prompt, SDK, tool, or workload-distribution changes.
  • 为每条准确路由保存当前官方价格、模型 ID、上下文区间、Service Tier、区域与条款快照。
  • 使用固定 Schema、质量、延迟与安全标准回放同一代表性负载。
  • 记录原生 Request ID 与用量字段,不要只按文本长度估算 Token。
  • 把缓存输入、缓存写入、推理或输出、重试、失败尝试、工具与路由费用对账到账单。
  • 计算每个有效结果及每个完整用户工作流的成本,而不只是单次调用成本。
  • 模型、价格、Prompt、SDK、工具或负载分布变化后重新测试。

Measure cost after the quality gate在质量门禁后计算成本

A versioned workload suite calls each eligible route. The ledger records every attempt and billing dimension. A quality gate marks accepted results. The comparison engine calculates total bill divided by accepted outputs and retains source URLs and timestamps.

版本化负载套件调用每条合格路由;账本记录每次尝试与计费维度;质量门禁标记有效结果;比较引擎用总账单除以有效输出,并保留来源 URL 与时间戳。

Production rule: never publish a cheapest claim without an exact model, route, workload, date and acceptance method.

生产规则:未说明准确模型、路由、负载、日期与验收方法时,绝不发布“最便宜”结论。

Separate model cost from capability cost分离模型成本与能力成本

QVeris complements GPT access with governed discovery and execution of external APIs, tools, services and live data. Track those capability charges separately, then join them to the parent workflow when calculating end-to-end cost.

QVeris 通过治理化的外部 API、工具、服务与实时数据发现和执行补充 GPT 访问。能力费用应单独追踪,再在计算端到端成本时关联父工作流。

Current GPT-5.6 list-price ladder当前 GPT-5.6 目录价阶梯

Official OpenAI prices verified July 21, 2026, in USD per 1 million tokens. Luna is the lowest list-price GPT-5.6 option, but the cheapest accepted route depends on task quality, reasoning effort, cached input, output length, retries, and tool fees.

以下 OpenAI 官方价格于 2026 年 7 月 21 日核验,单位为美元/百万 Token。Luna 的 GPT-5.6 目录价最低,但真实最低合格成本仍取决于任务质量、Reasoning Effort、缓存输入、输出长度、重试与工具费用。

Model or route模型或路由Input / cost输入/成本Output输出Scope适用范围
GPT-5.6 Sol$5.00$30.00Cached input $0.50缓存输入 $0.50
GPT-5.6 Terra$2.50$15.00Cached input $0.25缓存输入 $0.25
GPT-5.6 Luna$1.00$6.00Cached input $0.10缓存输入 $0.10

Always verify the linked official pricing source immediately before a buying decision. Model quality, output length, cacheability, retries, tools, service tier, region, taxes, and discounts can reverse a token-price comparison.

采购决策前务必立即核验页面所链接的官方定价来源。模型质量、输出长度、可缓存性、重试、工具、服务层、区域、税费与折扣都可能逆转 Token 单价比较。

FAQ

Which GPT API is cheapest today?

It depends on the exact model, route, region and workload; verify current official pricing.

Should I choose the lowest token price?

No. Include output mix, retries, tools, quality acceptance and support terms.

How often should I compare?

After material model, price, prompt, feature or workload-distribution changes.

今天哪个 GPT API 最便宜?

取决于准确模型、路由、区域与负载;请核对当前官方价格。

应选择 Token 单价最低的吗?

不应。应纳入输出结构、重试、工具、质量通过率与支持条款。

多久比较一次?

模型、价格、提示词、功能或负载分布发生重大变化后。

Official sources and further reading官方资料与延伸阅读