LLM API Pricing Comparison
Normalize Before You Rank
LLM API 价格对比:先统一口径,再进行排序
A price table becomes useful only after routes, units, workloads and acceptance criteria match. Keep current official source data separate from measured workload evidence.
只有路由、单位、负载与验收标准一致后,价格表才有意义。应将当前官方来源数据与实测负载证据分开保存。
TL;DR
Fix prompt distribution, context, tools, modalities, latency and quality gates.
Map units for comparison but keep every provider-native billing field.
Store source URL, retrieval time, currency, region, model and route.
Calculate total delivered cost after retries, tools and quality failures.
固定提示词分布、上下文、工具、模态、延迟与质量门禁。
为比较映射单位,同时保留每个供应商原生计费字段。
保存来源 URL、获取时间、币种、区域、模型与路由。
在纳入重试、工具与质量失败后计算交付总成本。
Why simple price tables fail 为什么简单价格表会失真
Providers may separate input, output, cached input, cache storage, batch, priority, long context, media and tool charges. Access paths can add platform fees, commitments, taxes or regional differences.
供应商可能分别计费输入、输出、缓存输入、缓存存储、批处理、Priority、长上下文、媒体与工具。不同接入路径还会增加平台费、承诺、税费或区域差异。
Quality and latency change how much work is delivered. A candidate that needs longer prompts, more retries or human review may cost more per accepted outcome despite a lower listed unit price.
质量与延迟会改变实际交付量。即使标价更低,若候选需要更长提示词、更多重试或人工审查,每个有效结果的成本仍可能更高。
Comparison layers 比较层次
| Layer 层次 | Best fit 最适合 | Verify before choosing 选择前验证 |
|---|---|---|
| Official list price 官方标价 | Build a timestamped reference for exact models and routes. 为准确模型与路由建立带时间戳参考。 | Units, thresholds, currencies, tiers, regions and special modes. 核对单位、阈值、币种、层级、区域与特殊模式。 |
| Measured usage 实测用量 | Estimate each route with representative workload evidence. 用代表性负载证据估算每条路由。 | Native usage, cache, tools, media, retries and failed calls. 核对原生用量、缓存、工具、媒体、重试与失败调用。 |
| Accepted-output cost 有效输出成本 | Compare delivered work after a shared quality gate. 在共享质量门禁后比较已交付工作。 | Sample size, evaluator version, review cost and uncertainty. 核对样本量、Evaluator 版本、审查成本与不确定性。 |
| Total commercial cost 商业总成本 | Support procurement and long-term operating decisions. 支持采购与长期运营决策。 | Support, SLA, commitments, credits, taxes, data terms and engineering. 核对支持、SLA、承诺、抵扣、税费、数据条款与工程成本。 |
Comparison worksheet fields 比较工作表字段
Provider, platform, model version, endpoint, region and processing mode.
Input, output, cache, batch, tools, modalities, storage and extras.
Latency, retry rate, valid schema, quality acceptance and review effort.
Official URL, timestamp, owner, currency, tax basis and expiry review.
供应商、Platform、模型 Version、端点、Region 与 Processing Mode。
输入、输出、缓存、批处理、工具、模态、存储与附加项。
延迟、重试率、结构定义有效性、质量通过率与审查工作量。
官方 URL、时间戳、Owner、币种、税基与到期复核。
Create a repeatable comparison 建立可重复比较
- Version the workload, acceptance rubric and provider configuration.
- Fetch and review current official pricing and commercial terms.
- Run candidates, capture every attempt and reconcile provider usage.
- Publish results with date, scope, assumptions, uncertainty and next review.
- 版本化负载、验收 Rubric 与供应商配置。
- 获取并审查当前官方价格与商业条款。
- 运行候选、记录每次尝试并对账供应商用量。
- 发布结果时说明日期、范围、假设、不确定性与下次复核。
Join price snapshots to workload runs 连接价格快照与负载运行
A source registry stores official pricing snapshots. A workload runner creates traceable native-usage events and acceptance results. The comparison layer normalizes units, adds retry and tool costs, applies the quality gate and reports cost per accepted output.
Source 注册表保存官方价格快照;工作负载 Runner 生成可追踪的原生用量事件与验收结果;比较层标准化单位、加入重试与工具成本、应用质量门禁,并报告每个有效输出的成本。
Production rule: never mix price data from different dates, model versions or commercial routes without labeling the difference.
生产规则:绝不能在未标记差异的情况下混合不同日期、模型版本或商业路由的价格数据。
Compare external capability costs on their own ledger 在独立账本中比较外部能力成本
QVeris discovers and calls external APIs, tools, services and live data. Their fees may use requests, records, compute time or subscription units rather than tokens. Keep the native units, then join end-to-end cost at the workflow level.
QVeris 发现并调用外部 API、工具、服务与实时数据。其费用可能按请求、记录、计算时间或订阅单位计算,而非 Token。保留原生单位,再在工作流层合并端到端成本。
A dated API price snapshot for workload tests 用于工作负载测试的定价快照
Verified against official pricing pages on July 21, 2026. Prices are USD per 1 million tokens and are not quality-equivalent recommendations. Use them to seed a workload calculator, then refresh before procurement or publication because providers can change models and rates.
以下数据于 2026 年 7 月 21 日根据官方定价页核验,单位为美元/百万 Token,并不表示模型质量等价。可用来初始化工作负载计算器,但采购或发布前必须再次刷新。
| Model or route 模型或路由 | Input / cost 输入/成本 | Output 输出 | Scope 适用范围 |
|---|---|---|---|
| OpenAI GPT-5.6 Luna | $1.00 | $6.00 | Cached input $0.10/MTok 缓存输入 $0.10/百万 Token |
| Claude Haiku 4.5 | $1.00 | $5.00 | Cache read $0.10/MTok 缓存读取 $0.10/百万 Token |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | Standard text/image/video input 标准文本/图像/视频输入 |
| DeepSeek V4 Flash | $0.14 | $0.28 | Cache-miss input; cache hit $0.0028 未命中缓存输入;命中为 $0.0028 |
When evaluating LLM API Pricing Comparison, verify the linked official pricing source immediately before a buying decision. Model quality, output length, cacheability, retries, tools, service tier, region, taxes, and discounts can reverse a token-price comparison.
评估“LLM API 价格对比”时,采购决策前务必立即核验页面所链接的官方定价来源。模型质量、输出长度、可缓存性、重试、工具、服务层、区域、税费与折扣都可能逆转 Token 单价比较。
Build an Effective Cost Model for an LLM API pricing comparison为LLM API 定价比较建立有效成本模型
When evaluating LLM API Pricing Comparison, list price is an input, not the decision. Compare the cost of an accepted production outcome after quality, retries, latency, operational work, and non-token charges are included.
评估“LLM API 价格对比”时,目录价只是输入,而不是最终决策。应在计入质量、重试、延迟、运营工作和非 Token 费用后,比较获得一个合格生产结果的成本。
Sample production-shaped tasks and record provider-specific token definitions, input and output ratios, cache accounting, long-context thresholds, batch discounts, modalities, tools, and currency or tax treatment. Use percentiles and task classes rather than one average prompt so long-context and output-heavy requests remain visible.
抽取接近生产形态的任务,并记录供应商专属 Token 定义、输入与输出比例、缓存计量、长上下文阈值、批处理折扣、模态、工具,以及币种或税费处理。使用分位数和任务类别,而不是单个平均提示词,确保长上下文与高输出请求不会被隐藏。
When evaluating LLM API Pricing Comparison, measure completion, rubric score, structured-output validity, tool accuracy, retry rate, and human-review time. Divide total run cost by accepted outcomes, not by raw requests.
评估“LLM API 价格对比”时,对每个模型与路由衡量完成率、评分标准得分、结构化输出有效率、工具准确率、重试率和人工复核时间,并用总运行成本除以合格结果,而不是原始请求数。
When evaluating LLM API Pricing Comparison, add gateway or aggregator fees, storage, egress, cache writes, evaluations, observability, support, engineering, incident handling, reserved capacity, unused credits, taxes, and the business cost of latency or failed work.
评估“LLM API 价格对比”时,加入网关或聚合费用、存储、流量、缓存写入、评估、可观察性、支持、工程、事故处理、预留容量、未用额度、税费,以及延迟或任务失败的业务成本。
When evaluating LLM API Pricing Comparison, store source URL, retrieval date, currency, region, service tier, thresholds, discounts, and model version. Recompute scenarios when a catalog changes and alert when observed invoice cost diverges from the estimate.
评估“LLM API 价格对比”时,保存来源链接、获取日期、币种、区域、服务层级、阈值、折扣和模型版本。目录变化时重新计算情景,并在实际发票成本偏离估算时发出告警。
FAQ
Only as a dated snapshot with an owner and review process; prices change.
Cost per accepted workload result, supported by full billing and quality evidence.
Yes when they materially affect production feasibility or total cost.
只能作为带日期、Owner 与复核流程的快照;价格会变化。
每个有效工作负载结果的成本,并有完整计费与质量证据支持。
当其显著影响生产可行性或总成本时,应纳入。