QVeris
DEEPSEEK COST GUIDEDeepSeek 成本指南

Cheapest DeepSeek API
Verify the Route, Measure Accepted Work
最便宜的 DeepSeek API:核对路由,测量有效工作量

DeepSeek pricing, models and availability can change. Compare exact authorized routes using current official sources, native usage and the same workload acceptance gate.

DeepSeek 的价格、模型与可用性会变化。应使用当前官方来源、原生用量与相同负载验收门禁比较准确的授权路由。

DeepSeek API authorized access route and true cost comparison worksheet

TL;DR

Pin the exact model

Record model ID, version or alias resolution, thinking mode, endpoint and date.

Verify official pricing

Use the current provider or platform page, not an old article or screenshot.

Count cache and retries

Separate cache hit and miss, input, output, failed calls and every attempt.

Rank accepted outcomes

Apply one quality, schema and latency gate before calculating true cost.

固定准确模型

记录模型 ID、Version 或别名 Resolution、Thinking Mode、端点与日期。

核对官方价格

使用当前供应商或平台页面,而不是旧文章或截图。

计算缓存与重试

区分缓存 Hit/Miss、输入、输出、失败调用与每次尝试。

按有效结果排序

先应用统一质量、结构定义与延迟门禁,再计算真实成本。

DeepSeek access routes are not identicalDeepSeek 接入路由并不相同

A first-party API, cloud model platform and managed gateway can differ in model version, protocol, context, cache behavior, tools, concurrency, region, data terms, support and commercial treatment.

第一方 API、云模型平台与托管网关可能在模型版本、协议、上下文、缓存、工具、并发、区域、数据条款、支持与商业处理上不同。

The current first-party V4 route distinguishes Flash and Pro and bills cache-hit input, cache-miss input, and output at different rates. That makes a claimed token price meaningless without the observed cache ratio. Thinking mode can also change output length, latency, and acceptance quality, while JSON output and tool calls must be tested against the exact model and protocol used in production.

当前第一方 V4 路由区分 Flash 与 Pro,并分别计费缓存命中输入、缓存未命中输入和输出。因此,如果没有真实缓存命中率,所谓 Token 单价并不能代表最终成本。Thinking Mode 也会改变输出长度、延迟与验收质量;JSON Output 和 Tool Call 则必须针对生产环境使用的准确模型与协议单独测试。

Compare only like-for-like routes. When a platform serves a different snapshot or adds services, keep those differences visible and evaluate the total workload outcome rather than a single token rate. Record the base URL, resolved model ID, context and output limits, concurrency policy, region, cache semantics, retries, and provider of record for every candidate.

只比较同口径路由。当平台提供不同 Snapshot 或附加服务时,应明确差异并评估完整负载结果,而不是只看某个 Token 单价。每个候选都要记录 Base URL、实际解析模型 ID、上下文与输出限制、并发策略、区域、缓存语义、重试以及合同供应方。

DeepSeek route comparisonDeepSeek 路由对比

Route路由Best fit最适合Verify before choosing选择前验证
First-party API一方 APIDirect service, current native docs and billing fit the workload.直连服务、当前原生文档与账单适合负载。Exact model, base URL, cache rules, limits, support and data terms.核对准确模型、基础地址(Base URL)、缓存规则、限制、支持与数据条款。
Cloud model platformCloud 模型 PlatformCloud identity, region, procurement or operations are decisive.云身份、区域、采购或运营是关键。Snapshot, provider of record, protocol, feature parity and platform price.核对 Snapshot、合同供应方、协议、功能等价与平台价格。
Managed gatewayManaged 网关Unified policy, telemetry or multi-provider routing adds value.统一策略、遥测或多供应商路由有价值。Markup, data path, adapter behavior, retries and provider evidence.核对加价、数据路径、适配器行为、重试与供应商证据。
Self-hosted weights自托管权重License and operations permit control of the serving stack.License 与运营允许控制 Serving Stack。Hardware, utilization, engineering, model identity, updates and security.核对硬件、利用率、工程、模型身份、更新与安全。

True-cost dimensions真实成本维度

Native usage

Input, cache status, output, reasoning or intermediate usage when exposed.

Delivery overhead

Retries, queueing, timeouts, rejected output and human review.

Route overhead

Gateway fee, cloud charges, currency, tax, commitments and support.

Operational fit

Throughput, latency, region, data policy, observability and recovery.

原生用量

输入、缓存状态、输出,以及公开时的推理或中间用量。

交付开销

重试、排队、超时、未通过输出与人工审查。

路由开销

网关费、Cloud Charge、币种、税费、承诺与支持。

运营适配

吞吐、延迟、区域、数据策略、可观测性与恢复。

Run a dated comparison运行带日期的比较

Start with a replay set that resembles production rather than a handful of short prompts. Include cacheable and unique prefixes, short and long outputs, thinking and non-thinking tasks, structured responses, tool calls, rate-limit pressure, and provider errors. Apply the same schema, factuality, latency, and safety gate before calculating cost per accepted result.

测试集应接近生产分布,而不是只放几个短 Prompt。它应同时覆盖可缓存与独特前缀、长短输出、Thinking 与 Non-thinking 任务、结构化响应、Tool Call、限流压力和 Provider Error。计算每个有效结果的成本前,必须统一应用 Schema、事实性、延迟与安全门禁。

  • Snapshot official model, pricing, limits and terms pages for each route.
  • Replay the same prompt distribution and acceptance criteria.
  • Record the resolved model ID instead of trusting a legacy alias or gateway label.
  • Reconcile native usage, cache hits and misses, retries, failed attempts, and route charges to invoices.
  • Measure end-to-end latency and throughput under the same concurrency, not only single-call response time.
  • Repeat after any model, price, alias, endpoint, SDK, prompt, or workload change.
  • 为每条路由保存官方模型、价格、限制与条款页面快照。
  • 回放相同提示词分布与验收标准。
  • 记录实际解析模型 ID,不要只相信旧别名或网关标签。
  • 把原生用量、缓存命中与未命中、重试、失败尝试及路由费用对账到账单。
  • 在相同并发下测量端到端延迟与吞吐,而不是只看单次响应时间。
  • 模型、价格、别名、端点、SDK、Prompt 或负载变化后重新比较。

Separate pricing snapshots from measured usage分离价格快照与实测用量

A source registry records the official URL, retrieval time, model and route. A workload runner captures native request IDs, usage, cache status, attempts and acceptance. The cost layer applies the dated price snapshot and reports total charges per accepted output.

Source 注册表记录官方 URL、获取时间、模型与路由;工作负载 Runner 捕获原生请求 ID、用量、缓存状态、尝试与验收;成本层应用带日期的价格快照并报告每个有效输出的总费用。

Production rule: never label a DeepSeek API route cheapest without an exact model, route, workload, date and acceptance method.

生产规则:未说明准确模型、路由、负载、日期与验收方法时,绝不能把某条 DeepSeek API 路由标为最便宜。

Keep capability calls outside the model price将能力调用与模型价格分开

QVeris complements DeepSeek inference with governed external APIs, tools, services and live data. Record those calls in their native units and join them only when calculating the parent workflow's end-to-end cost.

QVeris 通过治理化外部 API、工具、服务与实时数据补充 DeepSeek 推理。用原生单位记录这些调用,只在计算父工作流端到端成本时合并。

Current DeepSeek V4 prices and migration note当前 DeepSeek V4 价格与迁移提示

Official DeepSeek prices re-verified July 29, 2026, in USD per 1 million tokens. DeepSeek's documented July 24 deprecation date for the deepseek-chat and deepseek-reasoner aliases has now passed. New deployments should use the exact V4 model IDs. Existing deployments should inspect request logs to confirm the resolved model, migrate intentionally, and rerun output-schema, tool-calling, thinking-mode, latency, cache, and cost tests before increasing traffic.

以下 DeepSeek 官方价格于 2026 年 7 月 29 日重新核验,单位为美元/百万 Token。DeepSeek 文档所列 deepseek-chat 与 deepseek-reasoner 别名的 7 月 24 日弃用日期已经过去。新部署应直接使用准确的 V4 模型 ID;现有部署应检查请求日志确认实际解析模型,再有计划地迁移,并在扩大流量前重新验证输出 Schema、Tool Calling、Thinking Mode、延迟、缓存与成本。

Model or route模型或路由Input / cost输入/成本Output输出Scope适用范围
DeepSeek V4 Flash$0.14 cache miss / $0.0028 hit$0.28Concurrency limit shown as 2,500文档并发限制 2,500
DeepSeek V4 Pro$0.435 cache miss / $0.003625 hit$0.87Concurrency limit shown as 500文档并发限制 500

Always verify the linked official pricing source immediately before a buying decision. Model quality, output length, cacheability, retries, tools, service tier, region, taxes, and discounts can reverse a token-price comparison.

采购决策前务必立即核验页面所链接的官方定价来源。模型质量、输出长度、可缓存性、重试、工具、服务层、区域、税费与折扣都可能逆转 Token 单价比较。

FAQ

Which DeepSeek route is cheapest now?

It depends on the exact model, platform, region and workload; verify current sources.

Can I compare model aliases?

Only after recording which exact version each alias resolved to during the test.

Does cache change the result?

Yes. Measure real hit and miss behavior rather than assuming a cache ratio.

现在哪条 DeepSeek 路由最便宜?

取决于准确模型、平台、区域与负载;请核对当前来源。

可以比较模型别名吗?

只有记录测试时每个别名解析到的准确版本后才可以。

缓存会改变结果吗?

会。应测量真实 Hit/Miss 行为,而不是假设缓存比例。

Official sources and further reading官方资料与延伸阅读