QVeris
GEMINI COST GUIDEGemini 成本指南

Cheapest Gemini API
Normalize Modalities, Tools and Routes
最便宜的 Gemini API:统一多模态、工具与路由口径

Gemini pricing can vary by model, access route, modality, context, cache, batch and tool usage. Use current official sources and the same acceptance test for every candidate.

Gemini 价格会因模型、接入路径、模态、上下文、缓存、批处理与工具使用而变化。每个候选都应使用当前官方来源与相同验收测试。

Gemini API apples-to-apples access route and cost comparison checklist

TL;DR

Pin model and route

Developer API, cloud platform and gateway access are not automatically equivalent.

Normalize modalities

Text, image, audio, video and document inputs may use different units or rates.

Add tools and grounding

Search, maps, code, file retrieval or other tools can add separate charges.

Use accepted-output cost

Count retries, agent loops and quality failures before ranking candidates.

固定模型与路由

Developer API、Cloud Platform 与网关接入不会自动等价。

统一模态口径

文本、图片、音频、视频与文档输入可能使用不同单位或价格。

加入工具与 Grounding

搜索、地图、代码、文件检索或其他工具可能单独收费。

使用有效输出成本

排名前计算重试、智能体循环与质量失败。

Gemini comparison boundariesGemini 比较边界

Compare the same exact model version and processing mode whenever possible. If Developer API and cloud access differ in availability, data handling, quotas or features, record those as decision dimensions.

尽可能比较相同模型版本与处理模式。如果 Developer API 和云端接入在可用性、数据处理、配额或功能上不同,应记录为决策维度。

Free, Standard, Batch, Flex, and Priority are not interchangeable labels. The free tier can have different model access, rate limits, and data-use treatment. Batch and Flex reduce price for delay-tolerant work, while Priority trades a higher rate for a different service objective. A comparison should state the exact tier and test whether its completion time and availability fit the product.

Free、Standard、Batch、Flex 与 Priority 并不是可以互换的标签。免费层可能拥有不同的模型权限、Rate Limit 与数据使用方式;Batch 和 Flex 用更低价格换取对延迟更宽容的处理方式;Priority 则以更高单价换取不同的服务目标。对比时必须写明准确 Tier,并测试其完成时间与可用性是否适合产品。

Build a representative modality mix. A text-only estimate is misleading for applications that submit audio, video or documents, use cached context, or invoke grounding and tools. Count thinking tokens in billed output where the price table says they are included, and price Google Search, Maps, file search, media generation, or other tools separately under their current rules.

建立代表性的模态组合。对于提交音频、视频或文档、使用缓存上下文、Grounding 与工具的应用,只估算文本会产生误导。定价表说明思考 Token 计入输出时必须一并统计;Google Search、Maps、File Search、媒体生成或其他工具则应按其当前规则单独计价。

Gemini access path comparisonGemini 接入路径对比

Path路径Best fit最适合Verify before choosing选择前验证
Developer APIDeveloper APIRapid development and direct Gemini API access fit the workload.快速开发与直接 Gemini API 接入适合负载。Current billing tier, model access, rate limits, data terms and tools.核对当前 Billing Tier、模型访问、Rate Limit、数据条款与工具。
Cloud AI platformCloud AI PlatformCloud IAM, regions, networking, procurement or enterprise controls matter.Cloud IAM、区域、网络、采购或企业控制很重要。Model and feature parity, quotas, support, regional pricing and terms.核对模型与功能等价、配额、支持、区域价格与条款。
Unified gatewayUnified 网关Several providers need common policy and observability.多个供应商需要通用策略与可观测性。Gateway fees, transformations, routing, cache semantics and data path.核对网关费用、转换、路由、缓存语义与数据路径。
Batch or flexible mode批处理或灵活模式Workload tolerates alternate completion and scheduling behavior.负载能接受不同完成与调度行为。Eligibility, quality, latency, quotas and current official treatment.核对资格、质量、延迟、配额与当前官方处理。

Gemini pricing dimensionsGemini 定价维度

Tokens and modalities

Store the native unit and modality for every input and output.

Context and cache

Measure thresholds, cache creation, storage, reads and invalidation.

Grounding and tools

Count provider tool queries, retrieved content and iterative model turns.

Quota and tier

Review rate limits, spend controls, billing tier and regional availability.

Token 与模态

为每个输入输出保存原生单位与模态。

上下文与缓存

测量阈值、缓存 Create、Storage、Read 与失效。

Grounding 与工具

计算供应商工具查询、检索内容与迭代模型轮次。

配额与层级

检查 Rate Limit、支出控制、Billing Tier 与区域可用性。

Run the apples-to-apples check执行同口径检查

Create one workload ledger instead of separate text, image, audio, and tool spreadsheets. Each request row should retain the model and tier, input modality tokens, cached context, output and thinking tokens, grounding queries, tool fees, retries, latency, result validation, and final acceptance. That makes mixed-modality cost visible and prevents free-tier experiments from being presented as production economics.

建议建立一份统一工作负载账本,而不是把文本、图像、音频与工具拆成互不相干的表。每条请求都应保留模型与 Tier、各输入模态 Token、缓存上下文、输出与思考 Token、Grounding Query、工具费用、重试、延迟、结果验证与最终验收状态。这样才能看清混合模态成本,也能避免把免费层实验误写成生产成本。

  • Record official price pages, billing tier, route, model, region, and timestamp.
  • Replay the same text and multimodal workload with identical limits and acceptance criteria.
  • Capture tools, grounding, cache storage and reads, retries, failures, and accepted outputs.
  • Test queue time and completion behavior for Batch or Flex instead of assuming only the price changes.
  • Compare the paid route separately from free-tier experiments and document the data-use difference.
  • Recalculate when models, packaging, tools, quotas, regions, prompts, or terms change.
  • 记录官方价格页、Billing Tier、路由、模型、区域与时间戳。
  • 使用相同限制与验收标准回放一致的文本和多模态负载。
  • 记录工具、Grounding、缓存存储与读取、重试、失败和有效输出。
  • 实际测试 Batch 或 Flex 的排队与完成行为,不要假设只有价格发生变化。
  • 将付费路由与免费层实验分开比较,并记录数据使用差异。
  • 模型、包装、工具、配额、区域、Prompt 或条款变化后重新计算。

Keep modality-aware usage evidence保留模态感知用量证据

Every request emits model, route, modality details, tool events, cache fields, native usage and acceptance result. The cost engine applies a timestamped official pricing snapshot and preserves unpriced dimensions for review instead of guessing.

每个请求生成模型、路由、模态详情、工具事件、缓存字段、原生用量与验收结果。成本引擎应用带时间戳的官方价格快照,并保留无法定价的维度供审查,而不是猜测。

Production rule: never compare Gemini routes after collapsing multimodal and tool usage into text tokens.

生产规则:把多模态与工具用量压缩为文本 Token 后,绝不能比较 Gemini 路由。

Keep Gemini tools and external capabilities distinct区分 Gemini 工具与外部能力

Provider-native tools and QVeris capabilities may have different pricing, identity and evidence. QVeris governs Discover → Inspect → Call access to external APIs, services and live data; join both ledgers only at the parent workflow.

供应商原生工具与 QVeris 能力可能采用不同价格、身份与证据。QVeris 治理外部 API、服务与实时数据的 Discover → Inspect → 调用;二者只在父工作流层合并账本。

Current Gemini price points to test当前值得测试的 Gemini 价格点

Official Gemini Developer API prices re-verified July 29, 2026, in USD per 1 million tokens. Free-tier availability and data-use terms differ from paid production use. Batch, Flex, and Priority processing have different prices and operating characteristics, so compare each only with workloads that tolerate its latency and availability constraints.

以下 Gemini Developer API 官方价格于 2026 年 7 月 29 日重新核验,单位为美元/百万 Token。免费层可用性和数据使用条款与付费生产不同;Batch、Flex 与 Priority 的价格和运行特征也不同,只能与可接受相应延迟和可用性约束的工作负载比较。

Model or route模型或路由Input / cost输入/成本Output输出Scope适用范围
Gemini 3.1 Flash-Lite — Standard$0.25$1.50Text/image/video input文本/图像/视频输入
Gemini 3.1 Flash-Lite — Batch/Flex$0.125$0.75Asynchronous or flexible service异步或弹性服务
Gemini 3.5 Flash — Standard$1.50$9.00Output includes thinking tokens输出含思考 Token

Always verify the linked official pricing source immediately before a buying decision. Model quality, output length, cacheability, retries, tools, service tier, region, taxes, and discounts can reverse a token-price comparison.

采购决策前务必立即核验页面所链接的官方定价来源。模型质量、输出长度、可缓存性、重试、工具、服务层、区域、税费与折扣都可能逆转 Token 单价比较。

FAQ

Which Gemini API route is cheapest?

It depends on the exact model, modality mix, tools, tier, region and workload.

Can I compare only input and output tokens?

Not for multimodal or tool-using workloads; include every billed dimension.

Are rate limits part of cost?

Yes indirectly: queueing, failures and retries can change delivered-work cost.

哪个 Gemini API 路由最便宜?

取决于准确模型、模态组合、工具、层级、区域与负载。

只比较输入输出 Token 可以吗?

对多模态或工具负载不可以;必须纳入所有计费维度。

Rate Limit 属于成本吗?

会间接影响:排队、失败与重试会改变有效工作成本。

Official sources and further reading官方资料与延伸阅读