Cheapest Gemini API
Normalize Modalities, Tools and Routes
最便宜的 Gemini API:统一多模态、工具与路由口径
Gemini pricing can vary by model, access route, modality, context, cache, batch and tool usage. Use current official sources and the same acceptance test for every candidate.
Gemini 价格会因模型、接入路径、模态、上下文、缓存、批处理与工具使用而变化。每个候选都应使用当前官方来源与相同验收测试。
TL;DR
Developer API, cloud platform and gateway access are not automatically equivalent.
Text, image, audio, video and document inputs may use different units or rates.
Search, maps, code, file retrieval or other tools can add separate charges.
Count retries, agent loops and quality failures before ranking candidates.
Developer API、Cloud Platform 与网关接入不会自动等价。
文本、图片、音频、视频与文档输入可能使用不同单位或价格。
搜索、地图、代码、文件检索或其他工具可能单独收费。
排名前计算重试、智能体循环与质量失败。
Gemini comparison boundaries Gemini 比较边界
Compare the same exact model version and processing mode whenever possible. If Developer API and cloud access differ in availability, data handling, quotas or features, record those as decision dimensions.
尽可能比较相同模型版本与处理模式。如果 Developer API 和 Cloud 接入在可用性、数据处理、配额或功能上不同,应记录为决策维度。
Build a representative modality mix. A text-only estimate is misleading for applications that submit audio, video or documents, use cached context, or invoke grounding and tools.
建立代表性的模态组合。对于提交音频、视频或文档、使用 Cached 上下文、Grounding 与工具的应用,只估算文本会产生误导。
Gemini access path comparison Gemini 接入路径对比
| Path 路径 | Best fit 最适合 | Verify before choosing 选择前验证 |
|---|---|---|
| Developer API Developer API | Rapid development and direct Gemini API access fit the workload. 快速开发与直接 Gemini API 接入适合负载。 | Current billing tier, model access, rate limits, data terms and tools. 核对当前 Billing Tier、模型访问、Rate Limit、数据条款与工具。 |
| Cloud AI platform Cloud AI Platform | Cloud IAM, regions, networking, procurement or enterprise controls matter. Cloud IAM、区域、网络、采购或企业控制很重要。 | Model and feature parity, quotas, support, regional pricing and terms. 核对模型与功能等价、配额、支持、区域价格与条款。 |
| Unified gateway Unified 网关 | Several providers need common policy and observability. 多个供应商需要通用策略与可观测性。 | Gateway fees, transformations, routing, cache semantics and data path. 核对网关费用、转换、路由、缓存语义与数据路径。 |
| Batch or flexible mode 批处理或灵活模式 | Workload tolerates alternate completion and scheduling behavior. 负载能接受不同完成与调度行为。 | Eligibility, quality, latency, quotas and current official treatment. 核对资格、质量、延迟、配额与当前官方处理。 |
Gemini pricing dimensions Gemini 定价维度
Store the native unit and modality for every input and output.
Measure thresholds, cache creation, storage, reads and invalidation.
Count provider tool queries, retrieved content and iterative model turns.
Review rate limits, spend controls, billing tier and regional availability.
为每个输入输出保存原生单位与模态。
测量阈值、缓存 Create、Storage、Read 与失效。
计算供应商工具查询、检索内容与迭代模型轮次。
检查 Rate Limit、支出控制、Billing Tier 与区域可用性。
Run the apples-to-apples check 执行同口径检查
- Record official price pages, billing tier, route, model and timestamp.
- Replay the same text and multimodal workload with identical limits.
- Capture tools, grounding, cache, retries, failures and accepted outputs.
- Recalculate when models, packaging, tools, quotas or terms change.
- 记录官方价格页、Billing Tier、路由、模型与时间戳。
- 使用相同限制回放一致的文本与多模态负载。
- 记录工具、Grounding、缓存、重试、失败与有效输出。
- 模型、包装、工具、配额或条款变化后重新计算。
Keep modality-aware usage evidence 保留模态感知用量证据
Every request emits model, route, modality details, tool events, cache fields, native usage and acceptance result. The cost engine applies a timestamped official pricing snapshot and preserves unpriced dimensions for review instead of guessing.
每个请求生成模型、路由、模态详情、工具事件、缓存字段、原生用量与验收结果。成本引擎应用带时间戳的官方价格快照,并保留无法定价的维度供审查,而不是猜测。
Production rule: never compare Gemini routes after collapsing multimodal and tool usage into text tokens.
生产规则:把多模态与工具用量压缩为文本 Token 后,绝不能比较 Gemini 路由。
Keep Gemini tools and external capabilities distinct 区分 Gemini 工具与外部能力
Provider-native tools and QVeris capabilities may have different pricing, identity and evidence. QVeris governs Discover → Inspect → Call access to external APIs, services and live data; join both ledgers only at the parent workflow.
供应商原生工具与 QVeris 能力可能采用不同价格、身份与证据。QVeris 治理外部 API、服务与实时数据的 Discover → Inspect → 调用;二者只在父工作流层合并账本。
Current Gemini price points to test 当前值得测试的 Gemini 价格点
Official Gemini Developer API prices verified July 21, 2026, in USD per 1 million tokens. Free-tier availability and data-use terms differ from paid production use. Batch or Flex processing changes latency and availability, so compare it only with workloads that tolerate those constraints.
以下 Gemini Developer API 官方价格于 2026 年 7 月 21 日核验,单位为美元/百万 Token。免费层可用性和数据使用条款与付费生产不同;批处理或 Flex 会改变延迟与可用性,只适合可接受这些约束的工作负载。
| Model or route 模型或路由 | Input / cost 输入/成本 | Output 输出 | Scope 适用范围 |
|---|---|---|---|
| Gemini 3.1 Flash-Lite — Standard | $0.25 | $1.50 | Text/image/video input 文本/图像/视频输入 |
| Gemini 3.1 Flash-Lite — Batch/Flex | $0.125 | $0.75 | Asynchronous or flexible service 异步或弹性服务 |
| Gemini 3.5 Flash — Standard | $2.70 | $16.20 | Output includes thinking tokens 输出含思考 Token |
When evaluating Cheapest Gemini API, verify the linked official pricing source immediately before a buying decision. Model quality, output length, cacheability, retries, tools, service tier, region, taxes, and discounts can reverse a token-price comparison.
评估“最便宜的 Gemini API”时,采购决策前务必立即核验页面所链接的官方定价来源。模型质量、输出长度、可缓存性、重试、工具、服务层、区域、税费与折扣都可能逆转 Token 单价比较。
Build an Effective Cost Model for the cheapest Gemini API route为成本最低的 Gemini API 路径建立有效成本模型
When evaluating Cheapest Gemini API, list price is an input, not the decision. Compare the cost of an accepted production outcome after quality, retries, latency, operational work, and non-token charges are included.
评估“最便宜的 Gemini API”时,目录价只是输入,而不是最终决策。应在计入质量、重试、延迟、运营工作和非 Token 费用后,比较获得一个合格生产结果的成本。
Sample production-shaped tasks and record Gemini model tiers, multimodal input, context thresholds, caching, batch jobs, grounding or tool charges, quotas, and cloud versus developer access. Use percentiles and task classes rather than one average prompt so long-context and output-heavy requests remain visible.
抽取接近生产形态的任务,并记录Gemini 模型层级、多模态输入、上下文阈值、缓存、批处理、Grounding 或工具费用、配额,以及 Cloud 与开发者入口。使用分位数和任务类别,而不是单个平均提示词,确保长上下文与高输出请求不会被隐藏。
When evaluating Cheapest Gemini API, measure completion, rubric score, structured-output validity, tool accuracy, retry rate, and human-review time. Divide total run cost by accepted outcomes, not by raw requests.
评估“最便宜的 Gemini API”时,对每个模型与路由衡量完成率、评分标准得分、结构化输出有效率、工具准确率、重试率和人工复核时间,并用总运行成本除以合格结果,而不是原始请求数。
When evaluating Cheapest Gemini API, add gateway or aggregator fees, storage, egress, cache writes, evaluations, observability, support, engineering, incident handling, reserved capacity, unused credits, taxes, and the business cost of latency or failed work.
评估“最便宜的 Gemini API”时,加入网关或聚合费用、存储、流量、缓存写入、评估、可观察性、支持、工程、事故处理、预留容量、未用额度、税费,以及延迟或任务失败的业务成本。
When evaluating Cheapest Gemini API, store source URL, retrieval date, currency, region, service tier, thresholds, discounts, and model version. Recompute scenarios when a catalog changes and alert when observed invoice cost diverges from the estimate.
评估“最便宜的 Gemini API”时,保存来源链接、获取日期、币种、区域、服务层级、阈值、折扣和模型版本。目录变化时重新计算情景,并在实际发票成本偏离估算时发出告警。
FAQ
It depends on the exact model, modality mix, tools, tier, region and workload.
Not for multimodal or tool-using workloads; include every billed dimension.
Yes indirectly: queueing, failures and retries can change delivered-work cost.
取决于准确模型、模态组合、工具、层级、区域与负载。
对多模态或工具负载不可以;必须纳入所有计费维度。
会间接影响:排队、失败与重试会改变有效工作成本。