Alibaba Singapore cache pricing review¶
Editions: OSS, Cloud, Enterprise. Unless stated otherwise, everything on this page ships in OSS.
Retrieved beginning 2026-09-15T15:58:12Z from first-party public pages. These are list-price estimates, not invoice costs. No account credentials or private catalog calls were used in this review.
Sources:
- https://www.alibabacloud.com/help/en/model-studio/model-pricing
- https://www.alibabacloud.com/help/en/model-studio/context-cache
- https://www.alibabacloud.com/help/en/model-studio/list-models
The pricing page lists Singapore International qwen3.8-flash,
0<Token≤1M, input $0.15 and output $0.47 per million tokens.
The context-cache page names exact supported models separately for explicit
and implicit cache, and separately by serving region/deployment scope.
This change applies its ratios only to exact identifiers present in the
Singapore International supported-model lists and the existing USD seed.
It does not infer support from a family prefix or another region.
The explicit-cache Billing section states:
Cache creation: Content used to create a new cache is billed at 125% of the standard input token price.
Cache hit: Billed at 10% of the standard input token price.
Exception: The explicit cache hit price for qwen3.8-max, qwen3.8-max-0902, qwen3.8-flash, and qwen3.8-2.4t-a95b is not 10% of the standard input token price. For specific pricing, see the Model Studio console.
The generic creation formula would yield $0.1875, but the model-specific Singapore console quote below establishes $0.20 and supersedes that formula for Flash. The public page alone does not establish Flash's hit rates.
Operator-confirmed Singapore console tariff¶
Received and recorded 2026-09-15T16:26:50Z. An operator supplied the current
Alibaba Cloud Model Studio Singapore console quote for qwen3.8-flash:
| Billing dimension | USD per million tokens |
|---|---|
| Input | 0.15 |
| Output | 0.47 |
| Implicit cache read | 0.016 |
| Explicit cache read | 0.016 |
| Explicit cache creation | 0.20 |
This is operator-confirmed first-party console evidence, not a programmatic retrieval from a fabricated public URL. The linked public context-cache page points to the console for model-specific exceptions. No private account, workspace, or credential identifiers are included. This current snapshot does not establish an earlier effective date; the seed's cache evidence and reviewed feed use the receipt time as a conservative applicability boundary. Older cached rows require evidence covering their date or an explicit verified override. Uncached base prices are unchanged.
For 50,000 input tokens including 48,000 implicit cached tokens and 1,000 output tokens, the estimate is $0.001538. With explicit cache and an additional 1,000 creation tokens included in that same prompt total, it is $0.001588. Console inference capacity limits are not substituted for pricing-tier bounds. Tool-call prices are outside this token-tariff review and are not claimed as covered here.
The implicit-cache Billing section states:
For models other than deepseek-v4.1-flash, deepseek-v4-pro, qwen3.8-max, qwen3.8-max-0902, qwen3.8-flash, and qwen3.8-2.4t-a95b: The unit price of cached_token is 20% of the input_token unit price.
It specifies 10% for deepseek-v4.1-flash, and directs users to the console
for deepseek-v4-pro and the four Qwen exceptions above. The introductory
"typically 20%" text must not override these model-specific exceptions.
Public examples use other models; they do not establish Flash's rate.
For example, qwen3.7-flash is explicitly supported in Singapore and is not
an exception. Its first context tier has input $0.030 per million, implicit
read $0.006, explicit read $0.003, and creation $0.0375. Each context tier's
ratio is calculated from that tier's input price. Cache and reasoning tokens
are already included in prompt/output totals and are not added again.
The seed retains 92 model identifiers and now has explicit creation on 26
models, explicit read on 23 models, and implicit read on 28 models (including
the previously evidenced glm-5.2 rate). Flash's console-only hit exception is
now supplied by the confirmed quote above. Other console-only exceptions
remain missing until a native USD catalog or a reviewed regional feed provides
an exact price. The current Singapore list's separately deployed ZHIPU/*
models do not inherit the Alibaba-hosted GLM policy. Ambiguous newer GLM
wording is not treated as a newly verified price.
The list-models page documents Singapore's native catalog only at
https://dashscope-intl.aliyuncs.com/api/v1/models. Its US example uses
https://{WorkspaceId}.us-east-1.maas.aliyuncs.com/api/v1/models.
It does not document the corresponding Singapore workspace path. Workspace
credentials therefore stay scoped to their configured host. A trusted
reviewed regional catalog is the supported way to deliver verified prices
to Singapore workspace models without sending their keys to another host.
Unsupported regions/currencies, context lengths, missing cache dimensions, and unknown historical cache modes remain explicit pricing failures. A 92-model catalog is not a claim to cover every model, modality, deployment scope, or billing policy.
Reviewed feed entries retain their per-model effective date and provenance. Historical requests before the only available reviewed revision fail closed; the estimator does not silently use the current price or fall back to a seed for those requests. This implementation does not archive prior Alibaba tariff revisions. Native catalogs do not establish historical effective dates; native estimates continue to represent the currently observed list tariff.
After reviewed evidence expires, a fresh native tariff may replace it. If no
fresh native tariff exists, the last reviewed tariff remains an estimate and
its usage snapshot records stale: true plus the original expiry date. This
does not silently revert to an older release seed. Process restarts still use
the shipped seed until the first valid reviewed feed is received.
If a native listing omits a required cache dimension, the estimator may select
the whole verified seed tariff when the base prices, tier bounds, and every
known native cache rate agree. A single flat native row with no context bound
may use a single-tier seed only within that seed's bound. Conflicting prices
or tier policies stay unpriced; rates are never spliced across policies.
Usage provenance identifies this fallback as seed, and the console evidence
date still applies. Direct Fetch price continues to expose the native quote.
Downloaded HTML SHA-256:
- context-cache:
6e2206ad92cbcb6272d8c8d1b58e4f5a745805e5586e1180aadbbf1b980fd204 - list-models:
dbc6aa024c7954cb1c4cef74e90678de5f76ff8e999a1d913f0dba5ee24f9494 - model-pricing:
fe8c8b1e402c2ab21fb9c55b419e9fe81e34222c94961f4f73ec52a2eec17f09