GLM 5.3 on Amazon Bedrock: An Open-Weight Coding Model Behind a Managed API, Tested
written by Stefan Christoph
- 21 minutes readmax reasoning effort it was slower and, on one-shot tasks, more expensive per task than Opus despite a lower per-token price. In the effort sweep on one coding task, reasoning_effort produced the largest latency and output-token differences I measured: low cut latency about 35x and output tokens about 28x on a coding task while still passing its tests, though it dropped some edge cases elsewhere. Prompt caching worked on both models; GLM also cached implicitly without being asked. Before adopting, check access, where your requests may be processed, the license on the weights, and how you will bound reasoning and agent runaways.In 2025, when OpenAI published its first open-weight models, I wrote that models are becoming the engine of the car, and that an engine you can get both as weights and as a managed service is a good place to be [12]. GLM 5.3 is a much larger engine of that kind. Z.ai published the weights on Hugging Face, and on October 5, 2026 AWS made the model available on Amazon Bedrock [1] [2]. I wanted to know three things: what this changes for teams on AWS, what you should check before you build on it, and how it actually behaves next to Claude Opus 5.5 when I point both at the same coding work in my own account.
This post is long. It has three parts: implications, a practical guide with code, and my measured results. If you only want the numbers, jump to the hands-on section.
What GLM 5.3 is
The AWS model card describes GLM 5.3 as Z.ai’s flagship open-weight text model for complex software engineering and agentic tasks, built on the same base model as GLM 5.2, with the improvements coming from scaled post-training [3]. The facts that matter for planning:
| Property | Value | Source |
|---|---|---|
| Architecture | Mixture of experts, about 40B active parameters per token | [1] [3] |
| Total parameters | 753,329,940,480 per the Hugging Face weights metadata; the AWS What’s New post and launch blog say 753B, the AWS model card says 744B | [1] [2] [3] [7] |
| Context window / max output | 1M tokens / 128K tokens | [3] |
| Reasoning | Always on; reasoning_effort accepts low, high, max, default max | [1] [7] |
| Tools and output | Function calling, structured JSON output, streaming including streamed reasoning | [3] |
| Bedrock APIs | Converse, Invoke, Chat Completions, Responses on bedrock-runtime | [3] |
| Inference | Cross-Region only: us.zai.glm-5.3 and global.zai.glm-5.3 | [3] |
| Service tiers | Standard, Priority, Flex (Reserved not offered for this model) | [3] |
| Access | Limited to eligible customers, via your AWS account team | [1] [3] |
The parameter count deserves one sentence of care because three AWS pages disagree by 9B. I went to the source: the safetensors metadata of the published weights adds up to 753.3B parameters, of which about 751B are stored in FP8 [7]. “About 750B” is the honest round number.
On capability, I am not going to restate leaderboards. Z.ai reports open-source state of the art on Terminal Bench 3.0 and a 50% improvement over GLM 5.2 on its in-house coding benchmark [7]. Those are Z.ai’s claims. My own measurements come later, and they answer a narrower question.
Part 1: What it changes
A frontier open-weight model without an inference fleet
The weights are public, but at about 751B FP8 parameters they are roughly 750 GB before you count the KV cache for a 1M-token context [11]. Serving that yourself is a multi-GPU, multi-node exercise with its own capacity planning, patching and on-call. Bedrock removes that: you call an API and pay per token, with no persistent resources to clean up [2]. The launch blog frames it the same way, saying that using open-weight models for demanding coding work has historically meant provisioning and operating your own inference infrastructure [2].
The practical effect is that an open-weight model becomes one more option in your model portfolio. You can evaluate it with the same IAM roles, the same guardrails and the same observability you already use for other Bedrock models. Guardrails, model evaluation, prompt management, flows, agents and structured outputs are listed as supported for GLM 5.3; intelligent prompt routing, count tokens and knowledge bases are not [3].
Access is gated
Access is limited to eligible customers and goes through the AWS account team [1] [3]. Plan for that step if you want to test it. The test account I used had both inference profiles active.
Cross-Region inference only, and what that means for where requests run
GLM 5.3 has no in-Region endpoint. You call it from a source Region through one of two inference profiles [3]:
us.zai.glm-5.3(geographic): available as a source from four US Regions, and requests are routed within US Regions.global.zai.glm-5.3(global): available as a source from 31 Regions in the model card, including the US, Canadian, European, Asia Pacific, Israel, Africa, Mexico and South America Regions, and requests can be processed in any supported commercial Region.
There is no EU geographic profile for this model today, so a team in Frankfurt or Zurich reaches it through the Global profile [3]. The cross-Region inference guide spells out the trade-off: geographic profiles keep processing within the geography and are the choice for data residency requirements, while global profiles route worldwide and cost about 10% less [5]. Two operational details are easy to miss: CloudTrail records each cross-Region request in your source Region with an additionalEventData.inferenceRegion field that shows where it was processed, and service control policies must allow the destination Regions (for global profiles, aws:RequestedRegion of unspecified) [5]. If you have residency requirements, decide which profile your organization accepts before the first prototype, not after.
Always-on reasoning puts tokens and latency into the cost model
Reasoning cannot be switched off. What you can choose is how much of it you buy. The Hugging Face card says the default is max, that you must pass low or high explicitly, and that any other value falls back to max [7]. There is no medium, so an effort setting copied from another model can silently buy you the most expensive level. That default matters more than any price-per-token comparison, as my results show.
Prompt caching for long-context agent loops
Bedrock supports both implicit and explicit prompt caching for GLM 5.3, with a minimum of 1,024 tokens per cache checkpoint and a retention of at least 30 minutes [3]. The model card recommends explicit cache controls because they can raise the hit rate [3]. For agents that resend a large system prompt, tool definitions and repository files every turn, this is where most of the input cost goes, which is the same prefill arithmetic I walked through in an earlier post on token costs [11].
Low switching cost through OpenAI-compatible APIs
Chat Completions and Responses are supported on the bedrock-runtime endpoint, and the launch blog recommends them for new applications because they expose the most complete feature set for this model [2] [3]. If your client code already speaks the OpenAI protocol, trying GLM 5.3 is a base URL, a credential and a model ID. That is my opinion of the switching cost, grounded in the API table [3]; your prompts and tool schemas still need testing.
Lifecycle and license
The model card states at least 45 days of EOL notice and a legacy period of at least 45 days [3]. On licensing, read two documents. The Bedrock model card links its end user license to the Apache 2.0 license in Z.ai’s GLM-5 GitHub repository [3]. The weights on Hugging Face carry a separate “GLM-5.3 License” [7] [8]. It is a permissive grant (use, modify, distribute, sublicense, fine-tune) with one notable condition: an organization that operates a “Model as a Service” business and has more than USD 10 billion in revenue over 12 months must pass a Z.ai security review before commercial use of the weights [8]. Whether and how that matters for your use is a question for your legal team, not for a blog post.
Part 2: A practical guide
Pricing at a glance
Standard-tier list prices in USD per 1M tokens from the Bedrock pricing page, fetched on 2026-10-07 [4]:
| Model and profile | Input | Output | Cache read | Cache write |
|---|---|---|---|---|
| GLM 5.3, US profile | 1.848 | 5.808 | 0.3432 | 2.31 (30 min) |
| GLM 5.3, Global profile | 1.68 | 5.28 | 0.312 | 2.10 (30 min) |
| Claude Opus 5.5, US profile | 4.40 | 22.00 | 0.22 | 5.50 (5 min) |
Per token, GLM 5.3 input is about 2.4x cheaper than Opus 5.5 and output about 3.8x cheaper. For GLM 5.3, the pricing page states that Priority is a 75% premium and Flex and Batch a 50% discount on Standard [4]. Two caveats make the per-token table misleading on its own, and both come from my measurements below: the two models tokenize the same text very differently, and reasoning tokens at the default effort can dominate the bill.
First call with Converse
The Converse request shape is the one from the model card [3]. To pick the reasoning effort through Converse, I passed it in additionalModelRequestFields; the field name is from the Invoke sample in the model card and the Hugging Face card [3] [7], and it worked through Converse in my account (the shorter answers and lower token counts confirmed it took effect).
import boto3
client = boto3.client("bedrock-runtime", region_name="us-east-1")
resp = client.converse(
modelId="us.zai.glm-5.3",
messages=[{"role": "user", "content": [{"text": "Refactor this function to be iterative: ..."}]}],
additionalModelRequestFields={"reasoning_effort": "low"}, # low | high | max (default)
)
for block in resp["output"]["message"]["content"]:
if "reasoningContent" in block:
print("reasoning:", block["reasoningContent"]["reasoningText"]["text"][:200])
elif "text" in block:
print(block["text"])
print(resp["usage"])
The response carries a reasoningContent block before the answer, and outputTokens in usage includes the reasoning tokens.
Chat Completions with short-lived credentials
The launch blog recommends short-lived tokens generated from your normal AWS credentials over long-lived API keys [2]:
from aws_bedrock_token_generator import provide_token
from openai import OpenAI
region = "us-east-1"
client = OpenAI(api_key=provide_token(region=region),
base_url=f"https://bedrock-runtime.{region}.amazonaws.com/openai/v1")
resp = client.chat.completions.create(
model="us.zai.glm-5.3",
messages=[{"role": "user", "content": "Explain this stack trace: ..."}],
reasoning_effort="low",
)
print(resp.choices[0].message.content)
The Responses API worked the same way with reasoning={"effort": "low"}.
Prompt caching
With Converse, a cachePoint block marks the end of the reusable prefix [6]:
system = [
{"text": REPOSITORY_CONTEXT}, # large and stable: files, tool docs, conventions
{"cachePoint": {"type": "default"}}, # cache everything above this point
]
resp = client.converse(modelId="us.zai.glm-5.3", system=system,
messages=[{"role": "user", "content": [{"text": question}]}])
u = resp["usage"]
print(u["cacheWriteInputTokens"], u["cacheReadInputTokens"])
With the Responses or Chat Completions API, the launch blog shows extra_body={"prompt_cache_options": {"mode": "explicit"}} plus a "prompt_cache_breakpoint": {"mode": "explicit"} on the content blocks that end a reusable prefix [2]. One documentation gap to know about: the cross-model table of explicit caching limits in the prompt caching guide did not list GLM 5.3 when I checked [6]; the model card and the launch blog are the authoritative sources for this model [2] [3].
Service tiers
Converse takes a serviceTier parameter in the current SDK. A Flex request is one line:
client.converse(modelId="us.zai.glm-5.3", messages=msgs, serviceTier={"type": "flex"})
The service tier guide describes Flex as the discounted tier for workloads that tolerate longer processing, such as evaluations and agentic workflows, and Priority as the premium tier for latency-critical requests [9]. My single Flex test call succeeded; I did not measure Flex latency or queueing behavior beyond that one call.
Part 3: Hands-on results
Setup
Everything below ran on 2026-10-07 in one AWS account, from us-east-1, through the US inference profiles of both models, on the Standard tier (Experiment 1 additionally tried the Global profile and the Flex tier). The comparison model is Claude Opus 5.5 at its default effort, which its model card gives as medium [10]. GLM 5.3 ran at its default (max) unless I say otherwise. That is a defaults-versus-defaults comparison, and it is deliberate: it is what you get if you swap the model ID and change nothing else.
The coding tasks and the agent’s seed repository are self-written and unpublished, which lowers but does not rule out the chance that either model has seen something equivalent in training. The long-context corpora are CPython 3.9 standard library sources (PSF license). I graded everything with hidden unit tests and smoke-tested each test suite against my own reference solutions and against a deliberately broken one, to reduce the risk that a zero score reflects a grader bug rather than the model. Costs are computed from the usage of every call and the list prices above. All experiments together cost USD 17.57 at list prices.
The sample sizes are small. Treat every number here as “in my tests, N as stated”, not as a benchmark.
Experiment 1: every API surface worked
Converse on both profiles, Converse with reasoning_effort, InvokeModel with the chat-style body from the model card, Chat Completions and Responses through the OpenAI SDK, and Converse with serviceTier Flex all returned answers on the first try. A trivial prompt took 1.3 seconds on Converse with the default effort, and 0.7 seconds with low.
Experiment 2: reasoning effort is the cost and latency dial
I gave GLM 5.3 one coding task three times at each effort level: implement an arithmetic expression evaluator with precedence, right-associative power, unary minus and error handling, graded by 16 hidden tests.
| Effort | N | Tests passed | Wall time per call (s) | Output tokens | Cost per call (USD) |
|---|---|---|---|---|---|
| low | 3 | 16/16 every run | 4.6 (3.7β7.5) | 1,014 (819β1,055) | 0.006 |
| high | 3 | 16/16 every run | 18.0 (16.9β22.8) | 2,716 (2,622β3,942) | 0.016 |
| max (default) | 3 | 16/16 every run | 162 (131β318) | 28,053 (21,336β41,959) | 0.163 |
Median (minβmax).
In my tests every run passed all 16 tests; max effort bought time and tokens, not correctness, on this task.
On this task, max bought nothing measurable: every run at every level passed. It cost about 35x the wall time and 28x the output tokens of low, and at the extreme a single answer took over five minutes. That does not mean low is always enough; Experiment 5 shows where it is not. It does mean that leaving the default in place for routine calls is an expensive choice you should make on purpose.
Experiment 3: prompt caching, and a tokenizer surprise
I put about 250,000 characters of CPython asyncio source into the system prompt, added a cache point, and asked three different questions per trial (one cold call, two warm calls), three trials per model with a fresh prefix each time. As a control I sent the same prefix twice without a cache point.
| Model | Call | Prompt tokens (written / read from cache) | Wall time (s) | Cost (USD) |
|---|---|---|---|---|
| GLM 5.3 | cold | 51,612 written | 3.5 (3.4β3.5) | 0.120 |
| GLM 5.3 | warm | ~51,586 read | 2.0β3.5 (1.6β8.3) | 0.018 |
| GLM 5.3 | no cache point, 2nd call | 51,580 read | 3.7 | 0.018 |
| Claude Opus 5.5 | cold | 82,382 written | 7.1 (6.5β7.6) | 0.457 |
| Claude Opus 5.5 | warm | 82,382 read | 4.9β6.8 (4.8β7.3) | 0.023β0.026 |
| Claude Opus 5.5 | no cache point, 2nd call | 0 read, 82,432 uncached | 6.9 | 0.369 |
First, caching cut the cost of a warm call by about 85% for GLM 5.3 and about 95% for Opus 5.5 in my runs. The latency effect was small and noisy at these output lengths; the cost effect was not.
Second, GLM 5.3 cached implicitly. Without any cache point, its second call read the whole prefix from cache, and the first call reported the prefix as cache-written tokens. I priced those at the listed cache-write rate, which makes a one-off long prompt about 25% more expensive than if it had been billed as plain input; verify against your own bill before you rely on that number. Opus 5.5 did not cache anything without a cache point in my runs.
Third, the same text is not the same number of tokens. The identical system prompt was 51,612 tokens for GLM 5.3 and 82,382 for Opus 5.5, a ratio of 1.6 on Python source. A per-token price comparison therefore understates the gap on input-heavy work by that factor, at least for code like this. Measure on your own data before you build a cost model.
Experiment 4: an agent fixing bugs
This is the experiment I cared about most. I wrote a small Python package (a toy ledger with money parsing, transfers, compound interest and a statement report), seeded it with five bugs, and gave each model four tools through Converse tool use: list files, read a file, write a package file, run the tests. The tests were read-only. Each run had at most 25 model turns, and the final score was a fresh run of the original tests against the agent’s code. At the start, 8 of the 16 tests failed.
| Arm | Runs | Tests passed per run | Turns | Wall time (s) | Cost per run (USD) |
|---|---|---|---|---|---|
| Claude Opus 5.5, default effort | 3 | 16, 16, 16 | 5 (5β5) | 32 (30β37) | 0.150 (0.148β0.155) |
GLM 5.3, default (max) | 3 | 16, 16, 16 | 22 (13β22) | 115 (83β1,144) | 0.226 (0.162β3.017) |
GLM 5.3, high | 3 | 16, 8*, 16β | 24, β, 25 | 39, β, 116 | 0.091, 2.251*, 0.511 |
GLM 5.3, low | 3 | 16, 12*, 16β | 19, β, 8 | 34, β, 10 | 0.072, 2.036*, 0.030 |
Median (minβmax). Turns and wall time over completed runs. * = run I aborted; score is the test result of its working copy at that point. β = run with the write-size limit and USD 1 cap added, so the high and low rows show per-run values instead of a median across different harness settings.
Agentic bug fixing, three runs per arm: medians are close, GLM has a long tail.
Both models fixed all five bugs whenever they ran to completion, and their final summaries described the same root causes. The differences were in how they got there:
- Opus 5.5 was consistent. Five turns every time, about half a minute, about 15 cents. It batched about three tool calls per turn, so its low turn count partly reflects style.
- GLM 5.3 worked in smaller steps. It took 8 to 25 turns, usually one or two tool calls per turn, so compare cost and time rather than turns. At the default effort its median cost was USD 0.226 against USD 0.150 for Opus, about 50% higher. Among completed runs, GLM at
lowhad a lower median cost and latency than Opus, butlowcompleted only two of three runs, the comparison is not effort-matched (Opus ran at its default), and only GLM benefited from implicit caching. - GLM 5.3 had a long tail. One default-effort run took 19 minutes and cost USD 3.02. It produced 244,000 output tokens (about USD 1.42); the rest was input, because every turn re-sent a growing context. I did not log per-turn reasoning in this run, so I cannot say how much of that was reasoning. For comparison only: in Experiment 2, reasoning made up 92β96% of the characters of a
maxanswer. And two of six runs atlowandhighwent off the rails: in one, awrite_filecall produced a 178,614-byteledger.pyconsisting of the same two lines repeated thousands of times, and from then on every turn re-sent a context of more than 100,000 tokens at about 30 cents a turn. I stopped both runs by hand. For the last run at each level I added a 20,000-character limit on written files and a USD 1 per-run cap, which any production agent harness should have anyway. Both of those runs completed with all tests passing and under the cost cap.
Across all nine GLM runs, three were cheaper than every Opus run, two were within about eight cents of Opus, and four cost between 3x and 20x more. With N=3 per arm I cannot tell you how often such runs happen, only that they happened in a small sample, and that the guards that stop them are cheap.
I also noticed something about caching here: the four completed GLM agent runs I logged in full read roughly 80,000 to 225,000 tokens from cache each without a single cache point in my code, because of the implicit caching from Experiment 3. The Opus loops read none, because I did not add cache points. That asymmetry favors GLM in the cost column; in a real agent you would add cache points for Opus, which would lower its cost further.
Experiment 5: six one-shot tasks
Six self-written tasks with hidden tests (expression evaluator, interval merging, semantic-version ordering, an LRU cache with TTL, a CSV line splitter, Roman numerals with strict validation), three repeats each, 18 calls per arm.
| Arm | Assertions passed | Task-runs fully solved | Wall time per call (s) | Output tokens per call | Cost per call (USD) |
|---|---|---|---|---|---|
GLM 5.3, default (max) | 198/198 (100%) | 18/18 | 55 (17β188) | 14,176 (3,319β29,935) | 0.083 (0.020β0.174) |
| Claude Opus 5.5, default | 198/198 (100%) | 18/18 | 12 (5β24) | 1,022 (357β2,139) | 0.023 (0.009β0.049) |
GLM 5.3, low | 185/198 (93%) | 12/18 | 2.3 (0.9β8.6) | 431 (140β1,412) | 0.003 (0.001β0.009) |
Six one-shot tasks, three repeats each, 18 calls per arm.
At the defaults, both models solved everything. GLM 5.3 needed about 14x more output tokens and 4.6x more time per task, and cost about 3.6x more per task than Opus 5.5, despite the lower list price. At low, GLM 5.3 was about 7x cheaper than Opus and very fast, but it missed edge cases: the strict validation rules of the Roman numeral task (two of three runs), the “characters after a closing quote” rule of the CSV splitter (all three), and one cache-eviction case. In this sample max passed all of them; the experiment was not designed to show that the extra reasoning is what caught them.
A fair reading: on short, well-specified coding tasks in my sample, Opus 5.5 at its default effort was the most efficient way to get a fully correct answer, and GLM 5.3 matched its correctness when allowed to think at length. How the two compare at matched effort is a different experiment that I did not run.
Experiment 6: one long-context lookup
I planted a small function with an unusual constant at about 60% depth in roughly 1.4 million characters of CPython standard library source and asked for its value and file. One call per model.
| Model | Prompt tokens | Correct | Wall time (s) | Cost (USD) |
|---|---|---|---|---|
GLM 5.3 (low) | 328,113 | yes | 22.7 | 0.76 |
| Claude Opus 5.5 | 512,894 | yes | 15.0 | 2.26 |
Both found it. One needle at one depth is a smoke test of long-context retrieval, not an evaluation of it. The cost line repeats the tokenizer effect: the same input was 56% more tokens for Opus.
What to check before you adopt it
- Access and profile. Ask your account team about eligibility. Decide between the US and the Global profile with your compliance team, because that decides where requests can be processed. Check your SCPs.
- Set
reasoning_efforton every call. Do not inheritmaxby accident. Start atloworhigh, measure on your own tasks, and raise it where you see quality gaps. - Measure cost per task, not per token. Tokenizer differences and reasoning tokens decided the cost ranking in my tests, not the price list.
- Bound your agents. Per-run cost caps, turn limits, output size limits on file writes, and an alarm on context size. My two runaway runs would each have been caught by a 20 KB write limit or a USD 1 cap.
- Use cache points. Implicit caching helped GLM 5.3 automatically, but explicit cache points are the recommended path [3] and they let you control what is cached.
- Run your own evaluation. Six tasks and one toy repository are a starting point. Use Bedrock model evaluation or your own harness on your own code before you switch production traffic.
- Read the licenses that apply to how you use it: the Bedrock end user terms and, if you ever use the weights directly, the GLM-5.3 License [3] [8].
Closing
GLM 5.3 on Bedrock is a real option for coding and agent workloads, available without operating the infrastructure a model of this size needs. In my tests it was as correct as Claude Opus 5.5 at its default settings, and its price per token is lower, but its default reasoning effort spends that advantage and more. The teams that get value from it will be the ones that treat reasoning_effort as a configuration they tune per workload, put guardrails on their agents, and measure cost per solved task.
Which workload would you test first: a long-context code review, a bulk refactoring job on Flex, or an agent loop with cache points?
AI attribution: this post was drafted with AI assistance and reviewed, edited, and fact-checked by the author. All claims are sourced; the opinions and final wording are mine. The experiments ran in my own AWS account.
Sources
- [1] AWS What’s New: GLM 5.3 by Z.ai is now generally available on Amazon Bedrock, October 5, 2026: 753B total and about 40B active parameters, always-on reasoning with selectable effort, explicit prompt caching, eligible customers, US and Global profiles.
- [2] AWS Machine Learning Blog: Introducing GLM 5.3 on Amazon Bedrock, October 5, 2026: OpenAI-compatible APIs, short-lived tokens, explicit caching with
prompt_cache_options. - [3] Amazon Bedrock User Guide: GLM 5.3 model card: capabilities, APIs, caching limits, service tiers, regional availability, lifecycle, sample code.
- [4] Amazon Bedrock pricing, fetched 2026-10-07: GLM 5.3 and Claude Opus 5.5 on-demand prices and tier premiums and discounts.
- [5] Amazon Bedrock User Guide: Cross-Region inference: geographic versus global profiles, CloudTrail
inferenceRegion, SCP requirements. - [6] Amazon Bedrock User Guide: Prompt caching: implicit and explicit caching,
cachePointin Converse, billing of cached tokens. - [7] Hugging Face: zai-org/GLM-5.3 model card and weights:
reasoning_effortlevels and default, Z.ai’s benchmark claims, safetensors parameter counts. - [8] GLM-5.3 License on Hugging Face: permissive grant and the Model-as-a-Service security review condition.
- [9] Amazon Bedrock User Guide: Service tiers: Reserved, Priority, Standard and Flex.
- [10] Amazon Bedrock User Guide: Claude Opus 5.5 model card: always-on adaptive thinking, effort levels, default
medium. - [11] Stefan Christoph, “Why AI Tokens Are So Expensive, and What Actually Makes Them Cheaper”: prefill, decode and why prompt caching changes the input bill.
- [12] Stefan Christoph, “WOW! Yesterday OpenAI released two open-weight models”: the model-as-engine analogy and open-weight models on managed AWS services.
About the Author
Stefan Christoph is a Principal Solutions Architect at AWS, focused on agentic AI, media & entertainment, and helping builders move from demo to production. He writes about AI architecture, developer productivity, and the future of software.
This is a personal blog. Opinions expressed here are my own and do not represent the views or positions of my employer.
β€οΈ Created with the support of AI (Kiro)