<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>schristoph.online</title><link>https://schristoph.online/tags/openweights/</link><description>Personal homepage and blog of Stefan Christoph</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><copyright>Stefan Christoph. All rights reserved.</copyright><lastBuildDate>Thu, 08 Oct 2026 08:30:00 +0200</lastBuildDate><atom:link href="https://schristoph.online/tags/openweights/index.xml" rel="self" type="application/rss+xml"/><item><title>GLM 5.3 on Amazon Bedrock: An Open-Weight Coding Model Behind a Managed API, Tested</title><link>https://schristoph.online/blog/glm-5-3-on-amazon-bedrock/?utm=rss-feed</link><pubDate>Thu, 08 Oct 2026 08:30:00 +0200</pubDate><guid>https://schristoph.online/blog/glm-5-3-on-amazon-bedrock/</guid><description>&lt;div class="tldr" data-pagefind-weight="5" data-pagefind-meta="tldr" style="display:block;font-size:.875em;margin:2rem 0;border-left:4px solid #ccc;padding-left:1rem;line-height:1.5;">&lt;strong>TL;DR:&lt;/strong> GLM 5.3 from Z.ai is a mixture-of-experts coding model with roughly 750B total and about 40B active parameters, a 1M-token context and always-on reasoning, and it is now on Amazon Bedrock for eligible customers through US and Global cross-Region inference [1] [3] [7]. That gives you a frontier-class open-weight model without running a GPU fleet for about 750 GB of weights. In my tests (one account, small N, self-written tasks) at their default settings both models passed every one-shot task and every agent run, but at its default &lt;code>max&lt;/code> reasoning effort it was slower and, on one-shot tasks, more expensive per task than Opus despite a lower per-token price. In the effort sweep on one coding task, &lt;code>reasoning_effort&lt;/code> produced the largest latency and output-token differences I measured: &lt;code>low&lt;/code> cut latency about 35x and output tokens about 28x on a coding task while still passing its tests, though it dropped some edge cases elsewhere. Prompt caching worked on both models; GLM also cached implicitly without being asked. Before adopting, check access, where your requests may be processed, the license on the weights, and how you will bound reasoning and agent runaways.&lt;/div>
&lt;p>In 2025, when OpenAI published its first open-weight models, I wrote that models are becoming the engine of the car, and that an engine you can get both as weights and as a managed service is a good place to be [12]. GLM 5.3 is a much larger engine of that kind. Z.ai published the weights on Hugging Face, and on October 5, 2026 AWS made the model available on Amazon Bedrock [1] [2]. I wanted to know three things: what this changes for teams on AWS, what you should check before you build on it, and how it actually behaves next to Claude Opus 5.5 when I point both at the same coding work in my own account.&lt;/p></description></item></channel></rss>