GLM-5.3-Flash - The first natively multimodal model in GLM-5 series

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

Add a comment

Replies

Best

Hi everyone!

So was GLM-5.3-Flash 🐂

People were already running it on and OpenRouter before the name showed up. And it was the most used model of the week there:

The public version is 320B-A18B w/ native multimodal, 1M context, . Sparse plus linear attention is how they keep that long window cheaper to serve.

API is $0.15/M in and $0.50/M out. First two weeks are !

 The 1M context window at that price is definitely interesting for long tasks.

So, Ox Alpha is GLM-5.3-Flash!? lfg. been running it with for a few days. it's fast, holds up on long agentic runs, and it keeps finishing tasks in no time. the new default.

S/O to !

I wonder when GLM models will finally learn to understand/read images sent to them.