GLM-5.3-Flash, Z.ai's first natively multimodal model in the GLM-5 series, is now available through DigitalOcean Inference Engine. A hybrid sparse-and-linear attention architecture makes its 1M-token context window substantially cheaper to serve, delivering GLM-5.2-class-or-better coding and agentic performance at roughly one-tenth the cost, under the MIT License.