A developer has built Colibri, a lightweight inference engine that can run Z.ai’s GLM-5.2, a 744-billion-parameter Mixture-of-Experts model, on consumer hardware without an Nvidia GPU. The project is ...