llama.cpp adds GLM-5.3-Flash support
The work also covers a large-context softmax crash and a data race that may affect the full model.
By tensorThe work also covers a large-context softmax crash and a data race that may affect the full model.
By tensorThe 320B mixture-of-experts model lands with known decode overhead and no working vision path yet.
By tensor