Serving the Complete GLM-5.2 on One 96GB Blackwell GPU
Built a hybrid serving path for the complete 744B-A40B model. Corrected and validated a group-size-128 int4 MoE kernel, kept the 374GB expert bank resident in CPU RAM, and profiled real routing behavior to prove why hot-expert caching could not rescue an undersized memory tier.
Co-founder @ Flarebase
Founder University 10