Text Generation
Transformers
Safetensors
glm4_moe_lite
conversational
🇪🇺 Region: EU

q4 +mtp

#1
by victor11sk - opened

any plans to quant and add mtp?

MTP or some alternative solution would be perfect for GLM 4.7 Flash model. Honestly, I loved the original model, but it was too slow for me despite being a MoE. It's been a while so maybe the inference was slightly improved since then, but honestly when done just right, prediction models can make a big difference in speed, so I usually look for alternatives that come with them already bundled in.

Sign up or log in to comment