If Your Model Inference is Slow, MOE Can Fix it
Author(s): saniya jaswani Originally published on Towards AI. “Mixture of Experts makes model inference faster. To scale request volume, MoE optimizes token routing.” If you read about all AI developments, GPT-4 is rumored to use ~1.8 trillion parameters. Mixtral 8x7B punches well …