avx2 c99 cpu-inference deep-learning from-scratch inference-engine kimi-k3 linear-attention llm llm-inference machine-learning memory-efficient mixture-of-experts moe mxfp4 quantization simd systems-programming transformer zero-dependencies