SIMD/GPU pure C inference for Llama 2-family, GPT-2, Gemma-3n, and Qwen 3.x GGUF files. Faster and simpler than llama.cpp. Enhancements over parent repo are…
SIMD/GPU pure C inference for Llama 2-family, GPT-2, Gemma-3n, and Qwen 3.x GGUF files. Faster and simpler than llama.cpp. Enhancements over parent repo are 99% coded by Qwen 3.6-27b with my own lightweight custom agentic Pythonshit. No NodeJS cancer has been installed during the development. 0 money spent on cloudslop. Contains 0% ngxson - whoreson/picolm