llm
Description
A self-contained C module for LLM inference on top of ggml. No external dependencies besides the ggml library. Runs on CPU and on Vulkan GPU (NVIDIA / AMD / Intel).
The project has two parts: llm.h / llm.c - the inference module (GGUF loading, tokenizer, chat templates, tool calls, KV-cache reuse, generation); test_llm.c - a console demo (test_llm.exe): a one-shot prompt or an interactive chat on CPU or GPU.
Technologies
- C - the llm.h/llm.c module + test_llm.c console demo
- ggml - tensor library (CPU + Vulkan backend)
- GGUF - model format
- OpenMP - CPU parallelization
- License GNU AGPL v3 (Affero GPL)