LLMKube
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.
Health breakdown: 89/100
Five terms, recomputed nightly from the GitHub API. How this is calculated.
- Commit activity 35.0 / 35
534 commits in 90 days (30+ scores full marks)
- Release recency 25.0 / 25
v0.9.27 released today
- Stars 9.3 / 20
210 GitHub stars (logarithmic)
- Not archived 10.0 / 10
Repository is active
- Docker support 10.0 / 10
Docker image or Compose file available
Commit activity
- 534
- Last 90 days
- 909
- Last 12 months
- today
- Last push
Ranked #12 of 14 tools in Generative Artificial Intelligence (GenAI) by health.
Details
| Licence | Apache-2.0 |
| Platforms | Go, Docker, K8S |
| Latest release | v0.9.27 · today |
| Stars | 210 |
| Docker | Yes |
| Repository | defilantech/LLMKube ↗ |