Home › Projects › BeeLlama Recipes
Recipes for various models with BeeLlama.cpp for VRAM budgets of consumer GPUs.
Each recipe is a complete serving setup for one model and one VRAM budget. Inside are the weight quant, context size, KV cache types, ubatch size, and the llama-server command to run it. The cache types come from the data gathered from the KV cache quantization benchmarks, focusing on best possible quality for the budget. Click a row to open its recipe.
| VRAM | Model & Quant | Context Size | Spec Type | K Cache Type | V Cache Type | Precision Tail | Multimodal | Ubatch |
|---|---|---|---|---|---|---|---|---|
| 24 GB | Qwen 3.8 27B / UD-Q4_K_XL / Q4_K_M | 256K / 262144 | MTP | 5-bit / kvarn5 |
5-bit / kvarn5 |
1024 | Yes | 512 |
| 24 GB | Qwen 3.8 27B / UD-Q4_K_XL | 256K / 262144 | DFlash 2 | 4-bit / kvarn4 |
4-bit / kvarn4 |
1024 | Yes | 512 |