AnbeeldAnbeeld

Projects Articles SupportContact X

Anbeeld's BeeLlama Recipes

GitHub

Recipes for various models with BeeLlama.cpp for VRAM budgets of consumer GPUs.

Each recipe is a complete serving setup for one model and one VRAM budget. Inside are the weight quant, context size, KV cache types, ubatch size, and the llama-server command to run it. The cache types come from the data gathered from the KV cache quantization benchmarks, focusing on best possible quality for the budget. Click a row to open its recipe.

Support my work!

Recipes

VRAM Model & Quant Context Size Spec Type K Cache Type V Cache Type Precision Tail Multimodal Ubatch
24 GB Qwen 3.8 27B / UD-Q4_K_XL / Q4_K_M 256K / 262144 MTP 5-bit / kvarn5 5-bit / kvarn5 1024 Yes 512
24 GB Qwen 3.8 27B / UD-Q4_K_XL 256K / 262144 DFlash 2 4-bit / kvarn4 4-bit / kvarn4 1024 Yes 512
Back to projects