AI工具

beellama.cpp

anbeeld/beellama.cpp

KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM