Popular repositories Loading
-
llama-cpp-mtp-turboquant-sm120-blackwell-windows
llama-cpp-mtp-turboquant-sm120-blackwell-windows PublicWindows prebuilt of llama.cpp combining Multi-Token Prediction (MTP) + TurboQuant KV cache compression + native sm_120 (Blackwell consumer GPU, FP4 tensor cores). For RTX 5060 Ti / 5070 / 5080 / 5090.
-
beellama.cpp
beellama.cpp PublicForked from Anbeeld/beellama.cpp
DFlash & TurboQuant in llama.cpp with up to 3x faster generation and 7.5x more KV cache in same VRAM
C++ 1
-
knowghost
knowghost PublicAI interview assistant — stealth desktop app with speech recognition, LLM chat, vision, and card storage
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.