Introducing quantized Llama models with increased speed and a reduced memory footprint (via friendlysock) — discussion

#ai #performance