github.com · Inference Engines
Rare Pick

ExLlamaV3

Visit site
ExLlamaV3

About

ExLlamaV3 is an inference library designed for running large language models (LLMs) efficiently on consumer-grade hardware. It focuses on optimizing the execution of quantized models, particularly those in the ExLlamaV2 format, to achieve high performance with reduced memory footprint.

This tool is primarily for developers, researchers, and enthusiasts who need to deploy and experiment with large language models locally. It enables faster inference speeds and allows for the use of larger models on systems with limited GPU VRAM, making advanced AI capabilities more accessible.

Used by 1 person

More in Inference Engines

See all
Rare Pickgithub.com
SGLang

SGLang

Rare Pickgithub.com
MLX

MLX

Rare Pickgithub.com
vLLM

vLLM

Rare Pickgithub.com
llama.cpp

llama.cpp