github.com · Inference Engines
Rare Pick

llama.cpp

Visit site
llama.cpp

About

llama.cpp is a C/C++ port of Meta's LLaMA (Large Language Model Meta AI) that enables efficient inference of large language models on consumer hardware. It is designed for simplicity and performance, allowing users to run LLaMA and other compatible models locally with minimal setup.

This tool is primarily for developers, researchers, and enthusiasts who want to experiment with or deploy large language models on their own machines. It supports various quantization methods and hardware acceleration, making it suitable for a range of applications from local AI assistants to research projects.

Used by 1 person

More in Inference Engines

See all
Rare Pickgithub.com
SGLang

SGLang

Rare Pickgithub.com
ExLlamaV3

ExLlamaV3

Rare Pickgithub.com
MLX

MLX

Rare Pickgithub.com
vLLM

vLLM