github.com · Inference Engines
Rare Pick

llama.cpp

Visit site
llama.cpp

About

llama.cpp is a C/C++ port of Meta's LLaMA (Large Language Model Meta AI) that enables efficient inference of large language models on consumer hardware. It is designed for simplicity and performance, allowing users to run LLaMA and other compatible models locally with minimal setup.

This tool is primarily for developers, researchers, and enthusiasts who want to experiment with or deploy large language models on their own machines. It supports various quantization methods and hardware acceleration, making it suitable for a range of applications from local AI assistants to research projects.

More from 0xSero

See all
Rare Picknvidia.com
RTX PRO 6000 Blackwell

RTX PRO 6000 Blackwell

Rare Picknvidia.com
RTX 3090

RTX 3090

Rare Pickframe.work
Framework Desktop

Framework Desktop

Rare Pickgithub.com
vLLM

vLLM

Rare Pickgithub.com
SGLang

SGLang

Rare Pickgithub.com
MLX

MLX

More in Inference Engines

See all
Rare Pickgithub.com
SGLang

SGLang

Rare Pickgithub.com
ExLlamaV3

ExLlamaV3

Rare Pickgithub.com
MLX

MLX

Rare Pickgithub.com
vLLM

vLLM