llama_cpp-rs  by edgenai

Rust bindings for efficient CPU-based LLM inference

Created 3 years ago
251 stars

Top 99.9% on SourcePulse

GitHubView on GitHub
Project Summary

High-level, optionally asynchronous Rust bindings to the llama.cpp C++ library, llama_cpp-rs enables developers to integrate GGUF-based large language models directly into their Rust applications. It targets users seeking a simplified, CPU-centric approach to LLM inference, requiring no prior ML experience and allowing model execution in minimal code. The project offers a user-friendly API for loading models, managing inference sessions, and generating text completions.

How It Works

The library provides both idiomatic, high-level Rust bindings (crates/llama_cpp) and automatically generated low-level C API bindings (crates/llama_cpp_sys) to the core llama.cpp engine. It leverages llama.cpp's optimized inference capabilities for GGUF model formats, facilitating direct execution on the CPU. The design emphasizes ease of use, abstracting complex C++ interactions into a clean Rust interface for managing model weights, inference contexts, and token generation.

Quick Start & Requirements

  • Install/Run: Use cargo build --release or cargo run --release. Standard debug builds are not recommended due to significant performance degradation.
  • Prerequisites:
    • Rust toolchain.
    • GGUF format large language models.
    • Optional backend features require specific SDKs: CUDA Toolkit (for cuda), Vulkan SDK (for vulkan), ROCm (for hipblas). Metal backend is macOS-specific.
  • Links: No specific quick-start, documentation, or demo links are provided in the README.

Highlighted Details

  • Supports multiple hardware acceleration backends via Cargo features: CUDA, Vulkan, Metal (macOS), and HIPBLAS/ROCm.
  • Features an experimental capability for predicting context size in memory, though accuracy is not guaranteed.
  • Designed for simplicity, enabling LLM inference with approximately fifteen lines of Rust code.

Maintenance & Community

The README does not specify maintainers, sponsorships, or community channels (e.g., Discord, Slack). Contributions are explicitly welcomed, with an emphasis on maintaining a clean user experience.

Licensing & Compatibility

  • License: MIT or Apache-2.0, chosen by the user.
  • Compatibility: Both MIT and Apache-2.0 are permissive licenses, generally compatible with commercial use and integration into closed-source projects.

Limitations & Caveats

The context size prediction feature is explicitly noted as highly experimental and potentially inaccurate. Performance is heavily dependent on release builds; standard debug builds are impractical for inference due to computational intensity.

Health Check
Last Commit

2 years ago

Responsiveness

Inactive

Pull Requests (30d)
0
Issues (30d)
0
Star History
0 stars in the last 30 days

Explore Similar Projects

Starred by Andrej Karpathy Andrej Karpathy(Founder of Eureka Labs; Formerly at Tesla, OpenAI; Author of CS 231n), Anil Dash Anil Dash(Former CEO of Glitch), and
23 more.

llamafile by mozilla-ai

0.0%
26k
Single-file LLM distribution and runtime via `llama.cpp` and Cosmopolitan Libc
Created 3 years ago
Updated 1 day ago
Feedback? Help us improve.