Discover and explore top open-source AI tools and projects—updated daily.
edgenaiRust bindings for efficient CPU-based LLM inference
Top 99.9% on SourcePulse
High-level, optionally asynchronous Rust bindings to the llama.cpp C++ library, llama_cpp-rs enables developers to integrate GGUF-based large language models directly into their Rust applications. It targets users seeking a simplified, CPU-centric approach to LLM inference, requiring no prior ML experience and allowing model execution in minimal code. The project offers a user-friendly API for loading models, managing inference sessions, and generating text completions.
How It Works
The library provides both idiomatic, high-level Rust bindings (crates/llama_cpp) and automatically generated low-level C API bindings (crates/llama_cpp_sys) to the core llama.cpp engine. It leverages llama.cpp's optimized inference capabilities for GGUF model formats, facilitating direct execution on the CPU. The design emphasizes ease of use, abstracting complex C++ interactions into a clean Rust interface for managing model weights, inference contexts, and token generation.
Quick Start & Requirements
cargo build --release or cargo run --release. Standard debug builds are not recommended due to significant performance degradation.cuda), Vulkan SDK (for vulkan), ROCm (for hipblas). Metal backend is macOS-specific.Highlighted Details
Maintenance & Community
The README does not specify maintainers, sponsorships, or community channels (e.g., Discord, Slack). Contributions are explicitly welcomed, with an emphasis on maintaining a clean user experience.
Licensing & Compatibility
Limitations & Caveats
The context size prediction feature is explicitly noted as highly experimental and potentially inaccurate. Performance is heavily dependent on release builds; standard debug builds are impractical for inference due to computational intensity.
2 years ago
Inactive
Noeda
mozilla-ai