Skip to content

Top 100 Inference Engine Repositories

Back to Home

Ranking

RankingProject NameStarsForksLanguageOpen IssuesDescriptionLast Commit
1vllm90,92721,676Python2325A high-throughput and memory-efficient inference and serving engine for LLMs2026-09-04
2ds422,0542,065C240DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm2026-09-03
3web-llm18,9561,373TypeScript134High-performance In-browser LLM Inference Engine2026-09-03
4ml-engineering18,8901,236Python2Machine Learning Engineering Open Book2026-09-04
5MNN16,0172,427C++27MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.2026-09-03
6Paddle-Lite7,2731,619C++46PaddlePaddle High Performance Deep Learning Inference Engine for Mobile and Edge (飞桨高性能深度学习端侧推理引擎)2026-04-27
7gemma.cpp7,035659C++26lightweight, standalone C++ inference engine for Google's Gemma models.2026-09-03
8cactus5,979500C++41Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.2026-08-26
9shimmy5,825561Rust9⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.2026-08-30
10DALI5,751677C++191A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.2026-09-03
11CTranslate24,661526C++228Fast inference engine for Transformer models2026-08-31
12Tengine4,532982C++244Tengine is a lite, high performance, modular inference engine for embedded device2025-03-06
13Rapid-MLX3,649413Python45The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replace...2026-09-04
14TransformerEngine3,519818Python160A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance wi...2026-09-03
15spiceai3,075226Rust717Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.2026-09-04
16xDiT2,707342Python52xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism2026-09-03
17h3.c2,555193C26MiniMax H3 inference engine for Mac computers2026-08-11
18openlake2,405422Rust111OpenLake is a high performance storage engine for efficient LLM inference and GPU Training2026-09-03
19AI-Engineering.academy2,378277Jupyter Notebook6Mastering Applied AI, One Concept at a Time2026-02-27
20warp2,351174C7Run the full 2.78-trillion-parameter Kimi K3 model or GLM-5.3-Flash beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.2026-08-28
21gpu-perf-engineering-resources2,288252Python0A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.2026-08-23
22audio.cpp2,260272C++10An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependenc...2026-09-04
23tokenspeed2,088274Python13TokenSpeed is a speed-of-light LLM inference engine.2026-09-04
24ai-performance-engineering1,907263Python3Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning.2026-08-31
25sonar1,847209C++80Large-scale LLM inference engine2026-08-13
26Genie-TTS1,761120Python31GPT-SoVITS ONNX Inference Engine & Model Converter2026-08-30
27uzu1,70974Rust2A high-performance inference engine for AI models2026-09-03
28xllm1,555290C++83A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.2026-09-04
29Atomic-Chat1,424161TypeScript33Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V2026-09-03
30rtp-llm1,326271Cuda38RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.2026-09-04
31airunner1,314103Python0Offline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflows2026-08-29
32Jlama1,303164Java39Jlama is a modern LLM inference engine for Java2025-10-12
33cache-dit1,27081Python85A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.2026-09-02
34openrouter-runner1,258123Python0Deprecated inference engine2025-09-06
35FeatherCNN1,227275C++18FeatherCNN is a high performance inference engine for convolutional neural networks.2019-09-24
36ezkl1,221212Rust15ezkl is an engine for doing inference for deep learning models and other computational graphs in a zk-snark (ZKML). Use it from Python, Javascript, or the command line.2026-02-20
37tiny-vllm1,09385C++0Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM2026-08-23
38nobodywho1,09177Rust11NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.2026-09-03
39YOLOs-CPP1,081162C++0Cross-Platform Production-ready C++ inference engine for YOLO models (v5-v12, YOLO26). Unified API for detection, segmentation, pose estimation, OBB, and classification. Built on ONNX Runtime and Open...2026-08-23
40checkpoint-engine1,005107Python3Checkpoint-engine is a simple middleware to update model weights in LLM inference engines2026-08-12
41ssd99578Python2A lightweight inference engine supporting speculative speculative decoding (SSD).2026-05-10
42TinyChatEngine961102C++35TinyChatEngine: On-Device LLM Inference Library2024-07-04
43ZhiLight908104C++5A highly optimized LLM inference acceleration engine for Llama and its variants.2026-03-18
44kronk78057Go7Your personal engine for running open source models locally. Use Go for hardware accelerated local inference with llama.cpp, whisper.cpp, and stablediffusion.cpp directly integrated into your Go appli...2026-09-03
45emlearn75079Python16Machine Learning inference engine for Microcontrollers and Embedded devices2026-07-17
46atlas676102Rust83Pure Rust Inference Engine2026-09-04
47pegainfer670103Rust53Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K22026-09-03
48libonnx652114C16A lightweight, portable pure C99 onnx inference engine for embedded devices with hardware acceleration support.2026-07-07
49hipfire60265Rust90RDNA-native LLM inference engine in Rust.2026-09-04
50tidy59845Kotlin34Offline semantic Text-to-Image and Image-to-Image search on Android powered by quantized state-of-the-art vision-language pretrained CLIP model and ONNX Runtime inference engine2024-03-28
51swama59232Swift38High-performance MLX-based LLM inference engine for macOS with native Swift implementation2026-09-04
52WhisperS2T57875Jupyter Notebook31An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine2024-08-27
53qwen60055947Cuda0Static suckless single batch CUDA-only qwen3-0.6B mini inference engine2025-09-08
54FlashRT54473C++13FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Al...2026-08-31
55Anakin537135C++53High performance Cross-platform Inference-engine, you could run Anakin on x86-cpu,arm, nv-gpu, amd-gpu,bitmain and cambricon devices.2022-09-23
56VectorHub530135Jupyter Notebook1Deprecated historical repo. Superlinked now develops SIE, a self-hosted inference engine for embeddings, reranking, OCR, extraction, and document processing.2026-08-31
57dotLLM51460C#193LLM inference engine written in .NET2026-07-30
58OpenArc51344Python11Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.2026-09-03
59zinc51119Zig4Zig INferenCe Engine — Local LLM inference on AMD GPUs and Apple Silicon2026-09-03
60simple-llm48237Python0~950 line, minimal, extensible LLM inference engine built from scratch.2026-01-09
61crabml47045Rust24a fast cross platform AI inference engine 🤖 using Rust 🦀 and WebGPU 🎮2025-01-04
62ntransformer46519C++2High-efficiency LLM inference engine in C++/CUDA. Run Llama 70B on RTX 3090.2026-02-22
63Crane46253Rust26A Pure Rust based LLM, VLM, VLA, TTS, OCR Inference Engine, powering by Candle & Rust. Alternate to your llama.cpp but much more simpler and cleaner..2026-08-31
64flash-tokenizer45811C++7EFFICIENT AND OPTIMIZED TOKENIZER ENGINE FOR LLM INFERENCE SERVING2026-02-02
65JetStream45767Python14JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).2026-01-05
66InfiniTensor44774C++24InfiniTensor is a high-performance inference engine tailored for GPUs and AI accelerators. Its design focuses on effective deployment and swift academic validation.2026-09-01
67gpu-rest-engine42295C++6A REST API for Caffe using Docker and Go2018-07-20
68TensorSharp40639C#3A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It suppor...2026-09-04
69AutoGrad-Engine39950C#0A complete GPT language model (training and inference) in ~600 lines of pure C#, zero dependencies2026-02-14
70StockInference-Spark382194Java5Stock inference engine using Spring XD, Apache Geode / GemFire and Spark ML Lib.2016-06-03
71flex-nano-vllm35821Python1FlexAttention based, minimal vllm-style inference engine for fast Gemma 2 inference.2025-11-02
72sentis-samples35873C#11Inference Engine samples internal development repository. Contains example and template projects for Sentis package use.2026-08-12
73audio.cpp-webui34873C++3audio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generation. No Python dependency.2026-08-14
74rten33326Rust42ONNX neural network inference engine2026-09-02
75AMDMIGraphX329148C++247AMD's graph optimization engine.2026-09-03
76RL-Kernel29080Python81High-performance RL post-training infrastructure. Designed to achieve bitwise operator-level train-inference consistency across heterogeneous engines and extreme memory efficiency for GRPO, PPO, etc.2026-09-03
77elfi28362Python10ELFI - Engine for Likelihood-Free Inference2025-05-07
78yolov4-triton-tensorrt28261C++3This repository deploys YOLOv4 as an optimized TensorRT engine to Triton Inference Server2022-06-02
79awesome-edge-machine-learning28156Python1A curated list of awesome edge machine learning resources, including research papers, inference engines, challenges, books, meetups and others.2023-02-23
80dash-infer27328C7DashInfer is a native LLM inference engine aiming to deliver industry-leading performance atop various hardware architectures, including CUDA, x86 and ARMv9.2025-08-06
81tflite2tensorflow27242Python1Generate saved_model, tfjs, tf-trt, EdgeTPU, CoreML, quantized tflite, ONNX, OpenVINO, Myriad Inference Engine blob and .pb from .tflite. Support for building environments with Docker. It is possible ...2022-09-04
82whisper.el26625Emacs Lisp8Speech-to-Text interface for Emacs using OpenAI's whisper model and whisper.cpp as inference engine.2026-07-17
83ai-hardware-engineer-roadmap26539HTML0Master AI inference, AI agent harness systems, and hardware engineering — then design a physical AI chip. That is the goal.2026-09-03
84oramacore26123Rust9OramaCore is the complete runtime you need for your projects, answer engines, copilots, and search. It includes a fully-fledged full-text search engine, vector database, LLM interface, and many more u...2026-04-14
85TurboLLM25937TypeScript5Run any local LLM engine, auto-tuned to your GPU — polished web UI + OpenAI/Anthropic-compatible API. Point Claude Code at your own machine in one command. No Electron, no Python, offline-first.2026-09-03
86compute-engine25735C++17Highly optimized inference engine for Binarized Neural Networks2026-07-30
87inferflow25125C++8Inferflow is an efficient and highly configurable inference engine for large language models (LLMs).2024-03-15
88lm-inference-engines2419-8Comparison of Language Model Inference Engines2024-12-16
89KokoroSharp24130C#11Fast local TTS inference engine in C# with ONNX runtime. Multi-speaker, multi-platform and multilingual. Integrate on your .NET projects using a plug-and-play NuGet package, complete with all voices.2026-08-14
90llm-inference-engineering24028Markdown0Learn LLM Inference Engineering step by step - from KV cache, PagedAttention, and continuous batching to vLLM, SGLang, and GPUs.2026-09-03
91Awesome-LLM-Inference-Engine23722-12026-09-03
92amd_inference2348Python11Docker-based inference engine for AMD GPUs2024-10-07
93zse23414Python1The inference engine the open-source world built for itself.2026-08-02
94MIVisionX21792C++12AMD MIVisionX is a computer vision toolkit built around a highly optimized, conformant open-source implementation of the Khronos OpenVX™ 1.3.2 specification. As of the 4.0.0 release, MIVisionX ships t...2026-09-03
95pulsar21028Rust6SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places att...2026-09-01
96mlsub20421OCaml11Prototype type inference engine2025-01-31
97embedded-ai.bench20229Python17benchmark for embededded-ai deep learning inference engines, such as NCNN / TNN / MNN / TensorFlow Lite etc.2021-02-18
98rf-detr-cpp20121C++0Production-ready C++/TensorRT inference engine for RF-DETR. Object detection and instance segmentation with FP32/FP16/INT8 support. Optimized for NVIDIA GPUs, Jetson (Orin, AGX Thor).2026-08-14
99llm-systems-engineering-roadmap19326-0A practical roadmap for mastering LLM internals, training, inference, RAG, agents, evaluation, and production architecture.2026-07-27
100microflow-rs19030Rust3A robust and efficient TinyML inference engine.2026-05-26