Skip to content
#

continuous-batching

Here are 130 public repositories matching this topic...

Rapid-MLX

Rapid-MLX is an open-source (Apache 2.0) OpenAI- and Anthropic-compatible LLM inference server for Apple Silicon, built on MLX, focused on reliable tool calling for coding agents. Up to 4× faster than Apple's MLX (mlx-lm) on the same weights.

  • Updated Oct 10, 2026
  • Python

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

  • Updated Oct 9, 2026
  • Python

A high-performance API server that provides OpenAI-compatible endpoints for MLX models. Developed using Python and powered by the FastAPI framework, it provides an efficient, scalable, and user-friendly solution for running MLX-based vision and language models locally with an OpenAI-compatible interface.

  • Updated Oct 5, 2026
  • Python
dynalm

Run LLMs locally on your CPU: an open-source Ollama and llama.cpp alternative with no GPU needed. Llama, Qwen, Gemma, Phi, Mistral (GGUF, SafeTensors, GPTQ/AWQ), OpenAI-compatible API server. Linux, macOS, Windows.

  • Updated Oct 9, 2026
  • C++

Add this topic to your repo

To associate your repository with the continuous-batching topic, visit your repo's landing page and select "manage topics."

Learn more