vllm-project/vllm
GitHubA high-throughput and easy-to-use library for LLM inference and serving, designed for diverse hardware architectures and model structures.
Teams requiring high-throughput, low-latency LLM serving infrastructure for production environments.
Review deployment and security evidence before production adoption.
What matters before you adopt it
Security evidence without the noise
Trust remains a decision signal; CVEs and scanner evidence explain what is driving the risk.
View trust evidence & security findings▼
Architecture from code13 modules · 7 edges▼
Modules and dependency edges extracted from repository code. This is code evidence, not README inference.
Architecture evidence details
flowchart TD
%% vllm — high-level architecture (DRAFT, refine me)
n0["(root) · 2 files"]
n1["benchmarks · 104 files"]
n2["cmake · 1 file"]
n3["csrc · 5 files"]
n4["docs · documentation · 13 files"]
n5["examples · examples · 150 files"]
n6["rust · 1 file"]
n7["scripts · scripts · 2 files"]
n8["tests · tests · 1756 files"]
n9["tools · tool implementations · 26 files"]
n10["vllm · 394 files"]
n1 --> n8
n1 --> n10
n5 --> n10
n7 --> n10
n8 --> n9
n8 --> n10
n9 --> n10
class n4,n5 docs
class n7 infra
class n9 shared
class n8 test
classDef docs fill:#9d7660,color:#ffffff,stroke:#7c5d4c
classDef infra fill:#b35c00,color:#ffffff,stroke:#8f4a00
classDef shared fill:#79706e,color:#ffffff,stroke:#5d5654
classDef test fill:#499894,color:#ffffff,stroke:#397975Evidence, security & integrations▼
Nearby repositories worth comparing before adoption.
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.