cm-memory

← Back to advanced search | Prev | Next

Untitled

ID
ff6d979e-de48-4b0a-a605-cd97d98ae20c
Created at
2026-04-15T22:34:58.185520+00:00
Updated at
2026-04-15T22:34:58.185520+00:00
Created by
mart
Source
memory
Reference
Skip to content Navigation Menu Sign in • Platform • AI CODE CREATIO…
People
Location
longitude -71.89675032526051
location 728 Rue McManamy Sherbrooke QC J1H 2M8 Canada
latitude 45.38947070326357
Memory Status
Notifications
Linked Memories
Text
app Reddit
input Skip to content Navigation Menu Sign in • Platform • AI CODE CREATION • 
GitHub Copilot
Write better code with AI • 
GitHub Spark
Build and deploy intelligent apps • 
GitHub Models
Manage and compare prompts • 
MCP RegistryNew
Integrate external tools • DEVELOPER WORKFLOWS • 
Actions
Automate any workflow • 
Codespaces
Instant dev environments • 
Issues
Plan and track work • 
Code Review
Manage code changes • APPLICATION SECURITY • 
GitHub Advanced Security
Find and fix vulnerabilities • 
Code security
Secure your code as you build • 
Secret protection
Stop leaks before they start • EXPLORE • Why GitHub • Documentation
 • Blog
 • Changelog
 • Marketplace • View all features



 • Solutions • BY COMPANY SIZE • Enterprises • Small and medium teams • Startups • Nonprofits • BY USE CASE • App Modernization • DevSecOps • DevOps • CI/CD • View all use cases
 • BY INDUSTRY • Healthcare • Financial services • Manufacturing • Government • View all industries
 • View all solutions



 • Resources • EXPLORE BY TOPIC • AI • Software Development • DevOps • Security • View all topics
 • EXPLORE BY TYPE • Customer stories • Events & webinars • Ebooks & reports • Business insights • GitHub Skills
 • SUPPORT & SERVICES • Documentation
 • Customer support
 • Community forum • Trust center • Partners • View all resources



 • Open Source • COMMUNITY • 
GitHub Sponsors
Fund open source developers • PROGRAMS • Security Lab
 • Maintainer Community
 • Accelerator • GitHub Stars
 • Archive Program
 • REPOSITORIES • Topics • Trending • Collections • 
 • Enterprise • ENTERPRISE SOLUTIONS • 
Enterprise platform
AI-powered developer platform • AVAILABLE ADD-ONS • 
GitHub Advanced Security
Enterprise-grade security features • 
Copilot for Business
Enterprise-grade AI features • 
Premium Support
Enterprise-grade 24/7 support • 
 • Pricing Search or jump to... Sign up Resetting focus mostlygeek / llama-swap Public • 
Code

Issues
26

Pull requests
13

Discussions

Actions

Wiki

Security and quality

Insights mostlygeek/llama-swap  main Code Folders and files Name Last commit date Latest commit mostlygeek Apr 15, 2026 History .github Apr 15, 2026 ai-plans Dec 19, 2025 cmd Mar 19, 2026 docker Apr 11, 2026 docs Apr 15, 2026 event Jul 15, 2025 models Oct 3, 2024 proxy Apr 15, 2026 scripts May 23, 2025 ui-svelte Apr 12, 2026 View all files Repository files navigation • 
README

MIT license llama-swap Run multiple generative AI models on your machine and hot-swap between them on demand. llama-swap works with any OpenAI and Anthropic API compatible server and is used by thousands of people to power their local AI workflows. Built in Go for performance and simplicity, llama-swap has zero dependencies and is incredibly easy to set up. Get started in minutes - just one binary and one configuration file. Features: • ✅ Easy to deploy and configure: one binary, one configuration file. no external dependencies • ✅ On-demand model switching • ✅ Use any local OpenAI compatible server (llama.cpp, vllm, tabbyAPI, stable-diffusion.cpp, etc.) ◦ future proof, upgrade your inference servers at any time. • ✅ OpenAI API supported endpoints: ◦ v1/completions ◦ v1/chat/completions ◦ v1/responses ◦ v1/embeddings ◦ v1/audio/speech (#36) ◦ v1/audio/transcriptions (docs) ◦ v1/audio/voices ◦ v1/images/generations ◦ v1/images/edits • ✅ Anthropic API supported endpoints: ◦ v1/messages ◦ v1/messages/count_tokens • ✅ llama-server (llama.cpp) supported endpoints ◦ v1/rerank, v1/reranking, /rerank ◦ /infill - for code infilling ◦ /completion - for completion endpoint • ✅ SDAPI via stable-diffusion.cpp's server ◦ /sdapi/v1/txt2img ◦ /sdapi/v1/img2img ◦ /sdapi/v1/loras - requires model in request body to fetch the correct loras • ✅ llama-swap API ◦ /ui - web UI ◦ /upstream/:model_id - direct access to upstream server (demo) ◦ /models/unload - manually unload running models (#58) ◦ /running - list currently running models (#61) ◦ /log - remote log monitoring ◦ /health - just returns "OK" • ✅ API Key support - define keys to restrict access to API endpoints • ✅ Customizable ◦ Run concurrent models with a custom DSL swap matrix (#643) ◦ Automatic unloading of models after timeout by setting a ttl ◦ Reliable Docker and Podman support using cmd and cmdStop together ◦ Preload models on startup with hooks (#235) Web UI llama-swap includes a real time web interface with a playground for testing out all sorts of local models: View detailed token metrics: Inspect request and responses: Manually load and unload models: Real time log streaming: Installation llama-swap can be installed in multiple ways 1 Docker 2 Homebrew (OSX and Linux) 3 WinGet 4 From release binaries 5 From source Docker Install (download images) Nightly container images with llama-swap and llama-server are built for multiple platforms (cuda, vulkan, intel, etc.) including non-root variants with improved security. The stable-diffusion.cpp server is also included for the musa and vulkan platforms. $ docker pull ghcr.io/mostlygeek/llama-swap:cuda # run with a custom configuration and models directory $ docker run -it --rm --runtime nvidia -p 9292:8080 \ -v /path/to/models:/models \ -v /path/to/custom/config.yaml:/app/config.yaml \ ghcr.io/mostlygeek/llama-swap:cuda # configuration hot reload supported with a # directory volume mount $ docker run -it --rm --runtime nvidia -p 9292:8080 \ -v /path/to/models:/models \ -v /path/to/custom/config.yaml:/app/config.yaml \ -v /path/to/config:/config \ ghcr.io/mostlygeek/llama-swap:cuda -config /config/config.yaml -watch-config more examples # pull latest images per platform docker pull ghcr.io/mostlygeek/llama-swap:cpu docker pull ghcr.io/mostlygeek/llama-swap:cuda docker pull ghcr.io/mostlygeek/llama-swap:vulkan docker pull ghcr.io/mostlygeek/llama-swap:intel docker pull ghcr.io/mostlygeek/llama-swap:musa # tagged llama-swap, platform and llama-server version images docker pull ghcr.io/mostlygeek/llama-swap:v166-cuda-b6795 # non-root cuda docker pull ghcr.io/mostlygeek/llama-swap:cuda-non-root Homebrew Install (macOS/Linux) brew tap mostlygeek/llama-swap brew install llama-swap llama-swap --config path/to/config.yaml --listen localhost:8080 WinGet Install (Windows) Note WinGet is maintained by community contributor Dvd-Znf (#327). It is not an official part of llama-swap. # install C:\> winget install llama-swap # upgrade C:\> winget upgrade llama-swap Pre-built Binaries Binaries are available on the release page for Linux, Mac, Windows and FreeBSD. Building from source 1 Building requires Go and Node.js (for UI). 2 git clone https://github.com/mostlygeek/llama-swap.git 3 make clean all 4 look in the build/ subdirectory for the llama-swap binary Configuration # minimum viable config.yaml models: model1: cmd: llama-server --port ${PORT} --model /path/to/model.gguf That's all you need to get started: 1 models - holds all model configurations 2 model1 - the ID used in API calls 3 cmd - the command to run to start the server. 4 ${PORT} - an automatically assigned port number Almost all configuration settings are optional and can be added one step at a time: • Advanced features ◦ matrix to run concurrent models with a custom swap logic DSL ◦ hooks to run things on startup ◦ macros reusable snippets • Model customization ◦ ttl to automatically unload models ◦ aliases to use familiar model names (e.g., "gpt-4o-mini") ◦ env to pass custom environment variables to inference servers ◦ cmdStop gracefully stop Docker/Podman containers ◦ useModelName to override model names sent to upstream servers ◦ ${PORT} automatic port variables for dynamic port assignment ◦ filters rewrite parts of requests before sending to the upstream server See the configuration documentation for all options. How does llama-swap work? When a request is made to an OpenAI compatible endpoint, llama-swap will extract the model value and load the appropriate server configuration to serve it. If the wrong upstream server is running, it will be replaced with the correct one. This is where the "swap" part comes in. The upstream server is automatically swapped to handle the request correctly. In the most basic configuration llama-swap handles one model at a time. For more advanced use cases, using a matrix allows multiple models to be loaded at the same time. You have complete control over how your system resources are used. Reverse Proxy Configuration (nginx) If you deploy llama-swap behind nginx, disable response buffering for streaming endpoints. By default, nginx buffers responses which breaks Server‑Sent Events (SSE) and streaming chat completion. (#236) Recommended nginx configuration snippets: # SSE for UI events/logs location /api/events { proxy_pass http://your-llama-swap-backend; proxy_buffering off; proxy_cache off; } # Streaming chat completions (stream=true) location /v1/chat/completions { proxy_pass http://your-llama-swap-backend; proxy_buffering off; proxy_cache off; } As a safeguard, llama-swap also sets X-Accel-Buffering: no on SSE responses. However, explicitly disabling proxy_buffering at your reverse proxy is still recommended for reliable streaming behavior. Monitoring Logs on the CLI # sends up to the last 10KB of logs $ curl http://host/logs # streams combined logs curl -Ns http://host/logs/stream # stream llama-swap's proxy status logs curl -Ns http://host/logs/stream/proxy # stream logs from upstream processes that llama-swap loads curl -Ns http://host/logs/stream/upstream # stream logs only from a specific model curl -Ns http://host/logs/stream/{model_id} # stream and filter logs with linux pipes curl -Ns http://host/logs/stream | grep 'eval time' # appending ?no-history will disable sending buffered history first curl -Ns 'http://host/logs/stream?no-history' Do I need to use llama.cpp's server (llama-server)? Any OpenAI compatible server would work. llama-swap was originally designed for llama-server and it is the best supported. For Python based inference servers like vllm or tabbyAPI it is recommended to run them via podman or docker. This provides clean environment isolation as well as responding correctly to SIGTERM signals for proper shutdown. Star History Note ⭐️ Star this project to help others discover it! Releases 144 v202 Latest Apr 15, 2026 + 143 releases Packages 1 • 
llama-swap Contributors 41 • • • • • • • • • • • • • • + 27 contributors Languages • 
Go
62.9%

Svelte
22.3%

Shell
6.2%

TypeScript
5.9%

Dockerfile
1.3%

CSS
0.6%

Other
0.8% Footer © 2026 GitHub, Inc. Footer navigation • Terms • Privacy • Security • Status • Community • Docs • Contact • Manage cookies • Do not share my personal information
Audio
AI Transcription
AI Image Description
AI Title
AI Summary
Tags
AI Status
Pending
✎ Edit

Images

Image 1