bench
mlx-serve benchmarking methodology — bench.sh/llmprobe usage, comparison-trap rules…
Hook an app, game or script up to the local mlx-serve server for LLM chat, embeddings, image, speech, music, video and 3D generation, and Laya/Kev typed decisions. Use when code should call mlx-serve.
$ npx -y skills add ddalcu/mlx-serve --skill mlx-serve --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/mlx-serveContext preview
The summary Claude sees to decide when to auto-load this skill.
Hook an app, game or script up to the local mlx-serve server for LLM chat, embeddings, image, speech, music, video and 3D generation, and Laya/Kev typed decisions. Use when code should call mlx-serve.
name: mlx-serve description: Hook an app, game or script up to the local mlx-serve server for LLM chat, embeddings, image, speech, music, video and 3D generation, and Laya/Kev typed decisions. Use when code should call mlx-serve.
mlx-serve runs MLX models on this Mac behind one HTTP port: OpenAI, Anthropic and Ollama compatible chat, plus native endpoints for images, speech, music, video, 3D and decisions. Everything below is plain HTTP + JSON, so any language works.
config value or env var in the code you write, never a hardcoded LAN IP.
`Authorization: Bearer <key>` from other machines; a 401 means that.
directly.
curl -s "${MLX_SERVE_URL:-http://127.0.0.1:11234}/v1/models" \
| jq -r '.data[] | "\(.id)\t\(.capabilities | join(","))\t\(.state)"'Each row has `id` (like `org/name`), `capabilities`, `state` (`ready`, `unloaded`, `remote`) and `context_length`. Choose by capability:
| capability | endpoint | details | |---|---|---| | `chat` | `POST /v1/chat/completions` (also `/v1/messages`, `/v1/responses`) | chat.md | | `embeddings` | `POST /v1/embeddings` | chat.md | | `image` | `POST /v1/images/generations`, `POST /v1/images/edits` | media.md | | `audio` without `music` | `POST /v1/audio/speech` (TTS) | media.md | | `music` | `POST /v1/audio/music-generations` | media.md | | `video` | `POST /v1/video/generations` | media.md | | `3d` | `POST /v1/3d/generations` | media.md | | `decisions` | `POST /v1/decisions` | decisions.md |
Read the linked file (next to this one) before writing client code for that endpoint. If no model has the capability the user needs, say so and tell them to download one in the MLX Core app (Models). Do not invent an id.
the first request, which can take seconds to minutes. `POST /v1/load-model {"model": "<id>"}` pre-warms one; `POST /v1/unload-model {"model": "<id>"}` frees it.
something first (typically the media model when you are done with it).
replies for 3 minutes. Decisions and embeddings are milliseconds.
input path. Cache results on disk keyed by model + prompt + seed + size, and ship the cache (or regenerate lazily) instead of calling on every run.
music 30 s to minutes, 3D 1-5 min, video minutes. Set client timeouts to match (10+ minutes for video), and keep the game playable while waiting.
progress bar. See media.md.
message names the bad field. Show it; never retry a 400 unchanged.
swap models without touching game code.
BASE="${MLX_SERVE_URL:-http://127.0.0.1:11234}"
IMG=$(curl -s "$BASE/v1/models" | jq -r '[.data[] | select(.capabilities | index("image"))][0].id')
curl -s "$BASE/v1/images/generations" -H 'content-type: application/json' \
-d "{\"model\":\"$IMG\",\"prompt\":\"pixel art treasure chest on a plain white background\",\"size\":\"512x512\",\"seed\":7}" \
| jq -r '.data[0].b64_json' | base64 -d > chest.pngOpenAI- and Anthropic-compatible local inference for Apple Silicon — MLX and GGUF — faster than LM Studio on identical MLX weights. No Python. No cloud. No Electron.
Repo: ddalcu/mlx-serve
mlx-serve benchmarking methodology — bench.sh/llmprobe usage, comparison-trap rules…
Turn llmprobe reports into the charts a perf PR embeds (one line per arm across context size,…
mlx-serve pre-release validation checklist, CalVer versioning, release steps, and CHANGELOG…