company-ceo
Run a PenguinHarness organization as its CEO — turn the mission into a ticket tree, hire HR and finance first, partition the shared workspace, schedule the…
Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.
$ npx -y skills add Prism-Shadow/penguin-harness --skill ollama --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/ollamaContext preview
The summary Claude sees to decide when to auto-load this skill.
Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.
name: ollama description: Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.
Ollama runs open-weight models locally with automatic GPU detection and an OpenAI-compatible API on `http://localhost:11434`.
If the user's message only invokes this skill (e.g. "use ollama skill") without a concrete request, ask the user what they want. Do not run any command until the goal is clear.
Ask the user which model to run; if they have no preference, recommend the small default [Qwen/Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) (`ollama pull qwen3.5:0.8b`). The model must fit the machine's RAM/VRAM.
Ollama runs everywhere — macOS, Linux and Windows, on CPUs as well as NVIDIA/AMD GPUs — so engine choice follows the user's preference: Ollama is the simple default, while vLLM targets high-throughput GPU serving. Check the current state first:
ollama --version # is Ollama installed? ollama ps # is the service already serving models?
If port 11434 is already serving, reuse that instance — never kill an existing Ollama process.
1. Ask the user which model to run; with no preference, recommend [Qwen/Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) (`qwen3.5:0.8b`). 2. Pick the engine the user prefers: Ollama is the default; vLLM covers high-throughput GPU serving. 3. Install Ollama if missing, then `ollama pull qwen3.5:0.8b`. 4. Verify with `curl http://localhost:11434/v1/models`. 5. Register the endpoint: `penguin config model add ... --client-type openai-chat --base-url http://localhost:11434/v1` — a pulled Ollama model is not visible to Penguin until added. 6. Confirm the new entry with `penguin config model list`.
curl -fsSL https://ollama.com/install.sh | sh # Linux; macOS/Windows use the desktop app
The service then listens on `http://localhost:11434`.
ollama pull qwen3.5:0.8b # download a model ollama run qwen3.5:0.8b # interactive chat (pulls first if missing) ollama list # downloaded models ollama ps # models loaded in memory ollama stop qwen3.5:0.8b # unload a model
The endpoint is `http://localhost:11434/v1`; any non-empty API key is accepted (conventionally `ollama`):
curl http://localhost:11434/v1/models
The default context window is small, and agent sessions need a large one. Raise it in the server's environment:
OLLAMA_CONTEXT_LENGTH=32768 ollama serve # systemd service: set it via `systemctl edit ollama`
Or bake it into a model variant with a Modelfile:
FROM qwen3.5:0.8b PARAMETER num_ctx 32768
ollama create qwen3.5-32k -f Modelfile
Model configuration is the penguin CLI's job — `penguin config model add` registers an endpoint and `penguin config model list` shows what has been registered. A pulled Ollama model is not visible to Penguin until you add it:
penguin config model add --provider custom --client-type openai-chat \ --base-url http://localhost:11434/v1 --model-id qwen3.5:0.8b --api-key ollama penguin config model list # the new entry should now be listed
🐧 Harness for RSI. Let AI Build AI. Multi-Agent Auto-Dev Platform. Everything is Transparent.
Repo: Prism-Shadow/penguin-harness
Run a PenguinHarness organization as its CEO — turn the mission into a ticket tree, hire HR and finance first, partition the shared workspace, schedule the…
The protocol every employee of a PenguinHarness organization follows — read the handbook first, act on [org_trigger] work runs, schedule tickets from the desk…
Run finance for a PenguinHarness organization — set and adjust monthly budgets along the reporting line, audit spend daily with penguin org finance and penguin…
Run HR for a PenguinHarness organization — guarantee every employee has an enabled calendar event, hire and offboard employees (channels included), evaluate…
Run a PenguinHarness organization that mirrors a real company — one digital twin per real colleague, each twin's desk bound to that colleague's chat bot, twins…
Create a company-mode organization together with the user — collect its id, name, mission, shared workspace, model and CEO budget one question at a time,…