adhd-output-style
This skill should be used when the user asks for "ADHD output", "fewer output tokens", "short…
Deploys and operates a LiveKit agent in production: shipping a version to LiveKit Cloud and rolling it back, secrets and configuration, the worker process model and prewarming, safe async inside worker processes, provider timeouts and degradation, graceful shutdown, SDK
$ npx -y skills add fcakyon/claude-codex-settings --skill operating-livekit-agents --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/operating-livekit-agentsContext preview
The summary Claude sees to decide when to auto-load this skill.
Deploys and operates a LiveKit agent in production: shipping a version to LiveKit Cloud and rolling it back, secrets and configuration, the worker process model and prewarming, safe async inside worker processes, provider timeouts and degradation, graceful shutdown, SDK
name: operating-livekit-agents description: 'Deploys and operates a LiveKit agent in production: shipping a version to LiveKit Cloud and rolling it back, secrets and configuration, the worker process model and prewarming, safe async inside worker processes, provider timeouts and degradation, graceful shutdown, SDK upgrades, and observability. Use when the user says "deploy my agent", "roll back the deployment", "tail the agent logs", "the first call after a restart is slow", "attached to a different loop / event loop is closed", "prewarm the VAD", "shut down without dropping calls", "upgrade livekit-agents safely", "correlate logs by session", or is changing an agent codebase that is already live. Not for designing the agent (building-livekit-agents) or reproducing one bad conversation locally (debugging-livekit-agents).' license: MIT metadata: author: livekit
Everything after the agent works: getting a version onto LiveKit Cloud, keeping it fast and alive under real load, and changing it without breaking what's already running. The commands live under `lk agent`; read `lk agent --help` and each subcommand's help rather than trusting this skill for flags — it deliberately doesn't restate them. `reading-livekit-docs` has the deployment and observability docs.
On a self-hosted LiveKit server the `lk agent` deploy commands do not apply. Ship the agent as a container or process under your own tooling (Docker, Kubernetes, systemd), keep secrets in its environment, and look up the server side with `lk docs get-page /home/self-hosting`. The worker, shutdown, and observability guidance below applies either way.
The shape is stable even as the flags move:
1. **A project directory is bound to an agent** by a config file the CLI writes (`livekit.toml`). Commands run from that directory find the agent without an id. 2. **The agent ships as a container.** The CLI can generate a Dockerfile for the project, or you bring a prebuilt image. Run the container's start command locally (`lk agent start`) before the first deploy — it's production mode, with production logging and a shutdown drain, and it is not what `dev` mode runs. 3. **Each deploy creates a version** and rolls it out. `status`, `versions`, and `logs` (build logs and deploy logs are separate) tell you what's live and why a rollout failed. 4. **Secrets are injected as environment variables** and managed apart from the code — never baked into the image or committed. Changing secrets restarts the agent. 5. **Rollback returns to a previous version.** How instant that is depends on the plan; the docs say. Know the rollback command before you need it.
After a deploy, verify with the same tools you'd use on a stranger's agent: `status` for the rollout, `logs` for the first minutes, and a real conversation — `running-livekit-simulations` can run a scenario file against the deployed agent by name, which is the cheapest end-to-end check that the thing serving traffic is the thing you meant to ship.
Both SDKs run sessions in worker processes spawned from a parent. Misunderstanding this is the single most common source of production-only bugs.
**The parent prewarms; children inherit.** Load expensive, read-only resources — VAD and turn detection models, persistent clients — once in the parent through the SDK's prewarm hook, and each session inherits them without re-loading. Everything shared this way must be read-only or concurrency-safe; mutating parent state from a child is undefined behavior. Don't prewarm session-specific state, and don't prewarm what costs more memory than it saves — every byte in the parent is in every child's footprint. Then **verify the child actually uses the prewarmed instance**: the classic mistake is prewarming a model and having session code load a fresh one anyway, so the prewarm did nothing and startup is still slow.
**The framework owns the event loop.** Never create a new async runtime inside a worker, and never block on an async call from a synchronous constructor to force a result. If initialization needs async work, load lazily on first use from an already-async method, or split construction from an awaited `initialize` step. When you see errors about events bound to a different loop, a loop already running, tasks destroyed while pending, or a closed loop that can't be reused, the cause is almost always one of those two things — trace back to where a runtime was created or a sync path awaited something.
STT, TTS, LLM, VAD and any backend will fail in production: rate limits, timeouts, overload, outages. Set timeouts — a call that hangs is worse than one that fails fast. Distinguish transient failures worth retrying from persistent ones that need a fallback or escalation. Degrade to a meaningful spoken response, never to silence. Log provider response times, because rising latency is usually the first sign of an outage.
Adding a provider to an existing agent: check whether the SDK already has a plugin before writing one; follow how the codebase already initializes, configures, and handles errors for its other providers; run its config through the same pipeline; and test under realistic latency, not just happy-path responses.
The metric users feel is **time to first audio** — from the caller finishing to the agent starting to speak. Break it down before optimizing: connection, provider initialization, first inference, tool execution. Watch **context growth** over a long call; unbounded history means every turn is slower than the last.
The anti-patterns worth grepping for:
every concurrent operation in that process.
latency the call
Battle-tested Claude Code, OpenAI Codex, Cursor configs, plugins, hooks and agents with Kimi, MiniMax and GLM API support.
Repo: fcakyon/claude-codex-settings
This skill should be used when the user asks for "ADHD output", "fewer output tokens", "short…
Agent-browser usage guide. Read this before running any agent-browser commands. Covers the…
Reverse-engineer a website's internal API by recording browser traffic into a HAR file, then…
Systematically explore and test a web application to find bugs, UX issues, and other…
Automate Electron desktop apps (VS Code, Slack, Discord, Figma, Notion, Spotify, etc.) using…
Build and validate experimental WebMCP tools for an existing web page. Use when an agent…