adhd-output-style
This skill should be used when the user asks for "ADHD output", "fewer output tokens", "short…
Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them. Use when the user asks "what should I test", "generate simulation scenarios", "write scenarios for my agent", "add a scenario for X", "organize my scenario files", "my
$ npx -y skills add fcakyon/claude-codex-settings --skill writing-livekit-scenarios --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/writing-livekit-scenariosContext preview
The summary Claude sees to decide when to auto-load this skill.
Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them. Use when the user asks "what should I test", "generate simulation scenarios", "write scenarios for my agent", "add a scenario for X", "organize my scenario files", "my
name: writing-livekit-scenarios description: 'Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them. Use when the user asks "what should I test", "generate simulation scenarios", "write scenarios for my agent", "add a scenario for X", "organize my scenario files", "my simulations are flaky", "the scenario hits my real database", "seed state per scenario", "it passed but booked the wrong thing", "grade the final state", or wants to stress-test a flow before shipping. Covers generating a baseline with the LiveKit scenario generator and refining it, writing the cases generation misses, phrasing simulated-user instructions so they steer reliably, splitting scenarios into sets across files, and the agent-side code that seeds deterministic state from a scenario and fails a run on its final state. To run them use running-livekit-simulations.' license: MIT metadata: author: livekit
A scenario is a simulated user's **`instructions`** (who they are, what they want) plus **`agent_expectations`** (what counts as success). A simulation plays the scenario against the real agent, and an LLM judge grades the transcript.
Runs are disposable. The scenario file is what lasts: it gets reviewed in diffs, re-run for years, and it's what catches a prompt edit that breaks something without anyone noticing.
Before writing a file, look up the exact schema and CLI flags with `reading-livekit-docs`. Field names and commands change, so this skill doesn't restate them. Conceptually, a file is a named group of scenarios. Each scenario has the simulated user's instructions, the pass criteria, and optionally tags for grouping and per-scenario data the agent can read at runtime. If the CLI offers to add stable per-scenario ids to a file, accept and commit them.
Don't start from an empty file. LiveKit's generator reads the agent's source and produces a decent first pass much faster than you'd write ten scenarios by hand:
lk agent simulate text -n 10 # confirm the exact flags with --help
**Generating from source uploads the code to LiveKit Cloud** so the generator can read it, which is why the CLI asks for confirmation first. Tell the user that before running it and let them decline. Some code can't leave the machine, and in that case you write the scenarios by hand. A flag skips the prompt for non-interactive runs.
When the run finishes, the CLI either tells you where it saved the generated scenarios or offers to save them into the project. Either way, get that file into the repo as the starting point.
You can't steer the generator. It infers intent from the code, so it writes plausible conversations instead of the ones this agent's users have, and it doesn't know what the user is worried about. Treat its output as a draft:
1. **Read every scenario** and delete the ones that don't matter. A file you haven't read isn't a test suite. 2. **Sharpen vague expectations.** If the judge can't decide an `agent_expectations`, its verdict flips between runs. This is the biggest cause of flaky scenarios. 3. **Add what generation misses.** It drifts toward happy paths. `references/risk-coverage.md` explains how to turn the agent's constraints into a checklist and give each item a scenario. 4. **Add what the user is worried about.** If they haven't said, ask what they want stress-tested. That input is why a suite a person steers is better than a generated one.
The user knows which flow keeps breaking, which customer complained, and which change they're nervous about. Ask, and weight the suite toward that area with more and deeper scenarios.
Focus adds coverage. It never removes coverage of the agent's hard limits. If the user has no preference, generate broadly and tell them that's what you did.
Read the agent's code locally with your normal tools before writing anything. Look for what a user can ask for (**capabilities**), where requests get blocked (**constraints**: required steps, unavailable items by name, caps, eligibility), and what the agent must refuse.
Constraints matter most. A scenario that asks for something the agent can't do is only valid if the expectation is that the agent *says so*. Written the other way, the test fails when the agent behaves correctly.
Scope is the agent reachable from the session entrypoint, plus agents it hands off to and tasks it awaits. Ignore other classes in the directory, unused imports, and example files.
There are two shapes, and which one you use depends on what the scenario is for. Details and examples are in `references/scenario-craft.md`.
(persona, opening line, facts revealed only when asked, steps in order, conditional reactions) keeps the persona model on track and makes the run repeatable.
conversation, where you want the simulated user to improvise.
In both shapes, goals are requests *to* the agent, never the agent's own actions. Use real values from the agent's domain and assume no prior state. Vary persona, mood and difficulty across the suite so it isn't ten copies of the same cooperative caller.
The run command takes one scenario file, so **a file is a run**. Split along the lines you want to run separately, which usually isn't topic.
The most important axis is **how the set is graded**:
the conversation (see "Make the agent consume the scenario" below).
These need
Battle-tested Claude Code, OpenAI Codex, Cursor configs, plugins, hooks and agents with Kimi, MiniMax and GLM API support.
Repo: fcakyon/claude-codex-settings
This skill should be used when the user asks for "ADHD output", "fewer output tokens", "short…
Agent-browser usage guide. Read this before running any agent-browser commands. Covers the…
Reverse-engineer a website's internal API by recording browser traffic into a HAR file, then…
Systematically explore and test a web application to find bugs, UX issues, and other…
Automate Electron desktop apps (VS Code, Slack, Discord, Figma, Notion, Spotify, etc.) using…
Build and validate experimental WebMCP tools for an existing web page. Use when an agent…