adhd-output-style
This skill should be used when the user asks for "ADHD output", "fewer output tokens", "short…
Writes turn-level tests for a LiveKit agent in the user''s normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after
$ npx -y skills add fcakyon/claude-codex-settings --skill testing-livekit-agents --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/testing-livekit-agentsContext preview
The summary Claude sees to decide when to auto-load this skill.
Writes turn-level tests for a LiveKit agent in the user''s normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after
name: testing-livekit-agents description: 'Writes turn-level tests for a LiveKit agent in the user''s normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after building or changing agent behavior that needs regression coverage. Covers the SDK''s test session harness, assertions on messages, tool calls and handoffs, LLM judging of a reply against an intent, mocking tools, multi-turn tests, and judging whole conversations with the built-in judges. For interactive poking use debugging-livekit-agents. For grading whole conversations at scale use running-livekit-simulations.' license: MIT metadata: author: livekit
Turn-level tests are the cheapest lasting verification an agent can have. They run in the user's existing suite, in text mode, and they're fast enough for every commit. The framework's helper names and signatures change, so look them up with `reading-livekit-docs` before writing, and use the testing docs page for every API detail this skill leaves out.
Every test has the same shape: start a test session with the agent under test, run one user turn, and assert on the events that turn produced.
A turn produces a sequence of events. A simple turn is one message. A more typical one is a tool call, its output, maybe a handoff, and then a message. You write the test by walking that sequence in order, asserting on each event, and then asserting the turn has nothing more.
There are three kinds of assertion, and most of this skill is knowing which one to use.
**Structural** assertions check message roles, that a tool was called, its arguments, what it returned, and that a handoff to a specific agent happened. They're deterministic and fail for exactly one reason, so prefer them.
**Judged** assertions use an LLM-judge helper that gives one message and an intent string to a model and asks whether they match. Use them for the content of a reply, which you can't assert exactly. Describe the intent by outcome ("tells the user the booking is confirmed and gives the time"), not by wording.
**Whole-conversation** assertions use a judge-group helper that runs several built-in judges concurrently over the whole chat history and aggregates their verdicts. The built-in judges cover dimensions like grounding, relevance, safety, task completion, and tool use; the docs list the current set. Use this when the question spans several turns.
Rules of thumb:
structural is slower and flakier.
also does something you didn't intend.
Tests shouldn't hit real backends, since that makes a suite slow and nondeterministic. Override the tools for the agent under test and return fixed values.
Two things beyond the API:
lets you test what the agent says when a backend is down. Most agents lack this test, and it's the behavior users notice most.
a test needs. A session-scoped form also exists for a session that runs on its own and needs mocks active for its whole lifetime. Simulation entrypoints use that form, tests don't. See `writing-livekit-scenarios`.
Mocking only changes execution. The model still sees the real tool schemas, so tool selection is still under test.
There are two ways to test behavior that depends on earlier turns:
Use this when the path matters: collecting details across turns, the user changing their mind mid-flow, a handoff carrying context.
care about. Use this when the setup turns aren't what you're testing. It's faster and less brittle than replaying five turns to reach the sixth.
Roughly in order of value:
1. **The behavior the user asked for.** Every agent needs at least this. 2. **Tool invocation**: the right tool with the right arguments for a representative request. 3. **Tool failure**: what the agent says when a tool errors or returns nothing. 4. **Refusals and limits**: the agent declines what it should decline and doesn't make up data it can't have. The refusal is the pass condition. 5. **Handoffs**: the transition fires when it should, and the next agent has what it needs. 6. **Every bug you've fixed.** Without a test, a fixed bug can come back unnoticed.
A test suite for an agent that does things — books, edits, confirms — needs three kinds of test, and confusing them is how a green suite ships a broken agent:
storage failure, and retries, against the application core directly. Fast, exact, no model.
listener fires, that the tool receives what the session sends. A scripted provider proves *mechanics*, never language understanding — its exact sentences are fixtures, not requirements.
callers actually talk: paraphrases absent from the prompt, corrections mixed with agreement, the same short reply answering different questions.
Shortcuts that look lik
Battle-tested Claude Code, OpenAI Codex, Cursor configs, plugins, hooks and agents with Kimi, MiniMax and GLM API support.
Repo: fcakyon/claude-codex-settings
This skill should be used when the user asks for "ADHD output", "fewer output tokens", "short…
Agent-browser usage guide. Read this before running any agent-browser commands. Covers the…
Reverse-engineer a website's internal API by recording browser traffic into a HAR file, then…
Systematically explore and test a web application to find bugs, UX issues, and other…
Automate Electron desktop apps (VS Code, Slack, Discord, Figma, Notion, Spotify, etc.) using…
Build and validate experimental WebMCP tools for an existing web page. Use when an agent…