acquisition-channels
Help users identify unique distribution advantages and master the lifecycle of acquisition…
Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.
$ npx -y skills add RefoundAI/lenny-skills --skill ai-evals --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/ai-evalsContext preview
The summary Claude sees to decide when to auto-load this skill.
Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.
name: ai-evals description: Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.
Move beyond vibe checks to systematic, empirical measurement of AI product quality and reliability.
Help the user with ai evaluation strategy using insights from 11 guests and posts across Lenny's Podcast and Newsletter.
1. **Identify Failure Modes** - Help the user conduct error analysis on real traces to find where the system specifically breaks. 2. **Select Eval Methods** - Recommend the right mix of human, code, and LLM judges based on the specific technical use case. 3. **Build Gold Sets** - Assist in curating a reference dataset of high-quality examples to act as the ground truth for your application. 4. **Operationalize** - Guide the user in integrating these evaluations into a CI/CD pipeline for continuous quality improvement.
Brendan Foody: "I think that for enterprises especially, the core way to think about it is how can they build a test or systematic way to measure how well AI automates their core value chain? So if it's an architecture firm that's producing these architecture diagrams of what they provide to their end customer, how can they effectively measure that? And each company has its own value chain or maybe a handful of them if it's a multi-product company."
Identify the core deliverables unique to your business and develop systematic tests to measure how accurately AI can replicate those specific tasks.
Edwin Chen: "We are looking for a Nobel Prize-winning poetry. Is this poetry unique? Is it full of subtle imagery? Does it surprise you and target your heart? Does it teach you something about the nature of moonlight?"
True data quality is defined by deep, subjective human excellence, such as emotional resonance and uniqueness, rather than superficial binary checks.
Hamel Husain & Shreya Shankar: "Evals help you create metrics that you can use to measure how your application is doing and kind of give you a way to improve your application with confidence. That you have a feedback signal in which to iterate against."
Create systematic metrics to track application quality over time, allowing teams to iterate on prompts or models with the same confidence as traditional software.
From "Beyond vibe checks: A PM’s complete guide to evals": "Clearly articulating what you want your judge-LLM to measure isn’t just a step in the process; it’s the difference between a mediocre AI and one that consistently delights users. Building these writing skills requires practice and attention."
Write effective automated evaluations by using a structured prompt that defines the role, data, success criteria, and specific labels for the judge.
See `references/artifacts.md` for the full list with details.
76 product management and engineering skills, distilled from the full archive of Lenny's Podcast and Lenny's Newsletter: 597 episodes and posts, 4,019 sourced insights, every quote verified verbatim against its source. Curated by Refound AI.
Help users identify unique distribution advantages and master the lifecycle of acquisition…
Help users build functional product prototypes from natural language or visual mocks using AI…
Help users interact with probabilistic models by designing interfaces that manage fluidity,…
Help users decide where to apply AI effectively, manage the transition from deterministic to…
Help users convert massive volumes of qualitative and quantitative feedback into…
Help users identify the most viable path into product management, assess their role fit, and…