๐ Large variety of ready-to-use LLM eval metrics (all with explanations) powered by ANY LLM of your choice, statistical methods, or NLP models that run locally on your machine covering all use cases: Custom, All-Purpose Metrics: G-Eval โ a research-backed
> /plugin marketplace add confident-ai/deepeval> /plugin install deepeval@deepeval-plugins
Repo: confident-ai/deepeval
What's inside
DeepEval is a simple-to-use, open-source LLM evaluation framework, for evaluating large-language model systems. It is similar to Pytest but specialized for unit testing LLM apps. DeepEval incorporates the latest research to run evals via metrics such as G-Eval, task completion, answer relevancy, hallucination, etc., which use LLM-as-a-judge and other NLP models that run locally on your machine.
Whether you're building AI agents, RAG pipelines, or chatbots, implemented via LangChain or OpenAI, DeepEval has you covered. With it, you can easily evaluate:
Use these evaluations to determine the optimal models, prompts, and architecture to improve your AI quality, prevent prompt drifting, or even transition from OpenAI to Claude with confidence.
[!IMPORTANT] Want to compare iterations, share evaluation reports, and monitor your AI in production? Sign up for Confident AI, the enterprise AI evals and observability platform.
Want to talk LLM evaluation, need help picking metrics, or just to say hi? Come join our discord.
๐ Large variety of ready-to-use LLM eval metrics (all with explanations) powered by ANY LLM of your choice, statistical methods, or NLP models that run locally on your machine covering all use cases:
Custom, All-Purpose Metrics:
๐ฏ Supports both end-to-end and component-level LLM evaluation.
๐งฉ Build your own custom metrics that are automatically integrated with DeepEval's ecosystem.
๐ฎ Generate both single and multi-turn synthetic datasets for evaluation.
๐ Integrates seamlessly with ANY CI/CD environment.
๐งฌ Optimize prompts automatically based on evaluation results.
๐ Easily benchmark ANY LLM on popular LLM benchmarks in under 10 lines of code., including MMLU, HellaSwag, DROP, BIG-Bench Hard, TruthfulQA, HumanEval, GSM8K.
DeepEval plugs into any LLM framework โ OpenAI Agents, LangChain, CrewAI, and more. For enterprise teams standardizing evals and observability across the organization, Confident AI provides a native DeepEval integration.
Showing a partial view of a very large repo.
FAQ
deepeval is a Claude Code plugin with 3 hand-picked skills for testing work, indexed on Flowy. Install it with the command on its page. It includes deepeval-otel, deepeval-tracing, deepeval. Its skills do not fire on their own yet. Request auto-invocation to have Flowy route them as you prompt. Free and open source.
Is this plugin yours?
Claim it with GitHubSubmit a pluginPromote it