adr
Write an Architecture Decision Record documenting a significant technical decision.
Evaluate machine learning model performance with comprehensive metrics.
$ npx -y skills add rohitg00/awesome-claude-code-toolkit --agent claude-codeHow it fires
How this command gets triggered: by you, by Claude, or both.
/evaluate-modelContext preview
What this command does when you run it.
Evaluate machine learning model performance with comprehensive metrics.
Evaluate machine learning model performance with comprehensive metrics.
1. Ask the user for the model type: classification, regression, NLP, or generative 2. Load the model and test dataset from the specified paths 3. Run inference on the entire test dataset and collect predictions 4. For classification models, calculate: accuracy, precision, recall, F1-score, AUC-ROC 5. For regression models, calculate: MAE, MSE, RMSE, R-squared, MAPE 6. For NLP models, calculate: BLEU, ROUGE, perplexity, exact match 7. Generate a confusion matrix for classification tasks 8. Identify the worst-performing classes or data segments 9. Calculate calibration metrics: expected calibration error 10. Run performance profiling: inference time per sample, memory usage, throughput 11. Check for bias: evaluate performance across demographic subgroups if applicable 12. Generate a comprehensive evaluation report with all metrics and visualizations
The most comprehensive toolkit for Claude Code -- 135 agents, 35 curated skills (+400,000 via SkillKit), 42 commands, 176+ plugins, 20 hooks, 15 rules, 7 templates, 15 MCP configs, 26 companion apps, 53 ecosystem entries, and more.
Repo: rohitg00/awesome-claude-code-toolkit
Write an Architecture Decision Record documenting a significant technical decision.
Conduct a structured design review of a module, feature, or system component.
Create a structured implementation plan for the requested feature or change.