Skip to content
Development
Skill

/research-critique

Critical analysis of research papers evaluating methodology, claims-evidence alignment, and contribution significance. Triggers on: "critique this paper", "review this research", "analyze this study", "evaluate the methodology", "is this paper sound". NOT for formatting or

From plugin
armory
31181 skills2 agents1 command
Install
$ npx -y skills add Mathews-Tom/armory --skill research-critique --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/research-critique

Context preview

The summary Claude sees to decide when to auto-load this skill.

Critical analysis of research papers evaluating methodology, claims-evidence alignment, and contribution significance. Triggers on: "critique this paper", "review this research", "analyze this study", "evaluate the methodology", "is this paper sound". NOT for formatting or

SKILL.md

research-critique.SKILL.md
name: research-critique
description: 'Critical analysis of research papers evaluating methodology, claims-evidence alignment, and contribution significance. Triggers on: "critique this paper", "review this research", "analyze this study", "evaluate the methodology", "is this paper sound". NOT for formatting or submission readiness, use manuscript-review.'
metadata:
  version: 1.0.1
  category: review
  tags: [research, methodology, claims-evidence, peer-review]
  difficulty: intermediate
  phase: review

Research Paper Critique

Purpose

Critically evaluate research papers as an engaged, intellectually honest colleague — not an adversarial reviewer generating decorative objections. The goal is understanding what a paper contributes, whether the evidence supports the claims, and where genuine tension exists between what is demonstrated and what is argued.

The failure mode you must avoid

The default behavior when critiquing a paper is adversarial credentialism: treating the paper as something to defeat, generating a volume of objections to demonstrate rigor, and mistaking the enumeration of limitations for insight. This produces critiques that are shallow, performative, and often incoherent — holding papers to standards that no published work in the field meets, demanding ablations that would constitute separate papers, and flagging limitations the author already disclosed as if discovering them.

The tell: if you retreat easily from a point when challenged, you were never committed to it. It was decorative, not substantive. Do not generate points you would abandon under pressure.

Analytical principles

1. Understand before you evaluate

Read the paper as the author intended it. What problem does it identify? What is the proposed mechanism or contribution? What is the experimental design trying to isolate? What claims does the paper actually make, versus what you assume it is claiming?

Most bad critiques attack a paper the author did not write. Get the real paper straight first.

2. Proportionality

Every empirical paper has limitations. Uncontrolled confounds, limited domains, incomplete ablations, finite sample sizes. These exist in every paper ever published, including the landmarks of the field. Listing them is not critique. The question is always: **does this limitation actually undermine the paper's specific contribution?**

A limitation that the author discloses, contextualizes, and accounts for in their claims is not a flaw you discovered. A limitation that would require a different experiment entirely to resolve is not a flaw — it is future work. A limitation that exists in every comparable study is not a differentiating weakness.

Weight your critique by consequence. A confound that could fully explain the main result matters. A confound that affects a secondary finding the paper does not emphasize does not.

3. No phantom standards

Do not hold the paper against an imaginary ideal that nothing in the field achieves. Before asserting that the paper should have done X, ask: does any comparable work do X? If the answer is no, then you are requesting a methodological advance, not identifying a flaw. Name it as such.

The implicit comparison should be the realistic state of the field, not a platonic ideal of experimental design.

4. Mechanism versus evidence

Many papers propose a mechanism (why something works) and provide evidence (that it works). These are different claims with different standards. The evidence can be solid while the mechanistic story is underdetermined. This is normal and nearly universal in empirical work.

When critiquing the mechanism: identify what alternative explanations the data cannot distinguish between. Do not substitute your preferred alternative mechanism and assert it is "likely" doing the work — that is the same sin as the paper's, just pointing the other direction. If you cannot separate two explanations, say so honestly rather than privileging one.

5. Scope is not a weakness

A paper that tests in one domain and theorizes about generality is doing what papers do. The theory predicts what should happen elsewhere. Someone else can test it. Demanding that a single paper demonstrate generality across all domains it might apply to is demanding a research program, not a paper. Critique the scope if the paper overclaims relative to its evidence. Do not critique it for having a scope.

6. What survives

After your analysis, state clearly: what does this paper contribute that holds up? What finding, if confirmed, would the field benefit from knowing? Default to identifying the contribution, not to dismissal. A paper that advances understanding in a limited domain with honest limitations is a good paper. Say so.

Output structure

Do not produce a numbered checklist of weaknesses. Instead, write a coherent analytical response that:

  • Opens with what the paper is doing and why it matters (or does not)
  • Identifies where the evidence genuinely supports the claims
  • Identifies where the claims outrun the evidence, with specificity about _why_ and _how much_
  • Distinguishes between limitations that threaten the core contribution and limitations that are standard for the field
  • Closes with an honest assessment of what survives scrutiny

Tone: engaged, direct, respectful of the intellectual effort. You are a colleague reading carefully, not an adversary scanning for weaknesses.

Worked excerpt

The following shows the expected tone and structure — not a template to fill in.

> This paper proposes that chain-of-thought prompting improves arithmetic reasoning primarily through decomposition rather than retrieval, testing across three digit-multiplication benchmarks. The core evidence is solid: the ablation removing intermediate steps while preserving final-answer prompting shows a consistent 12-18% accuracy drop across all three benchmarks, which directly supports the decomposition hypothesis. > > Where the p

Read more
Ships witharmory

Curated, production-grade skills, agents, hooks, rules, commands, utilities, and presets for AI coding agents. No magic, no demos — battle-tested workflows built for developers who use AI seriously.

Get the whole plugin

Other skills on armory.