/bib-verify
Verify a BibTeX file for hallucinated or fabricated references by cross-checking every entry against CrossRef, arXiv, and DBLP. Reports each reference as verified, suspect, or not found, with field-level mismatch details (title, authors, year, DOI). Use when the user wants to
$ npx -y skills add agentscope-ai/OpenJudge --skill bib-verify --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/bib-verify
Context preview
The summary Claude sees to decide when to auto-load this skill.
Verify a BibTeX file for hallucinated or fabricated references by cross-checking every entry against CrossRef, arXiv, and DBLP. Reports each reference as verified, suspect, or not found, with field-level mismatch details (title, authors, year, DOI). Use when the user wants to
SKILL.md
bib-verify.SKILL.mdname: bib-verify
description: >
Verify a BibTeX file for hallucinated or fabricated references by cross-checking
every entry against CrossRef, arXiv, and DBLP. Reports each reference as
verified, suspect, or not found, with field-level mismatch details (title,
authors, year, DOI). Use when the user wants to check a .bib file for fake
citations, validate references in a paper, or audit bibliography entries for
accuracy.
BibTeX Verification Skill
Check every entry in a `.bib` file against real academic databases using the OpenJudge `PaperReviewPipeline` in BibTeX-only mode:
1. **Parse** — extract all entries from the `.bib` file 2. **Lookup** — query CrossRef, arXiv, and DBLP for each reference 3. **Match** — compare title, authors, year, and DOI 4. **Report** — flag each entry as `verified`, `suspect`, or `not_found`
Prerequisites
pip install py-openjudge litellm
Gather from user before running
| Info | Required? | Notes | |------|-----------|-------| | BibTeX file path | Yes | `.bib` file to verify | | CrossRef email | No | Improves CrossRef API rate limits |
Quick start
# Verify a standalone .bib file
python -m cookbooks.paper_review --bib_only references.bib
# With CrossRef email for better rate limits
python -m cookbooks.paper_review --bib_only references.bib --email your@email.com
# Save report to a custom path
python -m cookbooks.paper_review --bib_only references.bib \
--email your@email.com --output bib_report.md
Relevant options
| Flag | Default | Description | |------|---------|-------------| | `--bib_only` | — | Path to `.bib` file (required for standalone verification) | | `--email` | — | CrossRef mailto — improves rate limits, recommended | | `--output` | auto | Output `.md` report path | | `--language` | `en` | Report language: `en` or `zh` |
Interpreting results
Each reference entry is assigned one of three statuses:
| Status | Meaning | |--------|---------| | `verified` | Found in CrossRef / arXiv / DBLP with matching fields | | `suspect` | Title or authors do not match any real paper — likely hallucinated or mis-cited | | `not_found` | No match in any database — treat as fabricated |
**Field-level details** are shown for `suspect` entries:
- `title_match` — whether the title matches a real paper
- `author_match` — whether the author list matches
- `year_match` — whether the publication year is correct
- `doi_match` — whether the DOI resolves to the right paper
Additional resources
- Full pipeline options: [../paper-review/reference.md](../paper-review/reference.md)
- Combined PDF review + BibTeX verification: [../paper-review/SKILL.md](../paper-review/SKILL.md)
Read more
name: bib-verify description: > Verify a BibTeX file for hallucinated or fabricated references by cross-checking every entry against CrossRef, arXiv, and DBLP. Reports each reference as verified, suspect, or not found, with field-level mismatch details (title, authors, year, DOI). Use when the user wants to check a .bib file for fake citations, validate references in a paper, or audit bibliography entries for accuracy.
BibTeX Verification Skill
Check every entry in a `.bib` file against real academic databases using the OpenJudge `PaperReviewPipeline` in BibTeX-only mode:
1. **Parse** — extract all entries from the `.bib` file 2. **Lookup** — query CrossRef, arXiv, and DBLP for each reference 3. **Match** — compare title, authors, year, and DOI 4. **Report** — flag each entry as `verified`, `suspect`, or `not_found`
Prerequisites
pip install py-openjudge litellm
Gather from user before running
| Info | Required? | Notes | |------|-----------|-------| | BibTeX file path | Yes | `.bib` file to verify | | CrossRef email | No | Improves CrossRef API rate limits |
Quick start
# Verify a standalone .bib file python -m cookbooks.paper_review --bib_only references.bib # With CrossRef email for better rate limits python -m cookbooks.paper_review --bib_only references.bib --email your@email.com # Save report to a custom path python -m cookbooks.paper_review --bib_only references.bib \ --email your@email.com --output bib_report.md
Relevant options
| Flag | Default | Description | |------|---------|-------------| | `--bib_only` | — | Path to `.bib` file (required for standalone verification) | | `--email` | — | CrossRef mailto — improves rate limits, recommended | | `--output` | auto | Output `.md` report path | | `--language` | `en` | Report language: `en` or `zh` |
Interpreting results
Each reference entry is assigned one of three statuses:
| Status | Meaning | |--------|---------| | `verified` | Found in CrossRef / arXiv / DBLP with matching fields | | `suspect` | Title or authors do not match any real paper — likely hallucinated or mis-cited | | `not_found` | No match in any database — treat as fabricated |
**Field-level details** are shown for `suspect` entries:
- `title_match` — whether the title matches a real paper
- `author_match` — whether the author list matches
- `year_match` — whether the publication year is correct
- `doi_match` — whether the DOI resolves to the right paper
Additional resources
- Full pipeline options: [../paper-review/reference.md](../paper-review/reference.md)
- Combined PDF review + BibTeX verification: [../paper-review/SKILL.md](../paper-review/SKILL.md)
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
Other skills on openjudge.
- /auto-arena
Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and
Open skill - /claude-authenticity
Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project. Also extracts injected system prompts from providers that override Claude's identity. Fully self-contained
Open skill - /00-meta-eval
Use when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing at all. Also use when the user mentions evaluation, eval, benchmarking, testing LLM quality, measuring agent
Open skill - /01-eval-design
Use when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a labeled evaluation set. Also use when the user mentions test data design, eval coverage, difficulty
Open skill - /02-metric-design
Use when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining multiple metrics into a composite score, or building an automated evaluation pipeline. Also use when the
Open skill - /03-align-human
Use when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic evaluation can replace human review, or build a human-reduction roadmap. Also use when the user mentions
Open skill

