Skip to content
Data
Skill

/download-gated-pdfs

Download the actual PDF binary from bot-gated sites (taxpolicycenter.org, urban.org, SSRN-hosted mirrors, think-tank/publisher sites) via the Wayback Machine id_ URL form. Use when: (1) curl/WebFetch of a .pdf URL returns HTML instead of a PDF even with a browser User-Agent, (2)

From plugin
applied-micro-skills
2717 skills
Install
$ npx -y skills add kennethkhoocy/applied-micro-skills --skill download-gated-pdfs --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/download-gated-pdfs

Context preview

The summary Claude sees to decide when to auto-load this skill.

Download the actual PDF binary from bot-gated sites (taxpolicycenter.org, urban.org, SSRN-hosted mirrors, think-tank/publisher sites) via the Wayback Machine id_ URL form. Use when: (1) curl/WebFetch of a .pdf URL returns HTML instead of a PDF even with a browser User-Agent, (2)

SKILL.md

download-gated-pdfs.SKILL.md
name: download-gated-pdfs
description: |
  Download the actual PDF binary from bot-gated sites (taxpolicycenter.org, urban.org,
  SSRN-hosted mirrors, think-tank/publisher sites) via the Wayback Machine id_ URL form.
  Use when: (1) curl/WebFetch of a .pdf URL returns HTML instead of a PDF even with a
  browser User-Agent, (2) pypdf fails with "invalid pdf header: b'<!DOC'" or
  "EOF marker not found" on a freshly downloaded file, (3) Firecrawl can parse the PDF
  to markdown but you need the original file on disk (e.g., filing a reference copy).
author: Claude Code
version: 1.0.0
date: 2026-07-15

Download bot-gated PDFs via Wayback id_

Problem

Many think-tank and publisher sites (taxpolicycenter.org, urban.org, SSRN delivery) serve an HTML bot-challenge page instead of the PDF to non-browser clients. A browser User-Agent header does not help. The downloaded "PDF" is actually HTML.

Context / Trigger Conditions

  • `curl -o file.pdf <url>` succeeds but the file starts with `<!DOC`
  • pypdf raises `invalid pdf header: b'<!DOC'` or `PdfStreamError: Stream has ended unexpectedly`
  • Firecrawl `scrape` returns clean markdown for the same URL (its proxies get through),

but Firecrawl does not return the binary — only parsed content

Solution

1. Request the file through the Wayback Machine's raw-content (`id_`) endpoint, which serves the original archived binary without rewriting:

   curl -sL -A "Mozilla/5.0 ... Chrome/126.0 Safari/537.36" \
     "https://web.archive.org/web/<YYYY>id_/<original-pdf-url>" -o out.pdf

`<YYYY>` is any year likely to have a snapshot (e.g. publication year); Wayback redirects to the nearest capture. The `id_` suffix after the timestamp is what requests the untouched original. 2. Verify the download with pypdf — a bot page fails immediately:

   from pypdf import PdfReader
   r = PdfReader("out.pdf"); print(len(r.pages), "pages")

3. If Wayback has no capture, fall back to: another mirror found via search (Exa/Firecrawl), or Firecrawl scrape for the parsed text when the binary is not strictly needed.

Verification

`PdfReader` opens the file and reports a plausible page count; first-page text matches the expected title.

Example

Verified 2026-07-15: `taxpolicycenter.org/sites/default/files/publication/165884/ssrn-id4797771.pdf` and `urban.org/sites/default/files/publication/80621/2000790-...pdf` both bot-gated to direct curl (with UA), both downloaded intact via `https://web.archive.org/web/2024id_/<url>` and `.../web/2023id_/<url>` (18 and 12 pages).

Notes

  • Government data hosts (e.g. `ticdata.treasury.gov`) are usually NOT gated — try direct

curl first; Wayback is the fallback, not the default.

  • Wayback captures can be stale for frequently-revised documents; check the snapshot date

if currency matters.

  • See also: the `pdf` skill (parsing/extraction after download) and

`firecrawl:firecrawl-scrape` (parsed markdown when the binary is unnecessary).

Read more
Ships withapplied-micro-skills

Claude Code and Codex skills for empirical applied-microeconomics research: reproducibility auditing, LLM-assisted classification methods, event studies, data infrastructure (WRDS, Stata, pyfixest), and publication-grade tables, figures, and documents.

Get the whole plugin

Other skills on applied-micro-skills.