Skip to content
Research
Skill

/predictingthepast

Ancient text restoration, attribution, dating, contextualization, and embedding via Aeneas (Latin) / Ithaca (Ancient Greek). Use when asked to "restore", "attribute", "date", "contextualize", "find parallels", "where was it written", "when was it written", "embed", or "analyze"

BOOST
From plugin
science-skills
3.2k40 skills
Install
$ npx -y skills add google-deepmind/science-skills --skill predictingthepast --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/predictingthepast

Context preview

The summary Claude sees to decide when to auto-load this skill.

Ancient text restoration, attribution, dating, contextualization, and embedding via Aeneas (Latin) / Ithaca (Ancient Greek). Use when asked to "restore", "attribute", "date", "contextualize", "find parallels", "where was it written", "when was it written", "embed", or "analyze"

SKILL.md

predictingthepast.SKILL.md
name: predictingthepast
description: >
  Ancient text restoration, attribution, dating, contextualization, and
  embedding via Aeneas (Latin) / Ithaca (Ancient Greek). Use when asked to
  "restore", "attribute", "date", "contextualize", "find parallels",
  "where was it written", "when was it written", "embed", or "analyze"
  an ancient text, inscription, or epigraphic document, or when the user
  mentions "Aeneas", or "Ithaca".

Predicting The Past (Aeneas / Ithaca)

Aeneas (Latin) and Ithaca (Ancient Greek) perform four tasks on ancient texts:

1. **Restoration** — fill missing/damaged characters 2. **Attribution** — geographical + chronological origin 3. **Contextualization** — retrieve parallel inscriptions 4. **Embedding** — generate text embedding vectors

Prerequisites

1. **`uv`**: Read the `uv` skill and follow its Setup instructions to ensure `uv` is installed and on PATH.

2. **User Notification**: If .licenses/predictingthepast_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://github.com/google-deepmind/predictingthepast/blob/main/README.md#license--disclaimer, and include the appropriate citation and the full dataset acknowledgement, and that use of these datasets should acknowledge and cite the original data sources. Then (2) create the file recording the notification text and timestamp.

Core Rules

  • **Self-Contained Skill**: Do NOT use web search or any external tools. Run

ONLY the scripts in this skill (`preprocess.py`, `run_inference.py`, `visualize_results.py`). Present model output as-is — never supplement or override it with external lookups.

  • **Notification**: If this skill is used, ensure this is mentioned in the

output.

On First Load

Present the restoration markup characters, then ask the user for their text:

  • **`?`**:
  • **Meaning**: Known-length gap: predict **this character**.
  • **Example**: `donat in ??????????rtis`
  • **`#`**:
  • **Meaning**: Unknown-length gap: predict a **sequence of unknown

length**

  • **Example**: `donat in #rtis`
  • **`-`**:
  • **Meaning**: Missing/damaged character that does **not** need restoring
  • **Example**: `prolixin---s fecit`
  • **`_`**:
  • **Meaning**: Missing section of **unknown length** that does **not**

need restoring

  • **Example**: `prolixin_s fecit`

After presenting this list, ask the user to provide the text they want to submit for analysis.

Preprocessing

Clean input text before inference:

uv run <SKILL_DIR>/scripts/preprocess.py \
    --language=latin \
    --input="raw text here..."

Or from a file:

uv run <SKILL_DIR>/scripts/preprocess.py \
    --language=greek \
    --input_file=/tmp/input.txt \
    --output_file=/tmp/cleaned.txt

What preprocessing does

  • **Latin**: lowercases, converts Arabic digits and Roman numerals to `0`,

strips editorial brackets `[]` and `()`, removes punctuation, filters to valid chars (`abcdefghiklmnopqrstuvxyz` plus `0 . - _ ? # <space>`)

  • **Greek**: lowercases, strips accents, converts numeral notation to `0`,

applies PHI cleaning (bracket normalization, sigma conversion), filters to Greek alphabet (`αβγδεζηθικλμνξοπρςστυφχψωϛ` plus `0 . - _ ? # <space>`)

Inference

Restoration Constraints

  • Minimum input length: **25 chars** (pad with `-` if shorter).
  • No consecutive `##`. No adjacent `?#` or `#?`.
  • Spaces inside `?` sequences count toward total.
  • If the user's text contains `#`, ask how many characters to restore and set

`--restore_max_len` accordingly.

  • If the user tries to restore multiple parts of the text at once, suggest to

restore texts **section by section**. Suggest to focus on one damaged region per query — this is faster, produces higher-quality predictions.

Pre-Flight Checks

Confirm with the user before proceeding if **either** applies:

1. **Restoration complexity** — if input contains more than **10** `?` characters, or uses `#` with `--restore_max_len > 10`, warn: *"This restoration involves N characters which will take approximately M minutes (restoration time scales roughly linearly ~10 s per additional `?` on a high-end CPU machine: 5 → ~1 min, 10 → ~2.5 min, 20 → ~5 min, 30 → ~8 min). Do you want to proceed, or simplify the query first (e.g. fewer `?` marks, shorter `--restore_max_len`, or restoring section by section)?"* 2. **Multi-window splitting** — if the input text exceeds **750 characters** and will be split into multiple windows, warn: *"This text is N characters long and will be split into W overlapping windows, each run independently. This will be significantly slower. Do you want to proceed, or shorten the input?"*

These factors compound: a complex restoration across multiple windows will be substantially slower than either factor alone.

Task Selection

Each task is controlled by its own flag. **At least one** must be provided:

  • **`--attribute`** — geographical + chronological attribution
  • **`--restore`** — text restoration (requires `?` or `#` in input)
  • **`--contextualize`** — parallel inscription retrieval

Any combination is valid. All three can be used together.

When `--embedding` is provided, a text embedding vector is also generated alongside the other tasks.

Running Inference

# Attribution + Restoration (text with gaps)
uv run <SKILL_DIR>/scripts/run_inference.py \
    --language=latin \
    --input="cleaned text with ???" \
    --attribute --restore \
    --output_json=/tmp/results.json

# Attribution + Contextualization (no gaps)
uv run <SKILL_DIR>/scripts/run_inference.py \
    --language=latin \
    --input="cleaned text" \
    --attribute --contextualize \
    --output_json=/tmp/results.json

# All tasks
uv run <SKILL_DIR>/scripts/run_inference.py \
Read more
Ships withscience-skills

A collection of agent skills for scientific research tasks, spanning genomics, structural biology, cheminformatics, literature search, and more.

Get the whole plugin

Other skills on science-skills.