advanced-evaluation
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring,…
This skill should be used when a model gets read-write control over its own live context window instead of a harness-scheduled compaction policy: the context exposed as an editable file the model rewrites with code tools, model-driven eviction and in-place updates, the harness
$ npx -y skills add muratcankoylan/agent-skills-for-context-engineering --skill self-managed-context --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/self-managed-contextContext preview
The summary Claude sees to decide when to auto-load this skill.
This skill should be used when a model gets read-write control over its own live context window instead of a harness-scheduled compaction policy: the context exposed as an editable file the model rewrites with code tools, model-driven eviction and in-place updates, the harness
name: self-managed-context description: "This skill should be used when a model gets read-write control over its own live context window instead of a harness-scheduled compaction policy: the context exposed as an editable file the model rewrites with code tools, model-driven eviction and in-place updates, the harness invariants that keep self-editing safe (pinned prefix, edit gate, edit receipts, budget readouts, rollback on overflow), the prefix-cache cost of mid-context edits, and steering or training the model's own context-editing strategy. Route note content and fixed-threshold summarization to context-compression, cache-stable layout under harness control to context-optimization, file offloading to filesystem-context, and loop governance to self-improvement-loops."
This skill covers agents that manage their own context window: the model, not the harness, decides what stays in the live context, what is compacted, what is evicted, and what is rewritten in place. The reference implementation is Context Language Models (CLMs), which mirror the editable part of the conversation into a file that the model edits with ordinary code tools, then re-parse that file into the next prompt. Applied zero-shot to existing models on a shared agent backbone, this matched or beat harness-scheduled and action-based context management on accuracy and compute across deep-research, terminal-coding, and multi-hour repository-optimization tasks (claim-self-managed-context-zero-shot-results, claim-self-managed-context-long-horizon-results).
The controlling trade-off: model control buys adaptivity (verbatim retention, surgical in-place updates, eviction on demand) and creates three problems the harness must absorb. Edits break prefix caching, so a badly placed edit can cost more compute than the tokens it frees. Models do not know how full their context is. And a model-writable context is a persistence channel for whatever the model writes into it, including instructions. Most of this skill is the harness contract that makes the first property worth the other three.
Activate this skill when:
Do not activate this skill for adjacent work owned by other skills:
Context-management designs differ in who decides the transition from one context to the next. Append-only agents extend the context each step; a self-managing agent produces the whole next context as a function of the current one.
| Level | Who decides | Examples | Characteristic failure | | --- | --- | --- | --- | | Harness-scheduled | Harness, at a threshold or every turn | Threshold summarization, per-turn state rewrite | Wrong timing; summaries lose or invent verbatim state | | Action-based | Model chooses when; harness defines what | Self-compaction tools, offload-and-retrieve tools, context folding | Strategy bounded by the action set; offload without eviction | | Model-controlled | Model chooses when and what, with general tools | Context as an editable file | Edit cache cost, budget blindness, persisting self-written instructions |
A diagnostic built to isolate context management from reasoning (verbatim retention, in-place board updates, offload-and-evict) found no fixed strategy perfect even on simple synthetic tasks (claim-self-managed-context-contextbench-pilot). Each failure maps to a missing capability: summaries cannot hold exact values, append-only designs re-emit full state for every small update, and tool-based offloading cannot remove the original from the window.
The implementation needs no new model capability:
1. Split the context into a pinned prefix (system prompt and task) and an editable region (every later turn). 2. Before each command, render the editable region to a file, each turn preceded by a header such as `[[CTX_TURN 7 role=tool]]`. 3. The model edits the file with the same shell and code tools it uses for the task, locating text with code (header regexes, unique anchors) instead of retyping it. 4. After the command, re-read the file, parse it tolerantly back into messages, and replace the live context.
The load-bearing choice is general tools, not context tools. There is no summarize action and no eviction API; the model can delete, rewrite, merge, or annotate. Observed behaviors went beyond any predefined action set: a model-invented `notes` role, a self-defined compaction helper reused across a run, and an in-context subagent scoreboard maintained through in-place edits at small context size (claim-self-managed-context-emergent-behaviors).
Unrestricted edits are safe only because the safety contract lives outside the editable region.
| Invariant |
A comprehensive, open collection of Agent Skills focused on context engineering and harness engineering principles for building production-grade AI agent systems.
Repo: muratcankoylan/agent-skills-for-context-engineering
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring,…
This skill should be used when modeling agent mental states with BDI concepts: beliefs,…
This skill should be used when long-running agent sessions need context compression,…
This skill should be used for diagnosing and mitigating context degradation: lost-in-middle…
This skill should be used to explain or reason about the foundational concepts of context…
This skill should be used for improving context efficiency: context budgeting, observation…