agent-transcript
Requested GitHub PR/issue agent transcripts: redact, trim, preview, and insert safely.
Codex 1M context: direct OpenAI Responses API inference, safe Astra/Sol/Terra/Luna input headroom, Keychain delivery, and Mac fleet rollout.
$ npx -y skills add steipete/agent-scripts --skill codex-huge-context --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/codex-huge-contextContext preview
The summary Claude sees to decide when to auto-load this skill.
Codex 1M context: direct OpenAI Responses API inference, safe Astra/Sol/Terra/Luna input headroom, Keychain delivery, and Mac fleet rollout.
name: codex-huge-context description: "Codex 1M context: direct OpenAI Responses API inference, safe Astra/Sol/Terra/Luna input headroom, Keychain delivery, and Mac fleet rollout."
Use this skill when configuring, repairing, or auditing Codex's one-million-token context setup. The intended topology is a direct API inference route that preserves the normal ChatGPT login for Gmail, Calendar, and other connector OAuth:
Codex inference -> Keychain auth helper -> https://api.openai.com/v1/responses Codex connectors -> normal ChatGPT login in auth.json
This is not an HTTP proxy. The API remains authoritative for access, actual model limits, and billing.
Treat the provider selector, context window, compaction threshold, and custom model catalogue as one atomic configuration. In Peter's normal ChatGPT-authenticated setup, `model_provider = "openai"` selects the ChatGPT-backed route; never attach this skill's direct-API `922000` context window and `700000` compaction threshold to that route. The provider identifier alone does not determine transport: built-in `openai` with API-key authentication can also use the API. This skill and its preflight require the named `openai_api_direct` provider below so inference uses its dedicated Keychain helper while connector login stays separate.
That split configuration can leave a 700,000-token compaction threshold attached to a smaller provider or model-metadata limit. Codex may also clamp the requested window, so root numbers alone do not prove the usable context. A browser- or Computer Use-heavy thread can grow past the provider's real limit before compaction, receive `context_length_exceeded`, and become unable to compact because the compaction request itself no longer fits. Verify the fresh session's reported effective window as well as the files on disk.
The required preflight treats this mismatch as fatal. Do not launch or resume Codex after any config writer, app settings change, model change, or fleet sync until the preflight passes. If it reports an unsafe split configuration, restore `model_provider = "openai_api_direct"`, restart every Codex desktop/shared app server, and start or fork a fresh thread. Resuming the failed thread preserves its recorded provider.
[GPT-6 Astra](https://developers.openai.com/api/docs/models/gpt-6-astra) exposes a 1,050,000-token total context window and can produce up to 128,000 output tokens. Codex does not set a smaller output budget on normal Responses API turns, so the catalogue must describe the safe input allowance rather than the raw total:
1,050,000 total - 128,000 maximum output = 922,000 safe input
Use the same safe input policy for all four direct-provider catalogue models:
Preserve the operator's selected supported model when configuring context. The examples below use Astra; enabling large context does not authorize replacing another selected model.
Codex applies its normal 95% effective-window reserve to the 922,000-token input allowance, so it reports and guards about 875,900 usable tokens. Set automatic compaction to 700,000 total active tokens. That leaves about 175,900 tokens inside Codex's effective guard and 222,000 tokens before the provider's safe input ceiling for the next prompt, tool schemas and results, instructions, serialization overhead, and compaction itself. This larger margin is intentional: Codex 0.144.6 checks already-recorded context before adding the next user message and context updates, and a terminal response that crosses the threshold may not compact until the following turn. The observed large-context workload grew by about 144,000 tokens in one turn, which made the former 820,000 threshold too aggressive.
Long-context requests above 272,000 input tokens use the provider's higher long-context pricing. Do not enable this route accidentally for workloads that do not benefit from it.
`~/.codex/models-api-1m.json` must contain these values for all four model slugs while preserving the rest of each model entry:
{
"context_window": 922000,
"max_context_window": 922000,
"auto_compact_token_limit": 700000
}Leave `effective_context_window_percent` absent to use Codex's 95% default, or set it explicitly to the integer `95`. Null, floating-point, or other values are invalid.
Start from the complete native catalogue for the installed Codex release, including its reviewer models: a custom catalogue replaces the built-in catalogue rather than overlaying selected entries. Use every model's genuine metadata and a client satisfying its minimum version. Preserve instructions, tool capabilities, and every safety field, including required review behavior; override only the context and compaction fields above. Never invent Astra metadata or relabel a Sol entry as Astra. The API context contract above supports the direct-provider override; it does not expand ChatGPT entitlement or relax client safety requirements. The preflight validates context and credential delivery, not the provenance of the remaining catalogue metadata.
The root section of `~/.codex/config.toml` needs:
model = "gpt-6-astra" model_provider = "openai_api_direct" model_context_window = 922000 model_auto_compact_token_limit = 700000 model_auto_compact_token_limit_scope = "total" model_catalog_json = "/Users/steipete/.codex/models-api-1m.json" [model_providers.openai_api_direct] name = "OpenAI API direct" base_url = "https://api.openai.com/v1" wire_api = "responses" requires_openai_auth = false [model_providers.openai_api_direct.auth] command = "/Users/steipete/.codex/bin/fetch-openai-inference-key.zsh" timeout_ms = 5000 refresh_interval_ms = 300000
Replace legacy values such as `model_context_window = 1050000` or `model_auto_compact_token_limit = 233000`; do not leave duplicate ro
Shared agent instructions, skills, and small portable helpers for Peter's local workspaces.
Repo: steipete/agent-scripts
Requested GitHub PR/issue agent transcripts: redact, trim, preview, and insert safely.
Control the user's existing signed-in Chrome: callable Codex Chrome plugin first, then…
ClawSweeper status: URLs, workflow health, active workers, ops snapshot.
ClickClack ops: chat app, Cloudflare Workers deploy, DNS/docs/app, container rollout.
Cloudflare Registrar: domain availability, prices, registration via mcporter.