Skip to content

docs: authorize bounded P1 v8 retry - #299

Merged
ictechgy merged 1 commit into
mainfrom
authorize/p1-v8-retry-388
Aug 11, 2026
Merged

docs: authorize bounded P1 v8 retry#299
ictechgy merged 1 commit into
mainfrom
authorize/p1-v8-retry-388

Conversation

@ictechgy

@ictechgy ictechgy commented Aug 11, 2026

Copy link
Copy Markdown
Owner

Summary

  • record the user-approved cumulative P1 ceiling of 388 identities
  • retain all 170 prior terminal identities and authorize one fresh root with at most 218 identities
  • freeze sonnet, $0.75/process, $163.50 arithmetic maximum, exact stop rules, and no old-root replay
  • require a new exact merge/candidate before provider execution

Boundaries

  • no npm publish, next, or latest
  • no P2-P6 provider activity before P1-F
  • any integrity, privacy, ambiguity, or spend failure closes the fresh root to provider work and yields P1-X

Verification

  • Stage2 feasibility/protected surfaces: 9 passed
  • python3 scripts/prepublish_check.py --skip-tests
  • Gate-B status ok
  • diff check clean

Summary by CodeRabbit

  • Documentation
    • Updated authorization records to reflect v8 approval, a cumulative identity ceiling of 388, and 218 remaining identities.
    • Added v8 requirements covering integrity, privacy, ambiguity, spending limits, and execution safeguards.
    • Documented that the previous candidate is no longer reusable and that a newly reviewed, verified candidate is required.
    • Updated release-readiness guidance for v8 execution and continued promotion blocking until fresh evidence is available.

@coderabbitai

coderabbitai Bot commented Aug 11, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

The authorization packet and roadmap now record v8 approval, a 388-identity cumulative ceiling, 218 new identities, fresh candidate requirements, provider-free verification, mandatory stop conditions, and continued promotion blocking.

Changes

V8 authorization and execution readiness

Layer / File(s) Summary
V8 authorization contract
research/p1-live-authorization-packet.md
The packet records v8 approval, one fresh root, a 388-identity cumulative ceiling, 218 remaining identities, and a required new reviewed merge, retained ref, and exact candidate.
V8 execution readiness
research/token-savings-roadmap.md
The roadmap updates release readiness and execution steps for candidate attestation, provider-free verification, v8 execution, mandatory P1-X stops, and continued promotion blocking until fresh P1-F evidence exists.

Estimated code review effort: 1 (Trivial) | ~5 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the documented authorization for a bounded P1 v8 retry.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch authorize/p1-v8-retry-388

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@research/p1-live-authorization-packet.md`:
- Around line 111-127: Before live prepare, require the candidate commit_sha to
exactly match the approved MERGED_SHA and validate that the retained ref
resolves to that reviewed commit, rather than relying only on hex syntax or the
manifest digest. Update the candidate validation flow around prepare and
--study-v2-candidate-hash, preserving existing manifest and provenance checks
while rejecting mismatched or unresolved refs before provider execution.
- Around line 117-120: Update the v8 stop-path handling so every terminal
failure, including canary, run/resume integrity or ambiguity, and
error_max_budget_usd spend failures, persists the canonical claim-disabled P1-X
before returning. Ensure spend failure is not treated as valid_task_failure_v1:
stop retries and block all subsequent provider calls, while preserving the
existing study-invalid-decision.json behavior where applicable.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 82e6ce9a-da46-46c1-be66-7a558a9b95f3

📥 Commits

Reviewing files that changed from the base of the PR and between d4b6302 and 408d06f.

📒 Files selected for processing (2)
  • research/p1-live-authorization-packet.md
  • research/token-savings-roadmap.md

Comment on lines +111 to +127
- On 2026-08-11 the user approved one fresh v8 root under the unchanged finite
study shape below, raising the cumulative consumed/reserved ceiling from 307
to exactly 388. The grant retains all 170 terminal identities and authorizes
at most 218 new identities: two discarded canaries, 108 initials, and up to
108 policy-valid retries. The model remains `sonnet`, every process remains
capped at `$0.75`, and the fresh-root arithmetic maximum remains `$163.50`.
Old-root reuse, replay, repair, migration, optional stopping, npm publication,
and active P2-P6 work before P1-F remain forbidden. Integrity, privacy,
ambiguity, or spend failure must immediately close the root to provider work
and produce the canonical claim-disabled P1-X.
- PR #298 merged the provider-free P2-P6 evaluator as `d4b6302...`; candidate
run `31457488674` passed for that exact merge and its downloaded manifest and
tarballs passed paired smoke and attestation verification. Because this
authorization amendment changes the selected source revision, that earlier
candidate is delivery evidence only and cannot be reused for v8. A new
reviewed merge, retained ref, and exact candidate are required before live
`prepare`.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

ast-grep outline context-guard-kit/benchmark_runner.py \
  --match '_benchmark_study_v2_invalid_canary_decision' \
  --view expanded

rg -n -C 8 \
  'MERGED_SHA|git ls-remote|manifest|checksums|SRI|provenance|clean install|prepare' \
  --glob '*.py' --glob '*.yml' --glob '*.yaml' .

Repository: ictechgy/context-guard

Length of output: 50379


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- benchmark study symbols ---'
rg -n \
  'def (_benchmark_study_v2_invalid_canary_decision|.*prepare.*|prepare)|study-v2-action|candidate|MERGED_SHA|git ls-remote|provenance|SRI|checksum|clean install' \
  context-guard-kit/benchmark_runner.py \
  context-guard-kit \
  scripts \
  tests \
  --glob '*.py' --glob '*.yml' --glob '*.yaml' \
  | head -n 500

printf '%s\n' '--- candidate-related files ---'
git ls-files | rg -i 'benchmark|study|candidate|manifest|provenance|attestation|install'

printf '%s\n' '--- relevant function outlines ---'
ast-grep outline context-guard-kit/benchmark_runner.py | rg -i 'study_v2|prepare|candidate|manifest|provenance|install|canary'

Repository: ictechgy/context-guard

Length of output: 50378


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- candidate verification and prepare ---'
sed -n '12040,12655p' context-guard-kit/benchmark_runner.py

printf '%s\n' '--- prepared-manifest revalidation and provider boundary ---'
sed -n '12760,12810p' context-guard-kit/benchmark_runner.py
sed -n '13745,13945p' context-guard-kit/benchmark_runner.py

printf '%s\n' '--- command dispatch ---'
sed -n '14260,14375p' context-guard-kit/benchmark_runner.py

printf '%s\n' '--- candidate fixture and prepare tests ---'
sed -n '930,1100p' scripts/rehearse_measurement_study.py
rg -n -C 6 \
  'prepare_benchmark_study_v2_executable|candidate binding drift|candidate checksum|SRI mismatch|offline npm install|provider' \
  tests/test_benchmark_study_v2.py

Repository: ictechgy/context-guard

Length of output: 50378


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- candidate manifest production ---'
rg -n -C 10 \
  'commit_sha|candidate_manifest|provenance|attestation|MERGED_SHA|git ls-remote|remote' \
  scripts/build_npm_candidates.py \
  tests/test_npm_candidates.py \
  research/p1-live-authorization-packet.md

printf '%s\n' '--- v2 prepare and candidate-binding tests ---'
rg -n \
  'def test_.*(prepare|candidate|manifest|checksum|integrity|provenance)|prepare_benchmark_study_v2_executable|candidate_manifest_sha256|commit_sha|candidate binding' \
  tests/test_benchmark_study_v2.py \
  tests/test_npm_candidates.py \
  | head -n 300

printf '%s\n' '--- full candidate revalidation tail ---'
sed -n '12700,12805p' context-guard-kit/benchmark_runner.py
sed -n '14335,14380p' context-guard-kit/benchmark_runner.py

Repository: ictechgy/context-guard

Length of output: 40937


Bind commit_sha to the reviewed MERGED_SHA before live prepare.

prepare validates the manifest, checksums, SRI, package contents, provenance fields, and offline clean install. It only checks that commit_sha has valid hex syntax. --study-v2-candidate-hash binds the manifest digest, not the source revision. No runtime check resolves the retained ref with git ls-remote. An old candidate can therefore pass when its manifest digest is supplied. Compare the candidate commit_sha with the approved MERGED_SHA and validate the retained ref before provider execution.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@research/p1-live-authorization-packet.md` around lines 111 - 127, Before live
prepare, require the candidate commit_sha to exactly match the approved
MERGED_SHA and validate that the retained ref resolves to that reviewed commit,
rather than relying only on hex syntax or the manifest digest. Update the
candidate validation flow around prepare and --study-v2-candidate-hash,
preserving existing manifest and provenance checks while rejecting mismatched or
unresolved refs before provider execution.

Comment on lines +117 to +120
Old-root reuse, replay, repair, migration, optional stopping, npm publication,
and active P2-P6 work before P1-F remain forbidden. Integrity, privacy,
ambiguity, or spend failure must immediately close the root to provider work
and produce the canonical claim-disabled P1-X.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

rg -n -C 8 \
  'P1-X|claim_allowed|stop_reason|integrity|privacy|ambigu|spend|reservation|cancel|drift|HOME|credential' \
  --glob '*.py' .

Repository: ictechgy/context-guard

Length of output: 50379


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- contract ---'
sed -n '100,135p' research/p1-live-authorization-packet.md

printf '%s\n' '--- v2 symbols and P1-X references ---'
rg -n \
  'BENCHMARK_STUDY_V2|P1-X|invalid_.*decision|decision|stop_reason|claim_allowed|raise (ValueError|RuntimeError|SystemExit)|spend|privacy|integrity|credential|reservation|cancel|drift' \
  context-guard-kit/benchmark_runner.py \
  --max-columns 220 \
  | head -n 1600

printf '%s\n' '--- function outline ---'
ast-grep outline context-guard-kit/benchmark_runner.py | sed -n '1,260p'

Repository: ictechgy/context-guard

Length of output: 50378


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- decision builder and persistence references ---'
rg -n \
  '_benchmark_study_v2_invalid_|invalid_decision|P1-X|decision_path|decision.json|write.*decision|json.*decision' \
  context-guard-kit/benchmark_runner.py \
  --max-columns 240

printf '%s\n' '--- v2 lifecycle sections ---'
sed -n '13280,14380p' context-guard-kit/benchmark_runner.py

Repository: ictechgy/context-guard

Length of output: 50378


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- v2 dispatch and decision writing ---'
sed -n '14370,14580p' context-guard-kit/benchmark_runner.py

printf '%s\n' '--- v2 spend and stop-related branches ---'
rg -n -C 5 \
  'study-v2|BENCHMARK_STUDY_V2|budget|cost_usd|max_budget|reservation|launch_reserved|blocked_study_invalid|infrastructure-invalid|integrity|privacy|credential|HOME|cancel' \
  context-guard-kit/benchmark_runner.py \
  --max-columns 240 \
  | awk 'NR <= 1800'

printf '%s\n' '--- v2 tests mentioning invalid decisions and stop behavior ---'
rg -n -C 8 \
  'invalid decision|invalid_decision|P1-X|blocked_study_invalid|ambiguous.*process|infrastructure-invalid|claim_allowed|spend|budget' \
  tests context-guard-kit --glob '*test*.py' \
  --max-columns 240 \
  | head -n 1600

Repository: ictechgy/context-guard

Length of output: 50379


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- exact v2 action handler ---'
rg -n -C 20 \
  'args\.study_v2_action|execute_benchmark_study_v2|analyze_benchmark_study_v2_executable|study-invalid-decision\.json|report_name' \
  context-guard-kit/benchmark_runner.py \
  --max-columns 240

printf '%s\n' '--- exact decision output branch ---'
sed -n '14380,14470p' context-guard-kit/benchmark_runner.py

Repository: ictechgy/context-guard

Length of output: 21537


🏁 Script executed:

#!/bin/bash
set -euo pipefail

python3 - <<'PY'
import ast
from pathlib import Path

path = Path("context-guard-kit/benchmark_runner.py")
tree = ast.parse(path.read_text(encoding="utf-8"))

wanted = {
    "_execute_benchmark_study_v2_canary_unlocked",
    "_execute_benchmark_study_v2_unlocked",
    "_benchmark_study_v2_run_slot",
    "_benchmark_study_v2_invalid_canary_decision",
    "_benchmark_study_v2_invalid_analytic_decision",
    "_analyze_benchmark_study_v2_executable_unlocked",
    "main",
}

def text(node):
    return ast.get_source_segment(path.read_text(encoding="utf-8"), node) or ""

for node in ast.walk(tree):
    if not isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)) or node.name not in wanted:
        continue
    calls = []
    raises = []
    literals = []
    for child in ast.walk(node):
        if isinstance(child, ast.Call):
            fn = child.func
            if isinstance(fn, ast.Name):
                calls.append(fn.id)
            elif isinstance(fn, ast.Attribute):
                calls.append(fn.attr)
        elif isinstance(child, ast.Raise):
            raises.append((child.lineno, text(child)))
        elif isinstance(child, ast.Constant) and isinstance(child.value, str):
            if child.value in {
                "P1-X", "error_max_budget_usd", "valid_task_failure_v1",
                "study_infra_invalid", "ambiguous provider process state permanently blocks this study root",
                "terminal infrastructure-invalid evidence permanently blocks later provider launches",
                "v2 {arm} host PreToolUse canary failed",
            }:
                literals.append((child.lineno, child.value))
    print(f"\nFUNCTION {node.name} line {node.lineno}")
    print("CALLS", sorted(set(calls)))
    print("RAISES", raises[:12])
    print("KEY_LITERALS", literals)

print("\nP1-X persistence sites:")
for node in ast.walk(tree):
    if isinstance(node, ast.Call) and isinstance(node.func, ast.Name) and node.func.id == "_study_write_private":
        source = text(node)
        if "report" in source or "decision" in source:
            print(node.lineno, source.replace("\n", " "))

print("\nBudget classification source hits:")
source = path.read_text(encoding="utf-8").splitlines()
for i, line in enumerate(source, 1):
    if "BENCHMARK_STUDY_V2_BOUNDED_FAILURE_RESULT_CODES" in line or "bounded_failure_usage" in line or "valid_task_failure_v1" in line:
        if 13600 <= i <= 13920:
            print(i, line.strip())
PY

Repository: ictechgy/context-guard

Length of output: 7387


🏁 Script executed:

#!/bin/bash
set -euo pipefail

python3 - <<'PY'
import ast
from pathlib import Path

path = Path("context-guard-kit/benchmark_runner.py")
source = path.read_text(encoding="utf-8")
tree = ast.parse(source)

wanted = {
    "_execute_benchmark_study_v2_canary_unlocked",
    "_execute_benchmark_study_v2_unlocked",
    "_benchmark_study_v2_run_slot",
    "_benchmark_study_v2_invalid_canary_decision",
    "_benchmark_study_v2_invalid_analytic_decision",
    "_analyze_benchmark_study_v2_executable_unlocked",
    "main",
}

def segment(node):
    return ast.get_source_segment(source, node) or ""

for node in ast.walk(tree):
    if not isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)):
        continue
    if node.name not in wanted:
        continue

    calls = []
    raises = []
    literals = []
    for child in ast.walk(node):
        if isinstance(child, ast.Call):
            if isinstance(child.func, ast.Name):
                calls.append(child.func.id)
            elif isinstance(child.func, ast.Attribute):
                calls.append(child.func.attr)
        elif isinstance(child, ast.Raise):
            raises.append((child.lineno, segment(child)))
        elif isinstance(child, ast.Constant) and isinstance(child.value, str):
            if child.value in {
                "P1-X",
                "error_max_budget_usd",
                "valid_task_failure_v1",
                "study_infra_invalid",
                "ambiguous provider process state permanently blocks this study root",
                "terminal infrastructure-invalid evidence permanently blocks later provider launches",
                "v2 {arm} host PreToolUse canary failed",
            }:
                literals.append((child.lineno, child.value))

    print(f"\nFUNCTION {node.name} line {node.lineno}")
    print("CALLS", sorted(set(calls)))
    print("RAISES", raises[:12])
    print("KEY_LITERALS", literals)

print("\nP1-X persistence sites:")
for node in ast.walk(tree):
    if isinstance(node, ast.Call) and isinstance(node.func, ast.Name):
        if node.func.id == "_study_write_private":
            text = segment(node).replace("\n", " ")
            if "report" in text or "decision" in text:
                print(node.lineno, text)

print("\nBudget classification source hits:")
for i, line in enumerate(source.splitlines(), 1):
    if (
        13600 <= i <= 13920
        and (
            "bounded_failure_usage" in line
            or "valid_task_failure_v1" in line
            or "BENCHMARK_STUDY_V2_BOUNDED_FAILURE_RESULT_CODES" in line
        )
    ):
        print(i, line.strip())
PY

Repository: ictechgy/context-guard

Length of output: 7387


Route every v8 stop path through persisted P1-X.

study-invalid-decision.json is written only by analyze. Canary failures and run/resume integrity or ambiguity failures raise without writing P1-X. error_max_budget_usd is classified as valid_task_failure_v1, so retries and later provider calls continue. Persist claim-disabled P1-X before returning from every stop path, and block further provider calls after spend failure.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@research/p1-live-authorization-packet.md` around lines 117 - 120, Update the
v8 stop-path handling so every terminal failure, including canary, run/resume
integrity or ambiguity, and error_max_budget_usd spend failures, persists the
canonical claim-disabled P1-X before returning. Ensure spend failure is not
treated as valid_task_failure_v1: stop retries and block all subsequent provider
calls, while preserving the existing study-invalid-decision.json behavior where
applicable.

@ictechgy
ictechgy merged commit 538cc56 into main Aug 11, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant