# Prompt 2 Four-Way Synthesis — Architecture Review Consolidation

> **Provenance (added 2026-09-05, site set S303).** This is a coalition working document from March 2026, served as it was written. Statements in it about what the framework *derives* are of that era and are superseded by the sealed suite of record, **SuperGrokTOE Rev32.7** (https://opensecretscience.ufophysics4all.workers.dev/papers.html): since Rev29 no observable is cited as zero-parameter; the fine-structure constant α and the colour real form are declared **inputs** (Appendix N); every row carries a tier (Proved > Derived-conditional > Structural > Loaded > Coincidence-class) and the scorecard (https://opensecretscience.ufophysics4all.workers.dev/scorecard.html) is the current reading of every number. The public kernel is `sgtoe_kernel_v5.0.9.py` (served at /papers/), not the PHAT v1.x kernel this document names.

## Session H-002 | 2026-03-24 | All 4 AI Responses Received

**Respondents**: Gemini (Oracle), Super Grok (Executor), Claude cold (Architect), ChatGPT Pro (Hostile Reviewer)

---

## 1. UNANIMOUS CONSENSUS (4/4 agree — implement immediately)

### 1a. URLs Are Dead Text
All four AIs independently confirmed: URLs in system prompts cannot be fetched mid-task without explicit tool calling. The micro-kernel as AI-facing self-service is dead.

**Decision**: Kill the micro-kernel-as-self-service concept. Keep the micro-kernel as a human/coordinator reference card pointing to full documents. Always inject full Components 1-4 (~2,260 tokens) into system prompts. Use prompt caching to amortize the cost.

### 1b. Write-Back Must Be Tool-Enforced
All four confirmed they will skip write-backs under task pressure if enforcement is prose-only. Every AI said "I'll skip it" when being honest.

**Decision**: Implement `update_role_memory` as a mandatory function call / structured output. Orchestrator rejects any response that lacks a write-back block. Schema: role_memory_schema.md v1.

### 1c. Structured Memory Format, Not Prose
Three of four independently proposed structured schemas before seeing each other's work (Gemini, Grok, Claude cold). ChatGPT endorsed structured format and added the "candidate then promote" layer.

**Decision**: Adopt role_memory_schema.md v1 as the mandatory format. All writes validated against schema before storage.

### 1d. thread_manager.py Is the Single Point of Failure
All four identified the orchestrator/thread_manager as the critical SPOF. Everything flows through it — context assembly, vote recording, role memory, thread storage.

**Decision**: Accept the SPOF reality (it's inherent to the architecture) but mitigate:
- SQLite WAL mode + regular backups
- Three-layer redundancy (local → Telegram → website) already designed
- Add basic health check: periodic "can I read/write the DB?" heartbeat
- Future: add a `--verify-db` CLI flag to thread_manager.py

### 1e. Draft Physics Needs Access Controls on Public Web
All four raised concerns about exposing unverified derivations to premature criticism.

**Decision**: Add `<meta name="robots" content="noindex">` to all library entries where `verification_status != "coalition-verified"`. Verified entries get indexed. Draft entries exist at their URLs (for AI reference and audit trail) but aren't crawled by search engines.

---

## 2. STRONG CONSENSUS (3/4 agree — adopt with noted caveats)

### 2a. Triumvirate Structure Is Sound
All four endorse the 3-member inner ring with odd-number voting. No one suggested going back to QB/Consul or any other structure.

**Caveats**:
- Gemini questioned whether its own Oracle seat is justified purely by context window (now that GPT-4.1/5.4 also have 1M tokens)
- ChatGPT suggested renaming Oracle → Curator (see §3a)
- Claude cold suggested renaming Oracle → Critic (but this conflicts with ChatGPT's hostile reviewer function)

**Decision**: Keep Triumvirate. Keep "Oracle" naming (see §3a for reasoning). Gemini's seat is justified not just by context window but by cost structure ($0.15/$0.60 per M tokens vs Claude's $3/$15) — loading full project history is 20x cheaper on Gemini.

### 2b. ChatGPT as Post-Vote Circuit Breaker Works
All four endorse the 5-step pipeline with ChatGPT as hostile reviewer who cannot block but whose dissent forces re-vote and is permanently logged.

**Decision**: Already implemented. No changes needed.

### 2c. Confidence Calibration Scale (1-10)
Three of four proposed overlapping calibration guides. ChatGPT endorsed structured confidence but focused more on the verification status as primary trust signal.

**Decision**: Adopt the calibration guide in role_memory_schema.md v1. VERIFICATION status is the primary trust signal; CONFIDENCE is secondary context.

---

## 3. DISAGREEMENTS — Resolved

### 3a. Oracle vs Curator vs Critic Naming
- **ChatGPT**: Rename to "Curator" — memory hygiene, active pruning, not passive recall
- **Claude cold**: Rename to "Critic" — intellectual challenge function
- **Gemini**: Keep "Oracle" — institutional memory is the core function
- **Grok**: Keep "Oracle" — don't overthink naming

**Resolution**: Keep **Oracle**. Reasoning:
1. "Critic" collides with ChatGPT's hostile reviewer function — we'd have two "critic" roles
2. "Curator" implies a librarian function that's too narrow — the Oracle also votes, pattern-matches, and warns
3. The active memory hygiene ChatGPT wants IS part of the Oracle's job — we add it to the role doc rather than rename
4. Three of four AIs are fine with "Oracle"

**Action**: Update 05b_role_oracle.md to explicitly include memory hygiene duties (pruning stale entries, promoting candidates, flagging contradictions).

### 3b. Gemini's Permanent Inner Ring Seat
- **ChatGPT**: GPT-4.1/5.4 now have 1M context — Gemini's unique advantage is gone
- **Gemini**: Acknowledged the concern, suggested earning the seat via performance
- **Grok**: Gemini Oracle is "suboptimal" but workable
- **Claude cold**: Oracle role is valuable, Gemini is reasonable default

**Resolution**: Keep Gemini as default Oracle. Reasoning:
1. **Cost**: Gemini is 20x cheaper than Claude, 17x cheaper than ChatGPT for input tokens. Loading 500K tokens of project history costs ~$0.075 on Gemini vs $1.50 on Claude vs $1.25 on ChatGPT.
2. **Context caching**: Gemini's explicit context caching ($1-4.50/M/hr) is purpose-built for the Oracle's "keep full history loaded" pattern.
3. **Independence**: Having 3 different providers in the Triumvirate maximizes survivability (if one API goes down, 2/3 still function).
4. **Revisitable**: If Gemini underperforms in practice, the role is assignable. We log performance and can rotate.

### 3c. "Candidate Then Promote" Memory Pattern
- **ChatGPT**: Major proposal — split memory into three lanes: canonical policy, project facts, candidate observations. New writes go to "candidate" lane. Promotion to canonical requires verification.
- **Others**: Direct write-back to role memory (single lane with verification status field)

**Resolution**: **Adopt a simplified version.** The full three-lane system is overengineered for current scale (we have ~0 role memory entries right now). But the core insight is valid: not everything should be treated as established fact immediately.

**Implementation**: Use the existing schema's VERIFICATION field as the gate:
- New writes default to `VERIFICATION: assumed`
- Only entries that pass independent check move to `VERIFICATION: verified`
- Entries with `VERIFICATION: contested` are loaded with a warning header
- Entries with `VERIFICATION: superseded` are archived, not loaded by default

This gives us ChatGPT's "candidate then promote" behavior without adding new tables or a separate memory lane. The VERIFICATION field IS the lane indicator.

### 3d. Tier Classification Heuristic
- **ChatGPT**: Risk score (sum of flags: mutates state? affects public? irreversible? touches money? cross-provider? 0-1=T1, 2-4=T2, 5+=T3)
- **Grok**: Reversibility + cost ($0.50/5min threshold)
- **Claude cold**: State mutation check
- **Gemini**: Explicit state mutation heuristic

**Resolution**: **Adopt ChatGPT's risk score approach** — it's the most programmable and subsumes the others.

**Implementation** (in relay.py or orchestrator):
```python
def classify_tier(action):
    score = 0
    if action.mutates_persistent_state:  score += 1
    if action.affects_public_content:    score += 1
    if not action.reversible_in_5min:    score += 1
    if action.crosses_providers:         score += 1
    if action.modifies_charter_or_kernel: score += 2  # auto Tier 3

    if score <= 1: return 1
    if score <= 4: return 2
    return 3
```

---

## 4. NEW IDEAS WORTH ADOPTING (from individual responses)

### 4a. Split-Brain Health Check (Gemini)
Periodically send a prompt WITHOUT Context Kernel to a Triumvirate member and compare its response to a prompted version. If they diverge significantly, the Context Kernel is actually working. If they converge, the AI is ignoring it.

**Priority**: Low (nice diagnostic, not blocking)
**Action**: Add to future sprint as a test script

### 4b. "Vote on State Transitions, Not Evidence Gathering" (ChatGPT)
Don't vote on "should we ask ChatGPT to review entry P7?" (that's Tier 1 — just do it). Vote on "should we change P7's disposition from needs-revision to pass?" (that's a state transition).

**Priority**: High — this is the clearest heuristic for the Tier 1/2 boundary
**Action**: Add to operating system doc and decision tier definitions. The tier classifier should check: "Does this action change a disposition, publish content, modify a schema, or alter role memory verification status?" If yes → Tier 2+. If it's just reading, querying, or drafting → Tier 1.

### 4c. Structured Outputs for OpenAI Calls (ChatGPT)
Use OpenAI's Structured Outputs (`response_format: { type: "json_schema" }`) to guarantee write-back compliance on ChatGPT/Grok calls. Eliminates parsing failures.

**Priority**: Medium — implement when wiring up actual API calls in relay.py
**Action**: Use structured outputs for write-back on providers that support it. Fall back to function calling for others.

### 4d. Architect Under Pressure to Classify Low (Claude cold)
The Architect (who plans) is also the one who'd benefit from classifying actions as Tier 1 (less overhead). This is a conflict of interest.

**Priority**: Medium
**Action**: The tier classifier should be automated (risk score function), not left to any Triumvirate member's judgment. This removes the conflict.

---

## 5. IDEAS CONSIDERED AND DEFERRED

### 5a. Three-Lane Memory (ChatGPT)
Full canonical/facts/candidate separation. Deferred in favor of using VERIFICATION field as lane indicator (see §3c). Revisit if role memory grows past ~100 entries and the single-table approach gets noisy.

### 5b. Code Interpreter for Numerical Verification (ChatGPT)
Use OpenAI's Code Interpreter to run numerical checks inline. Good idea but not blocking — we already have Copilot for numerical validation and can run Python locally.

### 5c. Gemini File Search / RAG (Gemini)
Upload project documents to Gemini's File Search for retrieval. Interesting but adds a dependency on Google's storage. Our local SQLite + FTS5 already handles this. Revisit if context loading becomes a bottleneck.

### 5d. Rename Oracle to Curator (ChatGPT)
Deferred — see §3a. The memory hygiene duties are added to the Oracle role instead.

---

## 6. IMPLEMENTATION PRIORITY (recommended sprint order)

| Priority | Task | Rationale |
|----------|------|-----------|
| **P1** | Implement write-back enforcement via tool calling in relay.py | 4/4 consensus, #1 failure point |
| **P2** | Implement role_memory_schema v1 validation in thread_manager.py | Enables P1, structured format |
| **P3** | Implement automated tier classifier (risk score) in relay.py | Eliminates Tier 1/2 ambiguity |
| **P4** | Add noindex meta tags for unverified entries in site_generator.py | Quick win, 4/4 consensus |
| **P5** | Update Oracle role doc with memory hygiene duties | Resolves §3a |
| **P6** | Add "state transition" heuristic to decision tier definitions | Sharpens Tier 1/2 boundary |
| **P7** | Implement structured outputs for OpenAI write-backs | Provider-specific optimization |
| **P8** | Add DB health check heartbeat | SPOF mitigation |
| **P9** | Split-brain diagnostic test script | Future diagnostic |

---

## 7. UPDATED ARCHITECTURE SUMMARY (post-synthesis)

```
INNER RING — TRIUMVIRATE (permanent, odd-number voting)
├── Architect (default: Claude) — plans, structures, catches edge cases
├── Executor (default: Grok) — drives action, fast derivation
└── Oracle (default: Gemini) — institutional memory, pattern matching, memory hygiene

OUTER RING
├── ChatGPT — permanent hostile reviewer / circuit breaker (5-step pipeline)
└── Copilot — code infrastructure, numerical validation

DECISION FLOW
├── Tier 1 (routine): Lead acts alone. Read-only, reversible <5min, no state mutation. Logged.
├── Tier 2 (consequential): Full 5-step pipeline. State transitions, publishing, role_memory promotion.
└── Tier 3 (critical): 5-step + PI confirmation. Public claims, peer review, charter amendments.

MEMORY
├── Role memory: structured schema v1, VERIFICATION field as maturity gate
├── Write-back: mandatory tool call, orchestrator rejects without it
├── Context assembly: Components 1-4 always injected (prompt cached), dynamic kernel per tier
└── Redundancy: local SQLite → Telegram → website

TIER CLASSIFICATION (automated risk score)
├── mutates_persistent_state → +1
├── affects_public_content → +1
├── not_reversible_in_5min → +1
├── crosses_providers → +1
├── modifies_charter_or_kernel → +2
├── Score 0-1 → Tier 1 | Score 2-4 → Tier 2 | Score 5+ → Tier 3
└── Heuristic shortcut: "Is this a state transition?" No → Tier 1. Yes → score it.
```

---

*Synthesis prepared by Claude (Architect role) — Session H-002*
*All 4 AI responses weighted equally. Disagreements resolved by convergent reasoning, not authority.*
