Live-client continuity and completion — E2E P1 result¶
Date: 2026-09-23. HeyAira task: 635f5069-0529-42e7-aaf1-611a7e6069d5
(E2E P1 — live ChatGPT and Claude completion verification).
This report records step 7 only of that task: it summarises the durable
MCP evidence produced by the earlier client surfaces and recommends closing the task. It
does not re-run steps 2–6 and does not modify the fixture. All facts below
come from HeyAira MCP reads (server_identity, continuity_context,
task_get), not from any conversation transcript. Synthetic fixture values
(the nonce and the chosen words) are deliberately omitted; they remain only in
HeyAira.
Authenticated surface for this step¶
- Report author: Claude Code CLI with a native HeyAira MCP connection
(exact model
claude-opus-4-8), running in the isolated worktree branchcodex/live-client-evidence. This is a distinct surface from the earlier Claude consumer (Claude Code inside the Claude desktop app,claude-opus-5), from ChatGPT web, and from any Bridge-launched session. This session performed no producer, consumer or verifier assertion; it writes the report and RESULT memory. Coordinator closure is attributed below. - Preflight profile match (read-only, before any write):
server_identityreturnedresource=https://mcp.heyaira.eu/mcp,instance_id=heyaira-prod-f72d308011034fb3b50f32ff49c57fd5,project_id=75573220-a2eb-4b99-aab2-71cade5ae03e,environment=production,service_revision=e3827cbabf2be85328c3ee5e6cfd5903d463617c. All three identity fields equal the expected connection profile.continuity_contextandtask_getfor the parent and the fixture read successfully.
Surfaces and evidence recorded in HeyAira¶
| Role | Surface / model (as recorded) | Result |
|---|---|---|
| Producer (step 2) | ChatGPT web, self-reported GPT-6 Astra Pro | Created synthetic fixture 37653fd6-432c-4c4e-8f49-ea8c191cf97a and decision 7f4320f1-d44b-457f-9a79-9c5183f51603; left fixture todo v1, no receipt, parent status unchanged |
| Consumer (steps 3–5) | Claude Code (Claude desktop app), native HeyAira MCP, claude-opus-5 |
Recovered fixture + decision through MCP only; negative controls A/B rejected without state change; valid completion accepted |
| Fresh verifier (step 6) | ChatGPT web; self-reported GPT-5.6 Sol; UI model menu set to "Latest" (Najnowszy) | Fresh-session readback PASS on checks a–d |
Steps 3–5 assertions (consumer, on the fixture only)¶
- Recovery: fixture and decision recovered through
task_get/memory_get; the nonce, the five words and the transformation instruction were never in any prompt, transcript or shared file. - Negative control A —
status=doneatexpected_version=1with no handoff fields and no receipt → rejectedcompletion_handoff_required.task_getafter:todo, v1, 0 work sessions. - Negative control B —
status=doneatexpected_version=1with handoff fields but nowork_receipt→ rejectedcompletion_receipt_required.task_getafter:todo, v1, 0 work sessions; rejected handoff text was not persisted. Neither rejection consumed the version or added a receipt. - Positive completion —
task_updateatexpected_version=1with truthfulhandoff_completed/handoff_nextand one receipt → accepted. Fixture read back asstatus=done, version 2, exactly one work session777612a1-23ec-47d2-bbf7-50f76b1c2f04(harness="Claude Code (Claude desktop app), native HeyAira MCP",model=claude-opus-5,session_ref=claude-code-desktop:b25be4fc-cd80-468a-9d4b-47d2bb637554). No retry, so no duplicate receipt. This fixture state is confirmed unchanged by this session's read-onlytask_get.
Step 6 assertions (fresh ChatGPT verifier, 2026-09-23)¶
Fresh-session isolation PASS (no producer history; only the parent task ID and
the expected profile were supplied). Preflight matched the expected profile;
observed service_revision=e3827cbabf2be85328c3ee5e6cfd5903d463617c.
- (a) nonce identical across fixture description, fixture
handoff_completedand producer decision; same ordered input in both producer sources. - (b) verifier independently recomputed the required transformation and it matched the stored result with no spacing or spelling deviation.
- (c) fixture
status=done, version 2, work-session count exactly 1; sole receipt777612a1-23ec-47d2-bbf7-50f76b1c2f04,task_idmatches the fixture. - (d) sole receipt metadata read verbatim from HeyAira (harness / model / session as above).
Historical unsuccessful attempts (preserved)¶
- Bridge-launched Opus preflight (2026-09-22) — work session
98f9276d-b4b7-4a20-b1c8-df6d3a330d3a, harness "Codex coordinator / HeyAira Bridge / Claude Code", model "opus (provider response: claude-opus-4-8)". The receiving Opus reported no MCP tools; no acceptance assertion ran. This established that a Bridge-launched session did not by itself expose HeyAira MCP, and is why steps 3–5 used a natively pre-wired Claude Code connection instead. - Non-fresh ChatGPT verifier readback (2026-09-22) — a ChatGPT web readback (self-reported GPT-6 Astra Pro) confirmed checks a–d through fresh MCP reads but explicitly FAILED the fresh-session requirement because that conversation already held the step-2 producer history. It was retained verbatim as evidence and was not treated as sufficient to close the task.
Exact-model qualification¶
The fresh step-6 verifier's model string GPT-5.6 Sol is the ChatGPT
client's own self-report. The observing Codex UI session recorded that the
browser model menu was set to "Latest" (Najnowszy), not an explicitly pinned
GPT-5.6 Sol, and the exact served backend model is not independently
attested. The observed fresh browser chat reference is
https://chatgpt.com/c/6ab3c2d7-237c-83eb-b048-1212915dbd86. The fixture's
sole receipt records model=claude-opus-5 as stored receipt metadata, not a
separate provider-identity attestation. These are the honest limits of the
model attribution; the continuity and completion assertions above do not
depend on the exact backend model name.
Conclusion¶
All required real client surfaces and the negative/positive assertions have
passed and are durable in HeyAira: a ChatGPT producer created the synthetic
fixture and decision; a fresh Claude consumer recovered them through MCP alone,
proved that both invalid completions are rejected atomically without state
change, and completed the fixture with exactly one attributable receipt; and a
genuinely fresh ChatGPT verifier reconstructed the result through HeyAira and
confirmed all four readback checks. The only residual qualification is the
non-independently-attested ChatGPT backend model name, recorded above. On this
basis task 635f5069-0529-42e7-aaf1-611a7e6069d5 meets the continuity acceptance criteria. No product
code, deployment, merge, credential or external infrastructure change was part
of this step; the only writes are this report (local commit), one RESULT memory
and the separately attributed task closure.
Coordinator completion qualification¶
Claude authored commit c760eb38c06c51cf98d8aa8b2abc8fe664d01e8c and RESULT
memory 7ec2be59-2b21-445b-9f5f-f180148e2069. Its task-update attempts did not
advance parent version 11. The first CLI ended with error_during_execution
and an EDE diagnostic; a bounded retry also failed to update the parent and was
terminated after 240 seconds. Root cause is not established. The RESULT title's
claim that the task was already closed was premature and is superseded by this
qualification. No failed update is counted as a successful write.
Codex independently checked the report, native MCP readback and observed fresh
browser reference, then took over the final administrative closure via its
working native HeyAira connector. The producer/consumer/verifier assertions are
unchanged. This is a coordinator closure, not a claimed successful Claude
task_update. Native-client write reliability remains a separately recorded
operational limitation, not evidence against the demonstrated continuity.