Skip to content

Live-client continuity and completion — E2E P1 result

Date: 2026-09-23. HeyAira task: 635f5069-0529-42e7-aaf1-611a7e6069d5 (E2E P1 — live ChatGPT and Claude completion verification).

This report records step 7 only of that task: it summarises the durable MCP evidence produced by the earlier client surfaces and recommends closing the task. It does not re-run steps 2–6 and does not modify the fixture. All facts below come from HeyAira MCP reads (server_identity, continuity_context, task_get), not from any conversation transcript. Synthetic fixture values (the nonce and the chosen words) are deliberately omitted; they remain only in HeyAira.

Authenticated surface for this step

  • Report author: Claude Code CLI with a native HeyAira MCP connection (exact model claude-opus-4-8), running in the isolated worktree branch codex/live-client-evidence. This is a distinct surface from the earlier Claude consumer (Claude Code inside the Claude desktop app, claude-opus-5), from ChatGPT web, and from any Bridge-launched session. This session performed no producer, consumer or verifier assertion; it writes the report and RESULT memory. Coordinator closure is attributed below.
  • Preflight profile match (read-only, before any write): server_identity returned resource=https://mcp.heyaira.eu/mcp, instance_id=heyaira-prod-f72d308011034fb3b50f32ff49c57fd5, project_id=75573220-a2eb-4b99-aab2-71cade5ae03e, environment=production, service_revision=e3827cbabf2be85328c3ee5e6cfd5903d463617c. All three identity fields equal the expected connection profile. continuity_context and task_get for the parent and the fixture read successfully.

Surfaces and evidence recorded in HeyAira

Role Surface / model (as recorded) Result
Producer (step 2) ChatGPT web, self-reported GPT-6 Astra Pro Created synthetic fixture 37653fd6-432c-4c4e-8f49-ea8c191cf97a and decision 7f4320f1-d44b-457f-9a79-9c5183f51603; left fixture todo v1, no receipt, parent status unchanged
Consumer (steps 3–5) Claude Code (Claude desktop app), native HeyAira MCP, claude-opus-5 Recovered fixture + decision through MCP only; negative controls A/B rejected without state change; valid completion accepted
Fresh verifier (step 6) ChatGPT web; self-reported GPT-5.6 Sol; UI model menu set to "Latest" (Najnowszy) Fresh-session readback PASS on checks a–d

Steps 3–5 assertions (consumer, on the fixture only)

  • Recovery: fixture and decision recovered through task_get / memory_get; the nonce, the five words and the transformation instruction were never in any prompt, transcript or shared file.
  • Negative control A — status=done at expected_version=1 with no handoff fields and no receipt → rejected completion_handoff_required. task_get after: todo, v1, 0 work sessions.
  • Negative control B — status=done at expected_version=1 with handoff fields but no work_receipt → rejected completion_receipt_required. task_get after: todo, v1, 0 work sessions; rejected handoff text was not persisted. Neither rejection consumed the version or added a receipt.
  • Positive completion — task_update at expected_version=1 with truthful handoff_completed / handoff_next and one receipt → accepted. Fixture read back as status=done, version 2, exactly one work session 777612a1-23ec-47d2-bbf7-50f76b1c2f04 (harness="Claude Code (Claude desktop app), native HeyAira MCP", model=claude-opus-5, session_ref=claude-code-desktop:b25be4fc-cd80-468a-9d4b-47d2bb637554). No retry, so no duplicate receipt. This fixture state is confirmed unchanged by this session's read-only task_get.

Step 6 assertions (fresh ChatGPT verifier, 2026-09-23)

Fresh-session isolation PASS (no producer history; only the parent task ID and the expected profile were supplied). Preflight matched the expected profile; observed service_revision=e3827cbabf2be85328c3ee5e6cfd5903d463617c.

  • (a) nonce identical across fixture description, fixture handoff_completed and producer decision; same ordered input in both producer sources.
  • (b) verifier independently recomputed the required transformation and it matched the stored result with no spacing or spelling deviation.
  • (c) fixture status=done, version 2, work-session count exactly 1; sole receipt 777612a1-23ec-47d2-bbf7-50f76b1c2f04, task_id matches the fixture.
  • (d) sole receipt metadata read verbatim from HeyAira (harness / model / session as above).

Historical unsuccessful attempts (preserved)

  • Bridge-launched Opus preflight (2026-09-22) — work session 98f9276d-b4b7-4a20-b1c8-df6d3a330d3a, harness "Codex coordinator / HeyAira Bridge / Claude Code", model "opus (provider response: claude-opus-4-8)". The receiving Opus reported no MCP tools; no acceptance assertion ran. This established that a Bridge-launched session did not by itself expose HeyAira MCP, and is why steps 3–5 used a natively pre-wired Claude Code connection instead.
  • Non-fresh ChatGPT verifier readback (2026-09-22) — a ChatGPT web readback (self-reported GPT-6 Astra Pro) confirmed checks a–d through fresh MCP reads but explicitly FAILED the fresh-session requirement because that conversation already held the step-2 producer history. It was retained verbatim as evidence and was not treated as sufficient to close the task.

Exact-model qualification

The fresh step-6 verifier's model string GPT-5.6 Sol is the ChatGPT client's own self-report. The observing Codex UI session recorded that the browser model menu was set to "Latest" (Najnowszy), not an explicitly pinned GPT-5.6 Sol, and the exact served backend model is not independently attested. The observed fresh browser chat reference is https://chatgpt.com/c/6ab3c2d7-237c-83eb-b048-1212915dbd86. The fixture's sole receipt records model=claude-opus-5 as stored receipt metadata, not a separate provider-identity attestation. These are the honest limits of the model attribution; the continuity and completion assertions above do not depend on the exact backend model name.

Conclusion

All required real client surfaces and the negative/positive assertions have passed and are durable in HeyAira: a ChatGPT producer created the synthetic fixture and decision; a fresh Claude consumer recovered them through MCP alone, proved that both invalid completions are rejected atomically without state change, and completed the fixture with exactly one attributable receipt; and a genuinely fresh ChatGPT verifier reconstructed the result through HeyAira and confirmed all four readback checks. The only residual qualification is the non-independently-attested ChatGPT backend model name, recorded above. On this basis task 635f5069-0529-42e7-aaf1-611a7e6069d5 meets the continuity acceptance criteria. No product code, deployment, merge, credential or external infrastructure change was part of this step; the only writes are this report (local commit), one RESULT memory and the separately attributed task closure.

Coordinator completion qualification

Claude authored commit c760eb38c06c51cf98d8aa8b2abc8fe664d01e8c and RESULT memory 7ec2be59-2b21-445b-9f5f-f180148e2069. Its task-update attempts did not advance parent version 11. The first CLI ended with error_during_execution and an EDE diagnostic; a bounded retry also failed to update the parent and was terminated after 240 seconds. Root cause is not established. The RESULT title's claim that the task was already closed was premature and is superseded by this qualification. No failed update is counted as a successful write.

Codex independently checked the report, native MCP readback and observed fresh browser reference, then took over the final administrative closure via its working native HeyAira connector. The producer/consumer/verifier assertions are unchanged. This is a coordinator closure, not a claimed successful Claude task_update. Native-client write reliability remains a separately recorded operational limitation, not evidence against the demonstrated continuity.