Skip to content

workflow_get frozen task v16 result

  • Repository: h8v6/heyaira
  • Task: ee1b6fd5-078b-442f-9daa-4a3a857799ac, exact version 16 read through HeyAira MCP before changing files.
  • Execution: 69d8f410-23ce-4480-a913-cfbee25b4004
  • Branch: heyaira/ee1b6fd5
  • Starting candidate: fb23b53538d7c4545106d0b1655e47d03eb71f09
  • Original task base: 790cc1a31f7cf5b0e2db11c969fd89ad5996050a
  • Result: the commit containing this report, added on top of the starting candidate. Existing commits were not amended, rebased or reset.

The starting tree was clean. The carried candidate already implements the required response projection and documentation. This continuation adds targeted regressions and fresh verification evidence without changing unrelated runtime code. The previous report describes the earlier execution; this report records the current frozen v16 criteria and the current checks.

Behavior and regression coverage

Only the MCP workflow_get response is compact by default. The projection in src/heyaira/server.py replaces contract.profile.instructions, contract.verifier.instructions and contract.prompt with exact omitted, chars and UTF-8 sha256 markers. All other fields remain unchanged. detail=true returns the original full read result.

tests/test_workflow_view.py compares every field after restoring only the three omitted texts. It checks Unicode, empty text, input immutability and optional/missing fields. New cases verify that changing any of the three texts without changing its character count changes its hash. The registered MCP test now also checks compact/default/explicit-false and full output for a workflow with a null verifier, retaining that null and every other field.

tests/test_workflow_e2e.py contains the database-backed regression using a real MCP server and Bridge HTTP. It now runs for both queued and completed workflows, so comparison covers populated verifier verdicts and both execution attempts. For each phase it compares compact and full output field by field with default and explicit pagination, compares Bridge and internal reads against full MCP output, and verifies that stored contract text and its digest remain unchanged.

Internal caller audit

All production callers were read:

  • MCP workflow_get calls workflow.read and then the response-only projection.
  • workflow.read delegates directly to workflow.read_project without shaping.
  • Bridge /bridge/v1/tasks/{task_id}/workflow calls workflow.read_project directly and returns its full result.

src/heyaira/workflow.py and the Bridge route are unchanged from the original task base. No other production call site uses the projection.

Executed checks

No dependencies were installed. Temporary files and compiler module caches were placed under ignored .venv/task-tmp in the authorized workspace.

Targeted command (with workspace TMPDIR and the installed Git binary on PATH):

PYTHONPATH=src python -m pytest -q -p no:cacheprovider \
  tests/test_workflow_view.py tests/test_workflow_e2e.py \
  tests/test_mcp_runtime_contract.py

Actual result: 24 passed, 40 skipped in 1.95s, exit 0.

Documentation checks, both exit 0:

  • python scripts/export_mcp_contract.py --check: MCP contract is up to date.
  • python scripts/check_docs_contract.py: documentation contract passed: 30 MCP tools covered.

The first required suite invocation selected the Command Line Tools compiler and failed an unrelated Swift test because it could not load the macOS standard library: 1 failed, 395 passed, 233 skipped, 2 deselected in 16.42s. No code was changed to address that environment problem. Configuring the installed Xcode SDK and writable compiler caches made the isolated Swift test pass: 1 passed in 7.95s. The full required suite was then rerun as follows:

TMPDIR="$PWD/.venv/task-tmp" \
SDKROOT=/Applications/Xcode.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX.sdk \
SWIFT_MODULECACHE_PATH="$PWD/.venv/task-tmp/swift-module-cache" \
CLANG_MODULE_CACHE_PATH="$PWD/.venv/task-tmp/clang-module-cache" \
PATH="/Applications/Xcode.app/Contents/Developer/Toolchains/XcodeDefault.xctoolchain/usr/bin:/Applications/Xcode.app/Contents/Developer/usr/bin:$PATH" \
PYTHONPATH=src python -m pytest -q -p no:cacheprovider \
  --deselect tests/test_server_beat_audit.py::test_R6_02_tcp_connect_gets_only_the_budget_left_after_resolution \
  --deselect tests/test_server_beat_audit.py::test_F3_smtp_tarpit_is_cut_at_the_deadline

Actual final result: 396 passed, 233 skipped, 2 deselected in 16.30s, exit 0. Only the two tests explicitly excluded by task v16 were deselected. git diff --check also passed, exit 0.

Criteria and verification limits

Criterion Result Evidence
compact-default pass Shaping and registered MCP tests compare every field and exact markers; real MCP/PostgreSQL regression is committed and collected.
detail-full pass Unit and registered MCP tests assert full-output equality, including after compact calls and with a null verifier.
internal-callers pass Caller audit above; full read functions and Bridge route are unchanged.
contract-docs pass Tool description, generated contract and reference docs describe detail; both checks pass.
suite pass Required sandbox command completes with 396 passed, 233 skipped, 2 authorized deselections and no failures.

HEYAIRA_EXECUTION_E2E_DATABASE_URL is absent locally, so real database-backed regressions were skipped, not executed. CI's existing .github/workflows/ci.yml calls scripts/test_postgres.py, which supplies isolated database URLs, runs the complete suite and rejects skipped cases. Database behavior and independent acceptance remain for that verification; they are not claimed as locally tested.

Files changed in this continuation: tests/test_workflow_view.py, tests/test_workflow_e2e.py, and this report. Confidence is high in the MCP projection from passing field-by-field tests and the unchanged internal read paths. RESULT and receipt reference the final local commit through task_result. No push, PR, merge, installation, deployment or production mutation was performed.