← Choice Health 2026-07-10-unified-call-capture-replay-plan.md raw .md

Unified Call Capture and Replay Plan

Status: active implementation contract. Recorded navigation, local call-clock replay, automatic aligned per-channel transcription, shared HTTP/CLI artifact loading, state-boundary snapshots, durable annotations, causal late-turn ownership, and repeated-visit comparison are implemented. Full transport-neutral finalization, model/tool span joins, carrier recording, flow behavior fixes, and isolated branch simulation remain NO-GO. Goal: every browser and phone call can be replayed end-to-end in the State Program Debugger: what the caller said, what the agent generated and emitted, what audio was recorded, what the FSM decided, and why a state repeated.

2026-07-10 checkpoint: FX-126 through FX-130 implement retained playback lifecycle hydration, recent-session discovery, selectable raw tracks, an explicit call-clock replay track for new local captures, and calibrated Play/Step/Reverse browser tests. Legacy tracks without a retained clock mapping remain typed unverified. Three fresh phone calls now prove that the raw evidence is sufficient to diagnose real repetition, while also proving that the the former transcript hydrator and CLI loader could not present that evidence on one trustworthy timeline. FX-132 through FX-136 now close those replay-side gaps: the canonical loader, aligned role-separated Whisper observations, durable problem markers, causal turn ownership, and semantic visit comparison all use the same debugger engine. The measurements and revised work packets below supersede earlier assumptions about which WAV is canonical.

Implementation Checkpoint

Packet Current result
CP0 Shared manifest-aware loader is used by HTTP and CLI; explicit legacy import exists. Full domain checksums/builder version remain open.
CP1 Browser and phone retain the evidence used here, but one idempotent transport-neutral finalizer and outbound frame ledger remain open.
CP2 Implemented: audio_replay.wav is split by role, Whisper runs asynchronously, and manifest-backed rows retain source/channel/model/clock provenance.
CP3 Bounds and channel separation are implemented; complete disagreement grouping and energy adjudication remain open.
CP4 One recorded cursor and strict visit frames are implemented. Late locked-choice, guardrail, transcription, and turn-complete events now retain causal and observed visits. Complete model/tool span links and barge-in decisions remain open.
CP5 Recorded Entry/Exit boundary rows self-verify; clears are removed from folds and exact differing slots are named. Unsupported writers remain typed gaps.
CP6 Implemented: synchronized browser replay, append-only problem markers, compare_visits, and rewind_to_first_divergence are shared by UI/API/CLI.
CP7-CP9 Property/LAW behavior fixes, isolated state capsules, and authorized carrier capture remain open.

Outcome

A completed call produces one retained session directory and one manifest-backed evidence bundle. The debugger opens the bundle without guessing. It can follow the call downward through distinct state visits, play available recorded audio, compare Gemini and Whisper transcripts, and identify the transition that caused a repeated visit.

This is a diagnostic contract, not a promise that every channel is always available. Missing evidence must be explicit, typed, and visible.

Empirical WAV and Transcript Study

Question and method

The recorder currently retains three legacy agent-bearing perspectives:

  1. audio.wav: legacy stereo composite, caller on the left and generated agent audio on the right;
  2. audio_agent.wav: generated/pre-codec agent audio;
  3. audio_agent_played.wav: PCM reconstructed from the outbound Telnyx codec payload.

The study transcribed all three with mlx-community/whisper-small.en-mlx. It also transcribed audio_caller.wav and audio_replay.wav, then split audio_replay.wav into left and right mono channels and transcribed those separately. The replay split is required as an alignment control because the replay manifest is the only one that declares sample_zero_call_ms: 0 on the session call clock.

The sample comprises three finalized calls from 2026-07-10:

Short ID Flow and caller goal Manifest duration
v3:W0bj... LAW, start a new claim 119.773 s
v3:SsLr... LAW, existing case/status 61.359 s
v3:B-ly... Property management, read dishwasher status 206.507 s

Whisper text is an observation, not truth. The comparison therefore used four signals: WAV duration, PCM/channel identity, 20 ms speech-energy alignment, and normalized phrase/three-word coverage. A phrase is treated as genuinely sent only when it appears on the outbound played/replay channel or is supported by a playback lifecycle event. Downmixed Whisper output alone cannot prove a repeat.

Duration and clock results

Call audio_replay.wav audio.wav / audio_agent.wav Difference from call audio_agent_played.wav
LAW new claim 119.773 s 143.338 s +23.565 s 118.681 s
LAW status 61.359 s 65.629 s +4.270 s 61.114 s
Property status 206.507 s 255.227 s +48.720 s 206.513 s

audio.wav is not independent evidence: its left channel is a byte-exact copy of audio_caller.wav and its right channel is a byte-exact copy of audio_agent.wav, padded to the longer raw track. It inherits both raw clock problems. In the property call it is 23.6% longer than the call, so no event, state, or transcript timestamp may be aligned against it.

By contrast, audio_agent_played.wav and the right channel of audio_replay.wav agreed on speech/non-speech activity in 99.7% to 100.0% of 20 ms windows across the three calls. Sample bytes are not globally identical because the raw writer concatenates timestamped chunks while the replay writer places them on a fixed buffer with per-chunk rounding and clipping. The activity agreement and matching content establish the replay right channel as the call-clock view of RIFF's outbound audio.

The study also exposed missing derivation evidence. After the LAW new-claim menu barge-in, the raw played file and replay-right channel retain equivalent speech activity but are not sample-identical because queue clear and absolute timeline rendering handle overlapping chunks differently. RIFF retains no per-frame outbound ledger explaining send/reservation time, playback ID, clear generation, collision, overwrite, or clipping. The replay manifest currently says only aligned; it does not retain measured drift, collision counts, clipped samples, or confidence. Those derivation facts are required before the debugger can explain why two locally reconstructed waveforms differ.

Transcript and repetition results

LAW new claim (v3:W0bj...). audio_agent.wav contains the complete menu, including options 3, 4, and leave-a-message. The outbound played track stops after option 2 because the caller barged in. The generated track is therefore a content superset, but it includes speech the caller was not sent. Its later summary begins near raw-media 116.1 s while the played version begins near call-clock 90.0 s, a 26.1 s divergence. Whole-mix Whisper reported Goodbye twice, while the played mono channel and replay-right mono channel contain it once. That is an ASR/mix false duplicate, not an audible repeat.

LAW existing status (v3:SsLr...). Generated and played agent transcripts had 100% mutual normalized three-word coverage. There was no material generated-versus-sent omission in this short call. The aligned replay adds the caller's interruptions and makes their overlap with the menu visible. The caller said Status around call-clock 30.5 s while the full menu was playing, but the ledger contains no audio_interrupted event. The caller repeated the choice after the menu and then routed successfully. The evidence needs a barge_in_candidate decision lane so a rejected or suppressed interruption is as explainable as an accepted one.

Property status (v3:B-ly...). The generated clock diverges progressively: the first dishwasher action prompt starts at generated-media 88.6 s versus played call-clock 72.2 s; a later full menu starts at 202.6 s versus 156.2 s. The same repeated phrases occur on generated, played, and replay-right audio:

Repeated outbound phrase/group Count on played/replay evidence
Dishwasher action question 3
Full property menu, including the initial menu 3
What else can I help with? recovery family 4

The repetition is therefore real outbound behavior, not merely transcript duplication. Playback lifecycle evidence independently reports two matching static prompts requested 445 ms apart near the end. The caller made roughly six status attempts; the durable live transcript recognized only two as status and rendered others as unrelated phrases. Recognition failures amplified the loop, but they did not create the repeated outbound audio.

The raw caller Whisper passes also produced long repeated hallucinations such as one word repeated over silence, and one segment with an end before its start. Phrase blacklists are insufficient. Transcript acceptance must be energy-backed, time-bounded, source-labelled, and disagreement-preserving. Both LAW manifests also have null flow_git_sha and null model versions, so an otherwise complete transcript still cannot reproduce the exact flow/model combination that produced it.

Superset decision

There is no single truth-superset WAV:

Debugging question Canonical evidence Why
What did RIFF generate or intend? audio_agent.wav, labelled raw/generated and clock-unverified Contains generated content, including audio later interrupted or suppressed
What outbound audio did RIFF submit? right channel of aligned audio_replay.wav; raw audit in audio_agent_played.wav Post-codec content on the session call clock
What caller audio did RIFF receive? left channel of aligned audio_replay.wav Inbound media placed on the same call clock
What is the best synchronized local conversation? audio_replay.wav, transcribed per channel and merged by timestamps Only retained local stereo artifact with an explicit sample-to-call mapping
What did the handset actually play? Telnyx dual-channel carrier recording Independent carrier evidence; not yet enabled

audio_agent.wav is the generated-content superset. audio_replay.wav is the synchronized evidence superset. audio.wav is neither: it packages caller and generated content but does so on incompatible raw clocks. The implementation must preserve all perspectives, default playback and timed transcription to audio_replay.wav, and never transcribe the stereo mix as one speaker stream.

Implementation Status and Remaining Defects

This table is the execution source of truth. Later work packets retain the full contract, but agents must not redo completed items or infer that a partial item is trustworthy beyond its stated boundary.

Item Status on 2026-07-10 Remaining work
Canonical CLI/HTTP loader Complete (FX-132) Add domain-level checksums and artifact revision; no new client-specific loader.
Aligned per-channel Whisper Complete locally (FX-132) Automatic next-call proof is pending; energy coverage and uncertain/rejected rows remain CP3.
HTTP media seeking Complete (FX-133) Browser uses the artifact URL directly so native Range requests drive seeking.
Playback lifecycle and duplicate finding Complete for current captures Add outbound frame, queue-clear, collision, overwrite, and clip evidence.
Transcript state ownership Partial Clock-overlap attribution exists; capture-state/decision-state and selected-input evidence remain.
State boundary variables Partial (FX-134) New boundaries retain self-verifying safe rows and name differing unledgered slots. Instrument or centrally observe the remaining direct mutations; old fingerprint-only calls stay incomplete.
Cursor-following UI Partial (FX-135) Audio, transcript, state strip, step/reverse, and duplicate replay follow one cursor; model/tool/guard lanes and first-divergence comparison remain.
Durable problem markers Complete point-marker slice (FX-135) Browser, HTTP, and CLI share append-only revisions and hashes. Add note/category editing and explicit range selection in the UI.
Causal LLM/tool timeline Open Join prompt, generation, suppression, first-token/audio, tools, guards, and transitions.
Safe branch simulation Open Implement immutable state capsules, overlays, alternate input/audio, and isolated forward comparison.
Carrier-heard evidence Human GO required Local replay proves sent/queued audio only.
Barge-in decision evidence Open Capture candidate, threshold, accepted/rejected/suppressed reason, and queue-clear generation.
Reproducibility revisions Open Freeze flow, engine, prompt, and model revisions; current study manifests contain nulls.

Required Evidence Contract

Every finalized call, regardless of browser or Telnyx transport, retains:

Evidence Required artifact Meaning when unavailable
Call identity session_manifest.json call cannot be opened
FSM ledger bus_events.jsonl EVENT_LEDGER_UNAVAILABLE
State trace state_trace.jsonl or ledger equivalent STATE_TRACE_UNAVAILABLE
Caller audio audio_caller.wav, when transport supplies it CALLER_AUDIO_UNAVAILABLE
Agent emitted audio audio_agent.wav or audio_agent_played.wav, when capture is possible AGENT_AUDIO_UNAVAILABLE
Composite recording audio.wav, with declared channel map COMPOSITE_AUDIO_UNAVAILABLE
Call-clock replay audio_replay.wav, derived without altering raw tracks ALIGNED_REPLAY_UNAVAILABLE
Gemini transcript durable transcript.jsonl rows GEMINI_TRANSCRIPT_UNAVAILABLE
Whisper transcript durable whisper_transcript.jsonl rows or queued result WHISPER_TRANSCRIPT_UNAVAILABLE / WHISPER_PENDING
Timing per-turn/state timing events including first agent emit AGENT_EMIT_DELAY_UNAVAILABLE
Telnyx carrier recording dual-channel WAV, both tracks, from answer TELNYX_CARRIER_RECORDING_UNAVAILABLE
RIFF phone-wire reconstruction inbound caller plus exact outbound codec PCM RIFF_WIRE_RECORDING_UNAVAILABLE

No raw secret, card, authentication code, or unnecessary caller PII may enter text, slot, event, or derived-debug artifacts. Transcript and slot retention use the existing redaction policy; safe fingerprints preserve change evidence where raw values cannot be retained. Spoken audio is inherently unredactable and is governed by the separate Audio Privacy and Retention Contract below.

Cross-Transport Playback Lifecycle

Per-prompt playback evidence is required before the carrier-recording lift. Both browser and Telnyx paths emit the same lifecycle schema:

Event Required identity and evidence
audio_requested session, visit, state, playback ID, source kind, safe content fingerprint, request epoch
audio_started identifiers, actual start epoch, transport, client/wire acknowledgement source
audio_completed identifiers, end epoch, played duration, completion reason
audio_interrupted identifiers, epoch, elapsed duration, interruption reason
audio_failed identifiers, epoch, typed failure and safe diagnostic

The browser creates a playback ID before dispatch and posts client-side acknowledgements from the actual Web Audio/player start, stop, completion, and error boundaries. The server persists them; browser console logs are not evidence. Telnyx emits corresponding acknowledgements at outbound codec/send boundaries and retains carrier recording lifecycle separately.

This lifecycle is the required cross-transport signal for detecting a double request or replay. Carrier audio is an independent phone-only confirmation that the duplicate was audible on the Telnyx call. A plan may pass browser duplicate detection without a carrier file, but it may not claim caller-heard phone audio without carrier confirmation.

Audio Privacy and Retention Contract

Full caller and agent audio can contain names, addresses, health information, authentication codes, and other spoken PII that cannot be redacted reliably. Therefore:

The plan defines the engineering controls; enabling recording for real calls is a separate human GO and policy decision.

These controls do not block implementation or testing with synthetic audio, redacted fixtures, or explicitly authorized local recordings. Privacy/retention hardening is a production enablement gate, not a prerequisite for building the playback lifecycle, debugger replay, detector, or simulation engine.

Canonical Session Layout

data/sessions/<session-id>/
  session_manifest.json
  bus_events.jsonl
  state_trace.jsonl
  transcript.jsonl
  whisper_transcript.jsonl
  audio.wav
  audio_replay.wav
  audio_caller.wav
  audio_agent.wav
  audio_agent_played.wav
  state_program_artifact.json
  replay_index.json

session_manifest.json lists each artifact, checksum, availability, producer, channel map, time base, and unavailability reason. replay_index.json is a convenience index only; the ledger and audio artifacts remain source evidence.

Telnyx Two-Sided Recording Contract

Phone replay requires two independent capture points:

  1. Carrier truth: Telnyx records from answer with track both, channels dual, format wav, and no silence trimming. RIFF retains the recording ID, call-control/session IDs, start/stop/saved webhook epochs, checksum, channel count, sample rate, duration, and retrieval status. The downloaded carrier WAV is the default audible phone-call replay because it is independent of RIFF's local mixer.
  2. RIFF wire reconstruction: RIFF retains inbound media after decoding and the exact outbound PCM reconstructed from the codec bytes immediately before ws_send. This produces audio_telnyx_wire.wav, with a declared caller and agent channel map, plus the two mono source files.

The current recorder already has relevant inputs: write_caller() receives decoded inbound Telnyx audio and write_agent_played() receives PCM rebuilt from the exact outbound Telnyx payload. However, the existing composite audio.wav combines caller audio with audio_agent.wav, which is the agent source before the final phone codec path. Implementation must add the explicit wire composite rather than silently treating the existing composite as carrier truth.

The manifest must preserve all three agent perspectives when available:

Artifact Question answered
audio_agent.wav What agent/model audio did RIFF produce before the phone codec?
audio_agent_played.wav What outbound codec PCM did RIFF submit to Telnyx?
Telnyx dual-channel WAV What did the carrier recording contain on the two call legs?

Telnyx documents dual recording as first leg on channel A and other legs on channel B; RIFF must not assume that means caller-left/agent-right. A canary fixture with non-overlapping caller and agent tones verifies the channel map, and the verified mapping is written into each manifest per direction: inbound_channel, outbound_channel, and mapping_method. A channel whose direction was not verified is labelled unknown, not guessed.

Recording lifecycle

One replay clock

The carrier WAV, local wire WAV, local replay WAV, transcript segments, playback lifecycle, and FSM events are normalized onto one clock_ms axis from call answer. Each audio artifact declares its media-zero offset and any measured alignment correction. A local audio_replay.wav may use retained caller media timestamps plus a serialized outbound-queue estimate, but its manifest must label that agent lane as sent/queued evidence rather than carrier-heard truth. The UI defaults to the carrier WAV while allowing an operator to switch to caller mono, RIFF outbound-wire mono, or pre-codec agent audio without moving the cursor. Play, pause, scrub, step, and reverse therefore keep both sides of the Telnyx call aligned with state and transcript evidence.

The named alignment fiducial is the first outbound playback whose exact PCM is present in both RIFF's wire track and the carrier outbound channel. RIFF aligns those signals by bounded cross-correlation, then validates the inbound direction against the first caller-energy segment. If no outbound playback exists, it uses the first inbound media segment. The manifest records alignment_method, offset_ms, confidence, and measured drift; low-confidence alignment is unavailable, not a guessed offset.

One Finalization Pipeline

Create a single finalize_call_evidence() contract used by browser and phone transports.

  1. Freeze the session ID, flow revision, engine revision, start/end epoch, and transport metadata.
  2. Flush the event bus and state trace for that session into its directory.
  3. Finalize recorder streams and record channel availability rather than creating misleading empty artifacts.
  4. For Telnyx calls, wait for and download the dual-channel carrier recording, or retain a typed pending/unavailable result after bounded retries.
  5. Persist Gemini caller and agent transcript events into transcript.jsonl with role, source, turn, state/visit when known, and timing.
  6. Run or enqueue Whisper for each available permitted audio channel. Store source, model, segments, timestamps, language, and unavailable/pending reason.
  7. Build state_program_artifact.json from the same frozen ledger.
  8. Write manifest checksums and a completeness summary.
  9. Run the capture-completeness gate before declaring finalization successful.

Browser and Telnyx adapters may differ in what they can capture, but they must invoke this same finalizer and express gaps with the same schema. They must never silently omit a ledger or transcript.

Debugger Replay Model

The State Program Debugger consumes only retained artifacts.

Repeat Diagnosis

A repeat is a property of visits, not state IDs. The debugger must show:

  1. menu#1 → cap_leave_message__collect#1 → menu_announce#2 → menu#2 as separate frames.
  2. The transition reason and guard/verdict for each edge.
  3. Caller/Gemini/Whisper observations before the return edge.
  4. A repeat summary: first visit, repeated visit, intervening state path, and earliest evidence divergence.
  5. A typed conclusion only when supported: caller retry, no-match/retry policy, timeout/no-input, tool failure, or REPEAT_CAUSE_UNAVAILABLE.

The implemented comparison uses semantic projections rather than raw event objects. It compares entry context, boundary values, acted-on input, source-labelled caller observations, slot writes, playback, and exit decision; it ignores volatile visit/event/playback IDs, sequence numbers, and timestamps. Entry context remains repeat topology, not a claimed behavioral defect. The engine chooses the earliest candidate-side difference by recorded timeline position. If an unavailable required lane occurs earlier, it returns an earliest_known_divergence but withholds first_divergence and refuses to move the synchronized cursor. When telemetry for a turn arrives after synchronous transitions have re-entered the same state, the event carries both the causal visit and the visit active at observation time. Causal ownership wins; otherwise the debugger can fabricate a repeat by assigning the earlier input to the later occurrence.

Work Packets

These packets are deliberately small enough for lower-capability implementation agents. A packet owns its listed files, lands its red test first, and ends with one focused commit. Agents must not combine an evidence-contract change with a caller-facing flow change. No packet requires a real OTA call unless its gate explicitly says HUMAN GO.

CP0 — Canonical Manifest Loader and Cross-Client Parity

Objective: browser, CLI, Codex, and Claude build byte-identical artifacts from one session bundle.

Primary files: new riff/postcall/session_evidence.py, riff/state_program_debugger_service.py, scripts/state_program_debugger.py, tests/test_state_program_debugger_service.py, and tests/test_state_program_debugger_cli.py.

Steps:

  1. Introduce SessionEvidenceBundle.open(session_dir) as the only canonical, manifest-mediated loader. It validates paths, checksums, artifact kinds, availability, channel maps, and clock maps.
  2. Move _audio_refs() into that loader. Stop canonical code from globbing WAVs and inventing bare recorded references.
  3. Make the CLI and HTTP service call the same loader and the same build_state_program_artifact() function. Put legacy filename fallback behind an explicit import_legacy_session() adapter.
  4. Add artifact_revision, builder version, and source checksums to the derived artifact. Redefine completeness by domains: manifest, ledger, state, variables, Gemini transcript, Whisper transcript, playback, audio, and clock.

Gate: a golden session produces byte-identical JSON, completeness domains, audio refs, and determinism hash through CLI, HTTP, and direct Python. The current case where CLI says AUDIO_CLOCK_UNAVAILABLE while HTTP says aligned must fail the red test and pass after the change.

CP1 — One Transport-Neutral Evidence Finalizer

Objective: browser and phone sessions freeze the same complete evidence bundle, in the same order, without relying on teardown races.

Primary files: new riff/postcall/finalize_call_evidence.py, riff/audio/session_recorder.py, riff/live/session.py, riff/phone/telnyx_transport.py, and focused browser/phone finalization tests.

Steps:

  1. Implement idempotent finalize_call_evidence() with one session-scoped bus spool. Do not slice a global fixed-size event buffer at teardown.
  2. Stop and flush audio, freeze the ledger and trace, persist Gemini transcript rows, register typed absences, and then write the final manifest atomically.
  3. Enqueue Whisper after raw evidence is frozen. Record pending, complete, failed, or unavailable; a background result updates the manifest through an atomic revision and rebuilds the derived artifact.
  4. Mark empty WAV headers unavailable. Retain manifest finalized: false until all synchronous evidence is durable.
  5. Resolve and freeze flow revision, engine revision, transcript model versions, and audio producer versions at call start. A missing revision is a named completeness defect, not null with no explanation.
  6. Retain agent_media_frames.jsonl with playback ID, PCM fingerprint, source and reserved call times, actual send time, queue/clear generation, collision or overwrite result, and clipped sample count.

Gate: browser and phone fixtures produce the same layout and completeness schema; double finalization is byte-stable; a crash at each phase leaves a recoverable partial manifest rather than a falsely complete call.

CP2 — Automatic Aligned Per-Channel Transcription

Objective: every permitted finalized call receives a timed transcript on the call clock without an operator running a script.

Primary files: new riff/postcall/transcript_hydration.py, a thin scripts/hydrate_transcript.py wrapper, the finalizer, the manifest schema, and tests/test_hydrate_transcript.py.

Canonical transcript row: stable ID, role, source (gemini, whisper, static_text, dtmf, or selected), text or typed redaction, source artifact/channel, model, language, media start/end, call-clock start/end, clock status, confidence/no-speech evidence, turn index, state/visit attribution, and source references.

Steps:

  1. For new local calls, split audio_replay.wav into caller-left and outbound-agent-right in process; do not downmix and do not require an external ffmpeg binary.
  2. Transcribe the channels independently, apply the retained clock map, and merge only after both roles have timestamps.
  3. Keep audio_agent.wav as a separate generated/intended observation. Its transcript may be useful for suppression comparison, but stays untimed when its clock is unverified.
  4. Persist whisper_transcript.jsonl, list it in the manifest with its model and source checksum, and make the artifact loader consume it.
  5. Preserve Gemini and Whisper rows side by side. Never overwrite one with the other and never treat Whisper as the selected FSM input unless the recorded resolver actually selected it.

Gate: a synthetic stereo fixture with non-overlapping phrases at known offsets yields correct roles and call times within 100 ms. A barge-in fixture proves generated menu options can exist without appearing on the played lane. No accepted transcript segment may extend beyond its source duration.

CP3 — Transcript Validation and Disagreement Adjudication

Objective: silence hallucinations and downmix artifacts are visible as rejected hypotheses, not promoted as spoken words or duplicate defects.

Primary files: riff/postcall/transcript_hydration.py, riff/postcall/state_program_artifact.py, and focused validation fixtures.

Steps:

  1. Reject invalid bounds (start < 0, end <= start, or beyond duration plus tolerance) and non-finite values.
  2. Require overlapping channel energy for accepted speech, retain no_speech_prob and model confidence, and classify low-evidence rows as rejected_silence or uncertain rather than deleting them silently.
  3. Group Gemini, Whisper, static text, DTMF, and selected-input observations by overlap. Emit word diff, source confidence, agreement status, and the source the FSM acted on.
  4. Detect repeated audio from playback IDs, content fingerprints, and outbound channel spans. Text-only repetition is corroboration, never sufficient proof.

Gate: the LAW whole-mix false Goodbye. Goodbye. does not create a repeat finding because the outbound mono/lifecycle contains one occurrence. The property action/menu repeats remain findings because they exist on played/replay evidence. Silence fixtures containing repeated hallucinated words do not enter the selected conversation.

CP4 — One Causal Timeline and Explicit State Ownership

Objective: audio, transcript, LLM, FSM, variables, and tools become one step-able event stream.

Primary files: riff/postcall/state_program_artifact.py, live-session telemetry emitters, riff/state_program_debugger.py, and artifact/engine tests.

Steps:

  1. Produce one timeline_events array ordered by (clock_ms, causal_suborder, stable_event_id). Untimed evidence goes in an explicit untimed lane and never defaults to call zero.
  2. Persist capture-state ID, decision-state ID, turn, and visit when speech is received. Deterministic hydration may fill a missing visit by interval, but must retain that the attribution was derived.
  3. Join static playback fingerprints to immutable resolved text. Retain model request/instruction hash, output transcript, first token, first audio, suppression reason, playback lifecycle, tool start/result, guard verdict, slot writes, and transition commit as linked spans.
  4. Make conversation observations cursor events rather than a side list.
  5. Use recorded state entry/exit clocks for visit bands. Display visit ordinal (menu#1, menu#2) separately from total occurrence count.
  6. Add barge_in_candidate, barge_in_accepted, barge_in_rejected, and barge_in_suppressed events with energy, threshold, active playback ID, state/visit, decision reason, and queue-clear generation. Render outbound frame overwrite/clipping evidence on the same span.

Gate: stepping and reversing across a caller utterance selects the same audio sample, transcript group, owning visit, resolver verdict, slot write, and transition in browser, CLI, and direct API. A post-transition transcription cannot leak into the prior announce frame.

CP5 — State Boundary Variable Integrity

Objective: Entry, Current, and Exit variables are trustworthy enough to support mutation and regression promotion.

Primary files: slot/boundary emitters, state_program_artifact.py, and state-program artifact tests.

Steps:

  1. Reproduce each VARIABLE_ENTRY_FOLD_MISMATCH and VARIABLE_EXIT_FOLD_MISMATCH from the three study calls with a minimal fixture.
  2. Align snapshot and fold fingerprint schemes; capture every supported writer, including seeding, nested/dict mutation, tool writes, sets, and clears.
  3. If a mutation cannot be instrumented, emit a typed gap naming the writer and affected variables. Never label an unobserved variable unchanged or ignored.

Gate: fixtures and one browser/phone integration call have zero unexplained fold mismatches. Overlay editing changes only simulated Current/Exit values and never recorded Entry/Exit evidence.

CP6 — Synchronized UI and Durable Problem Markers

Objective: a human and an AI can open the same point, hear it, and inspect the same causal evidence without raw-log work.

Primary files: web/state-program-debugger.html, web/state-program-debugger.js, debugger HTTP routes, and Playwright tests.

Steps:

  1. Render fixed lanes for state visits, caller Gemini/Whisper observations, selected input/DTMF, agent resolved text, generated audio, played audio, model spans, tools, slots, guards, and transitions.
  2. Make Play/Pause/Scrub/Step/Reverse drive one cursor and auto-focus the active row. Raw audition remains visibly separate and cannot move an aligned cursor as if it were calibrated.
  3. Add append-only durable annotations with point/range, session and artifact revision, event/visit refs, category, note, author, status, timestamps, and evidence window. The share URL carries annotation ID.
  4. Add rewind_to_first_divergence and compare_visits; highlight the return edge and first changed input/variable/playback rather than merely coloring every repeated state.

Implemented checkpoint: the visit rail exposes ordinals and the immediately previous same-state occurrence. Selecting a repeat renders one compact comparison band, and the jump command moves the canonical cursor to the candidate-side event. Equal or unavailable evidence fails closed with typed errors and no cursor motion. inspect_pause_gaps is a deterministic companion projection that classifies caller wait, agent response delay, and back-to-back output with zero model calls; per-call model-token usage remains a typed capture gap until the live transport persists it.

Gate: an operator marks the property status defect, opens the marker through CLI/API, replays a bounded audio range, and sees the same state and evidence hash. Desktop and mobile Playwright screenshots show no overlap or clipped labels.

CP7A — First Behavior Regression: Property Status Loop

Objective: use the completed debugger to fix the observed call, then promote the evidence into a cold-rebuild regression.

Primary files: riff/business_pack/compile.py, locked-choice retry ownership in the FSM engine, tests/business_pack/test_replicant_compile.py, and a recorded/synthetic replay fixture. Do not hand-edit generated flows/property_management.yaml as the source fix.

Required defects to prove independently:

  1. cap_existing_request__status entered and exited in roughly 4 ms on when: always, with no playback lifecycle between them, so the status summary was skipped.
  2. __locked_choice_retry_cap_existing_request__choose_action survived return to the capability and immediately satisfied choice_retry_exhausted.
  3. _replicant_selected_ticket_id survived menu return, so later select_ticket visits auto-advanced without a new selection.

Steps:

  1. Add a compiler/runtime gate: a caller-facing state with resolved speech may not take an immediate edge until an observable speech-completion boundary. For dynamic status text, generate an announce boundary whose success/failure edges both preserve explicit playback evidence.
  2. Make locked-choice retry counters state-visit scoped or clear them on every matched/exhausted exit.
  3. Clear selected-ticket ID/title/status and capability-local choice/action state when returning to the main menu.
  4. Recompile the property flow from the pack and run the cold-rebuild test.

Gate: one status request speaks the recorded status once, then offers the next operation once. Re-entering existing requests waits for a fresh ticket choice and a fresh action. The debugger reports no immediate retry-exhausted loop and no unexplained duplicate playback.

CP7B — Second Behavior Regression: LAW Barge-In and Case Status

Objective: prove why an early Status was ignored and prevent a capability question from being mistaken for a completed case-status message.

Evidence from v3:SsLr...: the caller says Status during the full menu, but no accepted interruption event exists; the repeated post-menu choice routes. After the caller asks whether cases can be listed, the agent states its limitation, repeats the office-message question, and immediately returns to the main menu.

Steps:

  1. Use CP4 barge-in decision telemetry to distinguish insufficient energy, disabled barge-in, echo rejection, and accepted interruption. Do not infer a reason from the absence of audio_interrupted.
  2. Separate case identifier, caller message/status question, and capability question in the LAW capability schema. A question such as can you list them? must not satisfy the message-complete guard by itself.
  3. After a limitation response, offer a concrete recovery (tell me the case name/number, leave a message, or return to menu) and wait for a new caller decision.
  4. Fix the source generator/compiler and run a cold rebuild; do not patch only the generated LAW YAML.

Gate: the first menu-time Status either interrupts and routes or appears as a rejected candidate with the exact reason. A case-list capability question does not auto-complete the collection or return to the menu without a caller choice.

CP8 — Immutable State Capsule and Branch Simulation

Objective: let Codex, Claude, and the UI change an input, variable, tool fixture, or candidate audio and replay forward without altering recorded truth.

Primary files: new capsule schema/runner, debugger engine/API, CLI, UI, and simulation tests. Existing scripts/replay_branch.py remains explicitly legacy until replaced.

Steps:

  1. Export riff.state_test_capsule/v1: exact state definition and legal edges, recorded engine/flow revisions, Entry snapshot, selected and alternate transcript observations, DTMF/control input, deterministic tool fixtures, audio outcome, and evidence refs.
  2. Run the capsule in a bounded subprocess with network disabled, side-effecting tools rejected, fixture-only reads, and explicit event/time budgets. Missing recorded code must fail; never fall back silently to HEAD.
  3. Implement fork, set_overlay, set_input, set_audio_fixture, run, compare_branches, and promote_regression over the same structured API used by browser and AI clients.
  4. Store simulated output separately with purple provenance, fork root, overlays, candidate revision, exit/variable/audio deltas, and invariant results. Recorded artifacts remain immutable.

Gate: mutate the property status state from the misheard input to status, run forward with a fixture ticket, hear candidate audio, compare recorded and candidate exits/variables, reverse to the fork, and export a deterministic test. UI, CLI, Codex, and Claude return the same branch hash.

CP9 — Carrier-Heard Confirmation (HUMAN GO)

Objective: distinguish RIFF sent/queued evidence from what the phone carrier recorded on both legs.

Primary files: Telnyx recording lifecycle, finalizer/download worker, manifest channel/alignment schema, and synthetic/canary tests.

Steps: start dual-channel recording from answer, retain recording webhooks, download without silence trimming, verify direction with non-overlapping tones, align to the outbound PCM fiducial, and record offset, drift, confidence, and channel map. Low-confidence alignment remains unavailable.

Gate: caller-only, agent-only, overlap, barge-in, interrupted playback, and early hangup are audible on the expected carrier channels. Enabling this on a real line still requires explicit human authorization; all prior packets run on local or synthetic evidence.

Handoff Contract for Implementation Agents

Every assigned packet must include:

Lower-cost models are appropriate for CP0 loader consolidation, schema fixtures, HTML rendering, and narrow UI/test packets. Use stronger models for CP1 teardown concurrency, CP3 transcript adjudication, CP7 FSM ownership, and CP8 sandboxed simulation. Parallel agents must use isolated worktrees and must not edit the same schema or generated flow concurrently.

Acceptance Gates

  1. A finalized call cannot have audio without explicit ledger, transcript, and clock availability results.
  2. CLI, HTTP, browser, Codex, and Claude consume the same manifest-backed bundle and return the same artifact bytes and determinism hash.
  3. audio_replay.wav duration equals manifest call duration within 20 ms; its channel map and clock provenance are retained. Raw tracks remain unverified unless separately calibrated.
  4. Automatic Whisper transcribes replay channels independently. No accepted row has invalid bounds, exceeds source duration, or lacks role, source, model, clock status, and evidence reference.
  5. Every speech-energy span longer than one second is covered by an accepted or uncertain transcript row, or by a typed reason. Silence hallucinations do not become selected conversation.
  6. Gemini, Whisper, static text, DTMF, and selected input remain distinct. Disagreement shows what the FSM actually used.
  7. Play, pause, seek, step, and reverse keep audio, transcript, state visit, variable projection, playback, model/tool span, guard, and transition on the same call-time cursor.
  8. Every state repeat renders as a distinct visit ordinal with the return edge, total occurrence count, and earliest causal divergence shown separately.
  9. Entry and Exit folds match recorded boundary snapshots, or name the exact uninstrumented writer and affected variables.
  10. Barge-in candidates record accepted/rejected/suppressed outcomes and queue clear generation. Replay derivation records collisions, overwrites, clips, drift, and confidence.
  11. A durable issue marker opened by browser, CLI, or AI resolves to the same bounded audio range, event/visit references, and artifact revision.
  12. A controlled double-play produces two matching lifecycle/audio spans with owning visits; a downmix-only duplicate transcript produces no false defect.
  13. The property status regression speaks status once, resets retry/selection state, and does not enter an immediate menu loop.
  14. The LAW status regression explains the early barge-in decision and does not auto-complete a capability question as an office message.
  15. A simulated branch is immutable, network-isolated, separately labelled, reproducible by hash, and cannot overwrite or masquerade as recorded audio or events.
  16. Every call freezes flow and model revisions or reports a typed reproducibility gap.
  17. Carrier-heard claims require the separately authorized Telnyx dual-channel artifact; local replay remains labelled sent/queued evidence.

Rollout

  1. Land CP0 first. No additional UI or AI client may introduce another session loader while artifact parity is unresolved.
  2. Land CP1 and CP2 behind the retained-call feature flag. Run browser and phone fixtures; keep asynchronous Whisper failures typed and non-fatal.
  3. Land CP3 through CP5 and make domain completeness visible. Keep the gate warning-only for one release while collecting real gap counts.
  4. Land CP6, including durable markers, after the canonical timeline is stable. A human must replay the property repeat without consulting raw logs.
  5. Use that marker to implement CP7A and CP7B, cold-rebuild both generated flows, and promote the two calls into regression fixtures.
  6. Land CP8 only after recorded evidence gates pass. The Fork button remains disabled until capsule validation and isolation tests are green.
  7. Implement CP9 with synthetic/canary evidence, but do not enable real carrier recording without explicit HUMAN GO.

This sequence makes the reported double-play observable and attributable. It does not itself fix the duplicate; behavior changes follow only after replay evidence identifies the responsible request, client playback, state transition, or model narration.

Operator Validation

After CP0, run one browser simulator call and one phone-path fixture. For each:

.venv/bin/python scripts/state_program_debugger.py build data/sessions/<session-id> --out /tmp/<session-id>.json
.venv/bin/python scripts/state_program_debugger.py command /tmp/<session-id>.json --json '{"action":"list_visits"}'

Open /state-program-debugger.html, enter the session ID, and verify the same visit list, completeness status, and repeat markers. A call without required evidence is a capture failure to fix before using it to judge flow behavior.

To persist the exact defect point for browser, Codex, and Claude:

.venv/bin/python scripts/state_program_debugger.py session-command \
  data/sessions/<session-id> \
  --json '{"action":"record_issue","issue":{"issue_id":"issue:repeat-1","status":"confirmed","time_ms":123400,"event_id":"bus:42","visit_id":"menu#3","reporter":"codex","category":"duplicate_agent_playback","note":"Same outbound prompt twice."}}'

This appends an annotation; it never rewrites recorded audio, events, or state.

Then run automatic transcript hydration and verify that the browser timeline, inspect_timeline, and the CLI expose identical caller/agent transcript groups, audio refs, visit ownership, and artifact hash. The study session v3:B-ly... is the required negative/positive control: its real outbound repeats must remain visible, while raw-track timestamps past 206.507 s and downmix-only duplicate words must not enter the synchronized lane.