Unified Call Capture and Replay Plan
Status: active implementation contract. Recorded navigation, local call-clock replay, automatic aligned per-channel transcription, shared HTTP/CLI artifact loading, state-boundary snapshots, durable annotations, causal late-turn ownership, and repeated-visit comparison are implemented. Full transport-neutral finalization, model/tool span joins, carrier recording, flow behavior fixes, and isolated branch simulation remain NO-GO. Goal: every browser and phone call can be replayed end-to-end in the State Program Debugger: what the caller said, what the agent generated and emitted, what audio was recorded, what the FSM decided, and why a state repeated.
2026-07-10 checkpoint: FX-126 through FX-130 implement retained playback
lifecycle hydration, recent-session discovery, selectable raw tracks, an
explicit call-clock replay track for new local captures, and calibrated
Play/Step/Reverse browser tests. Legacy tracks without a retained clock mapping
remain typed unverified. Three fresh phone calls now prove that the raw
evidence is sufficient to diagnose real repetition, while also proving that the
the former transcript hydrator and CLI loader could not present that evidence
on one trustworthy timeline. FX-132 through FX-136 now close those replay-side
gaps: the canonical loader, aligned role-separated Whisper observations,
durable problem markers, causal turn ownership, and semantic visit comparison
all use the same debugger engine. The measurements and revised work packets
below supersede earlier assumptions about which WAV is canonical.
Implementation Checkpoint
| Packet | Current result |
|---|---|
| CP0 | Shared manifest-aware loader is used by HTTP and CLI; explicit legacy import exists. Full domain checksums/builder version remain open. |
| CP1 | Browser and phone retain the evidence used here, but one idempotent transport-neutral finalizer and outbound frame ledger remain open. |
| CP2 | Implemented: audio_replay.wav is split by role, Whisper runs asynchronously, and manifest-backed rows retain source/channel/model/clock provenance. |
| CP3 | Bounds and channel separation are implemented; complete disagreement grouping and energy adjudication remain open. |
| CP4 | One recorded cursor and strict visit frames are implemented. Late locked-choice, guardrail, transcription, and turn-complete events now retain causal and observed visits. Complete model/tool span links and barge-in decisions remain open. |
| CP5 | Recorded Entry/Exit boundary rows self-verify; clears are removed from folds and exact differing slots are named. Unsupported writers remain typed gaps. |
| CP6 | Implemented: synchronized browser replay, append-only problem markers, compare_visits, and rewind_to_first_divergence are shared by UI/API/CLI. |
| CP7-CP9 | Property/LAW behavior fixes, isolated state capsules, and authorized carrier capture remain open. |
Outcome
A completed call produces one retained session directory and one manifest-backed evidence bundle. The debugger opens the bundle without guessing. It can follow the call downward through distinct state visits, play available recorded audio, compare Gemini and Whisper transcripts, and identify the transition that caused a repeated visit.
This is a diagnostic contract, not a promise that every channel is always available. Missing evidence must be explicit, typed, and visible.
Empirical WAV and Transcript Study
Question and method
The recorder currently retains three legacy agent-bearing perspectives:
audio.wav: legacy stereo composite, caller on the left and generated agent audio on the right;audio_agent.wav: generated/pre-codec agent audio;audio_agent_played.wav: PCM reconstructed from the outbound Telnyx codec payload.
The study transcribed all three with
mlx-community/whisper-small.en-mlx. It also transcribed
audio_caller.wav and audio_replay.wav, then split audio_replay.wav into
left and right mono channels and transcribed those separately. The replay split
is required as an alignment control because the replay manifest is the only
one that declares sample_zero_call_ms: 0 on the session call clock.
The sample comprises three finalized calls from 2026-07-10:
| Short ID | Flow and caller goal | Manifest duration |
|---|---|---|
v3:W0bj... |
LAW, start a new claim | 119.773 s |
v3:SsLr... |
LAW, existing case/status | 61.359 s |
v3:B-ly... |
Property management, read dishwasher status | 206.507 s |
Whisper text is an observation, not truth. The comparison therefore used four signals: WAV duration, PCM/channel identity, 20 ms speech-energy alignment, and normalized phrase/three-word coverage. A phrase is treated as genuinely sent only when it appears on the outbound played/replay channel or is supported by a playback lifecycle event. Downmixed Whisper output alone cannot prove a repeat.
Duration and clock results
| Call | audio_replay.wav |
audio.wav / audio_agent.wav |
Difference from call | audio_agent_played.wav |
|---|---|---|---|---|
| LAW new claim | 119.773 s | 143.338 s | +23.565 s | 118.681 s |
| LAW status | 61.359 s | 65.629 s | +4.270 s | 61.114 s |
| Property status | 206.507 s | 255.227 s | +48.720 s | 206.513 s |
audio.wav is not independent evidence: its left channel is a byte-exact copy
of audio_caller.wav and its right channel is a byte-exact copy of
audio_agent.wav, padded to the longer raw track. It inherits both raw clock
problems. In the property call it is 23.6% longer than the call, so no event,
state, or transcript timestamp may be aligned against it.
By contrast, audio_agent_played.wav and the right channel of
audio_replay.wav agreed on speech/non-speech activity in 99.7% to 100.0% of
20 ms windows across the three calls. Sample bytes are not globally identical
because the raw writer concatenates timestamped chunks while the replay writer
places them on a fixed buffer with per-chunk rounding and clipping. The activity
agreement and matching content establish the replay right channel as the
call-clock view of RIFF's outbound audio.
The study also exposed missing derivation evidence. After the LAW new-claim
menu barge-in, the raw played file and replay-right channel retain equivalent
speech activity but are not sample-identical because queue clear and absolute
timeline rendering handle overlapping chunks differently. RIFF retains no
per-frame outbound ledger explaining send/reservation time, playback ID, clear
generation, collision, overwrite, or clipping. The replay manifest currently
says only aligned; it does not retain measured drift, collision counts,
clipped samples, or confidence. Those derivation facts are required before the
debugger can explain why two locally reconstructed waveforms differ.
Transcript and repetition results
LAW new claim (v3:W0bj...). audio_agent.wav contains the complete menu,
including options 3, 4, and leave-a-message. The outbound played track stops
after option 2 because the caller barged in. The generated track is therefore a
content superset, but it includes speech the caller was not sent. Its later
summary begins near raw-media 116.1 s while the played version begins near
call-clock 90.0 s, a 26.1 s divergence. Whole-mix Whisper reported Goodbye
twice, while the played mono channel and replay-right mono channel contain it
once. That is an ASR/mix false duplicate, not an audible repeat.
LAW existing status (v3:SsLr...). Generated and played agent transcripts
had 100% mutual normalized three-word coverage. There was no material
generated-versus-sent omission in this short call. The aligned replay adds the
caller's interruptions and makes their overlap with the menu visible. The
caller said Status around call-clock 30.5 s while the full menu was playing,
but the ledger contains no audio_interrupted event. The caller repeated the
choice after the menu and then routed successfully. The evidence needs a
barge_in_candidate decision lane so a rejected or suppressed interruption is
as explainable as an accepted one.
Property status (v3:B-ly...). The generated clock diverges progressively:
the first dishwasher action prompt starts at generated-media 88.6 s versus
played call-clock 72.2 s; a later full menu starts at 202.6 s versus 156.2 s.
The same repeated phrases occur on generated, played, and replay-right audio:
| Repeated outbound phrase/group | Count on played/replay evidence |
|---|---|
| Dishwasher action question | 3 |
| Full property menu, including the initial menu | 3 |
What else can I help with? recovery family |
4 |
The repetition is therefore real outbound behavior, not merely transcript
duplication. Playback lifecycle evidence independently reports two matching
static prompts requested 445 ms apart near the end. The caller made roughly six
status attempts; the durable live transcript recognized only two as status
and rendered others as unrelated phrases. Recognition failures amplified the
loop, but they did not create the repeated outbound audio.
The raw caller Whisper passes also produced long repeated hallucinations such
as one word repeated over silence, and one segment with an end before its
start. Phrase blacklists are insufficient. Transcript acceptance must be
energy-backed, time-bounded, source-labelled, and disagreement-preserving.
Both LAW manifests also have null flow_git_sha and null model versions, so an
otherwise complete transcript still cannot reproduce the exact flow/model
combination that produced it.
Superset decision
There is no single truth-superset WAV:
| Debugging question | Canonical evidence | Why |
|---|---|---|
| What did RIFF generate or intend? | audio_agent.wav, labelled raw/generated and clock-unverified |
Contains generated content, including audio later interrupted or suppressed |
| What outbound audio did RIFF submit? | right channel of aligned audio_replay.wav; raw audit in audio_agent_played.wav |
Post-codec content on the session call clock |
| What caller audio did RIFF receive? | left channel of aligned audio_replay.wav |
Inbound media placed on the same call clock |
| What is the best synchronized local conversation? | audio_replay.wav, transcribed per channel and merged by timestamps |
Only retained local stereo artifact with an explicit sample-to-call mapping |
| What did the handset actually play? | Telnyx dual-channel carrier recording | Independent carrier evidence; not yet enabled |
audio_agent.wav is the generated-content superset. audio_replay.wav is the
synchronized evidence superset. audio.wav is neither: it packages caller and
generated content but does so on incompatible raw clocks. The implementation
must preserve all perspectives, default playback and timed transcription to
audio_replay.wav, and never transcribe the stereo mix as one speaker stream.
Implementation Status and Remaining Defects
This table is the execution source of truth. Later work packets retain the full contract, but agents must not redo completed items or infer that a partial item is trustworthy beyond its stated boundary.
| Item | Status on 2026-07-10 | Remaining work |
|---|---|---|
| Canonical CLI/HTTP loader | Complete (FX-132) | Add domain-level checksums and artifact revision; no new client-specific loader. |
| Aligned per-channel Whisper | Complete locally (FX-132) | Automatic next-call proof is pending; energy coverage and uncertain/rejected rows remain CP3. |
| HTTP media seeking | Complete (FX-133) | Browser uses the artifact URL directly so native Range requests drive seeking. |
| Playback lifecycle and duplicate finding | Complete for current captures | Add outbound frame, queue-clear, collision, overwrite, and clip evidence. |
| Transcript state ownership | Partial | Clock-overlap attribution exists; capture-state/decision-state and selected-input evidence remain. |
| State boundary variables | Partial (FX-134) | New boundaries retain self-verifying safe rows and name differing unledgered slots. Instrument or centrally observe the remaining direct mutations; old fingerprint-only calls stay incomplete. |
| Cursor-following UI | Partial (FX-135) | Audio, transcript, state strip, step/reverse, and duplicate replay follow one cursor; model/tool/guard lanes and first-divergence comparison remain. |
| Durable problem markers | Complete point-marker slice (FX-135) | Browser, HTTP, and CLI share append-only revisions and hashes. Add note/category editing and explicit range selection in the UI. |
| Causal LLM/tool timeline | Open | Join prompt, generation, suppression, first-token/audio, tools, guards, and transitions. |
| Safe branch simulation | Open | Implement immutable state capsules, overlays, alternate input/audio, and isolated forward comparison. |
| Carrier-heard evidence | Human GO required | Local replay proves sent/queued audio only. |
| Barge-in decision evidence | Open | Capture candidate, threshold, accepted/rejected/suppressed reason, and queue-clear generation. |
| Reproducibility revisions | Open | Freeze flow, engine, prompt, and model revisions; current study manifests contain nulls. |
Required Evidence Contract
Every finalized call, regardless of browser or Telnyx transport, retains:
| Evidence | Required artifact | Meaning when unavailable |
|---|---|---|
| Call identity | session_manifest.json |
call cannot be opened |
| FSM ledger | bus_events.jsonl |
EVENT_LEDGER_UNAVAILABLE |
| State trace | state_trace.jsonl or ledger equivalent |
STATE_TRACE_UNAVAILABLE |
| Caller audio | audio_caller.wav, when transport supplies it |
CALLER_AUDIO_UNAVAILABLE |
| Agent emitted audio | audio_agent.wav or audio_agent_played.wav, when capture is possible |
AGENT_AUDIO_UNAVAILABLE |
| Composite recording | audio.wav, with declared channel map |
COMPOSITE_AUDIO_UNAVAILABLE |
| Call-clock replay | audio_replay.wav, derived without altering raw tracks |
ALIGNED_REPLAY_UNAVAILABLE |
| Gemini transcript | durable transcript.jsonl rows |
GEMINI_TRANSCRIPT_UNAVAILABLE |
| Whisper transcript | durable whisper_transcript.jsonl rows or queued result |
WHISPER_TRANSCRIPT_UNAVAILABLE / WHISPER_PENDING |
| Timing | per-turn/state timing events including first agent emit | AGENT_EMIT_DELAY_UNAVAILABLE |
| Telnyx carrier recording | dual-channel WAV, both tracks, from answer | TELNYX_CARRIER_RECORDING_UNAVAILABLE |
| RIFF phone-wire reconstruction | inbound caller plus exact outbound codec PCM | RIFF_WIRE_RECORDING_UNAVAILABLE |
No raw secret, card, authentication code, or unnecessary caller PII may enter text, slot, event, or derived-debug artifacts. Transcript and slot retention use the existing redaction policy; safe fingerprints preserve change evidence where raw values cannot be retained. Spoken audio is inherently unredactable and is governed by the separate Audio Privacy and Retention Contract below.
Cross-Transport Playback Lifecycle
Per-prompt playback evidence is required before the carrier-recording lift. Both browser and Telnyx paths emit the same lifecycle schema:
| Event | Required identity and evidence |
|---|---|
audio_requested |
session, visit, state, playback ID, source kind, safe content fingerprint, request epoch |
audio_started |
identifiers, actual start epoch, transport, client/wire acknowledgement source |
audio_completed |
identifiers, end epoch, played duration, completion reason |
audio_interrupted |
identifiers, epoch, elapsed duration, interruption reason |
audio_failed |
identifiers, epoch, typed failure and safe diagnostic |
The browser creates a playback ID before dispatch and posts client-side acknowledgements from the actual Web Audio/player start, stop, completion, and error boundaries. The server persists them; browser console logs are not evidence. Telnyx emits corresponding acknowledgements at outbound codec/send boundaries and retains carrier recording lifecycle separately.
This lifecycle is the required cross-transport signal for detecting a double request or replay. Carrier audio is an independent phone-only confirmation that the duplicate was audible on the Telnyx call. A plan may pass browser duplicate detection without a carrier file, but it may not claim caller-heard phone audio without carrier confirmation.
Audio Privacy and Retention Contract
Full caller and agent audio can contain names, addresses, health information, authentication codes, and other spoken PII that cannot be redacted reliably. Therefore:
- existing
record:offpolicy must cascade to local tracks, browser capture, carrier recording, downloads, Whisper jobs, excerpts, and branch fixtures; - carrier recording is disabled until a human-approved recording/consent policy and bounded retention value are configured for the deployment;
- audio and derived transcripts require authenticated, audited access and must not be exposed by a broadly accessible debugger route;
- the finalizer records retention class and
delete_after_utc; an expiry job removes carrier downloads, local WAVs, derived Whisper output, and cached excerpts while retaining only policy-approved non-audio metadata; - capacity monitoring must alert before the configured retention window can exhaust disk; indefinite local retention is not permitted;
- export or regression promotion uses the smallest redacted excerpt/fixture necessary and never silently copies the whole caller recording.
The plan defines the engineering controls; enabling recording for real calls is
a separate human GO and policy decision.
These controls do not block implementation or testing with synthetic audio, redacted fixtures, or explicitly authorized local recordings. Privacy/retention hardening is a production enablement gate, not a prerequisite for building the playback lifecycle, debugger replay, detector, or simulation engine.
Canonical Session Layout
data/sessions/<session-id>/
session_manifest.json
bus_events.jsonl
state_trace.jsonl
transcript.jsonl
whisper_transcript.jsonl
audio.wav
audio_replay.wav
audio_caller.wav
audio_agent.wav
audio_agent_played.wav
state_program_artifact.json
replay_index.json
session_manifest.json lists each artifact, checksum, availability, producer, channel map, time base, and unavailability reason. replay_index.json is a convenience index only; the ledger and audio artifacts remain source evidence.
Telnyx Two-Sided Recording Contract
Phone replay requires two independent capture points:
- Carrier truth: Telnyx records from answer with track
both, channelsdual, formatwav, and no silence trimming. RIFF retains the recording ID, call-control/session IDs, start/stop/saved webhook epochs, checksum, channel count, sample rate, duration, and retrieval status. The downloaded carrier WAV is the default audible phone-call replay because it is independent of RIFF's local mixer. - RIFF wire reconstruction: RIFF retains inbound media after decoding and
the exact outbound PCM reconstructed from the codec bytes immediately before
ws_send. This producesaudio_telnyx_wire.wav, with a declared caller and agent channel map, plus the two mono source files.
The current recorder already has relevant inputs: write_caller() receives
decoded inbound Telnyx audio and write_agent_played() receives PCM rebuilt
from the exact outbound Telnyx payload. However, the existing composite
audio.wav combines caller audio with audio_agent.wav, which is the agent
source before the final phone codec path. Implementation must add the explicit
wire composite rather than silently treating the existing composite as carrier
truth.
The manifest must preserve all three agent perspectives when available:
| Artifact | Question answered |
|---|---|
audio_agent.wav |
What agent/model audio did RIFF produce before the phone codec? |
audio_agent_played.wav |
What outbound codec PCM did RIFF submit to Telnyx? |
| Telnyx dual-channel WAV | What did the carrier recording contain on the two call legs? |
Telnyx documents dual recording as first leg on channel A and other legs on
channel B; RIFF must not assume that means caller-left/agent-right. A canary
fixture with non-overlapping caller and agent tones verifies the channel map,
and the verified mapping is written into each manifest per direction:
inbound_channel, outbound_channel, and mapping_method. A channel whose
direction was not verified is labelled unknown, not guessed.
Recording lifecycle
- Start carrier recording as part of answer, not after the greeting begins.
- Use an idempotent command ID and retain every recording lifecycle webhook.
- Do not trim silence; preserving call-zero and gaps is required for alignment.
- Finalization remains
pendinguntil the recording-saved webhook and download complete, or until a bounded retry policy records a typed unavailable reason. - Compare carrier duration and per-channel energy with RIFF's local wire tracks. Material drift, missing leading audio, empty channels, or duration mismatch creates a capture finding and blocks claims about what the caller heard.
One replay clock
The carrier WAV, local wire WAV, local replay WAV, transcript segments,
playback lifecycle, and FSM events are normalized onto one clock_ms axis from
call answer. Each audio artifact declares its media-zero offset and any measured
alignment correction. A local audio_replay.wav may use retained caller media
timestamps plus a serialized outbound-queue estimate, but its manifest must
label that agent lane as sent/queued evidence rather than carrier-heard truth.
The UI defaults to the carrier WAV while allowing an operator to switch to
caller mono, RIFF outbound-wire mono, or pre-codec agent audio without moving
the cursor. Play, pause, scrub, step, and reverse therefore keep both sides of
the Telnyx call aligned with state and transcript evidence.
The named alignment fiducial is the first outbound playback whose exact PCM is
present in both RIFF's wire track and the carrier outbound channel. RIFF aligns
those signals by bounded cross-correlation, then validates the inbound direction
against the first caller-energy segment. If no outbound playback exists, it uses
the first inbound media segment. The manifest records alignment_method,
offset_ms, confidence, and measured drift; low-confidence alignment is
unavailable, not a guessed offset.
One Finalization Pipeline
Create a single finalize_call_evidence() contract used by browser and phone transports.
- Freeze the session ID, flow revision, engine revision, start/end epoch, and transport metadata.
- Flush the event bus and state trace for that session into its directory.
- Finalize recorder streams and record channel availability rather than creating misleading empty artifacts.
- For Telnyx calls, wait for and download the dual-channel carrier recording, or retain a typed pending/unavailable result after bounded retries.
- Persist Gemini caller and agent transcript events into
transcript.jsonlwith role, source, turn, state/visit when known, and timing. - Run or enqueue Whisper for each available permitted audio channel. Store source, model, segments, timestamps, language, and unavailable/pending reason.
- Build
state_program_artifact.jsonfrom the same frozen ledger. - Write manifest checksums and a completeness summary.
- Run the capture-completeness gate before declaring finalization successful.
Browser and Telnyx adapters may differ in what they can capture, but they must invoke this same finalizer and express gaps with the same schema. They must never silently omit a ledger or transcript.
Debugger Replay Model
The State Program Debugger consumes only retained artifacts.
- The left state tree renders every visit, including
state#2, with the entry source, exit reason, and repeat marker. - The center timeline is ordered evidence: agent/caller transcript bubbles, emitted-audio events, slot writes, guard verdicts, and transition commits.
- Transcript is always visible in the timeline. Each bubble names its source
(
Gemini,Whisper,selected, orDTMF), role, turn, state visit, and timestamp; disagreements expand inline rather than replacing one source. - Each selected state has immutable Entry and Exit snapshots, cursor-derived Current values, legal edges, winner/shadowed verdicts, and evidence lanes.
- Audio controls play only artifact-backed spans. They show unavailable or pending rather than synthesizing replacement speech.
- If no audio track has a proven sample-to-call mapping, the debugger selects no track and disables Play. An operator may explicitly audition raw tracks, but the UI labels their media position nominal and never presents their mix as the synchronized conversation.
- Regenerated or re-recorded candidate audio is always a simulated branch artifact. It is playable for comparison but can never replace, edit, or be presented as carrier/recorded evidence from the original call.
- Replay is a shared time cursor, not a list of unrelated controls:
Playadvances through timestamped transcript, audio, and FSM events;Pausefreezes every lane; scrubbing seeks every lane together;StepandReversemove one canonical recorded event; andRewind to divergenceseeks the earliest transcript/audio/transition mismatch for the selected repeated path. The timeline keeps past evidence visible above the cursor and future evidence below it, so an operator can stop precisely before or after the defect. - The inspector shows Gemini, Whisper, selected input, and DTMF independently; disagreement is a first-class finding.
agent_emit_delay_msis labelled server-side emit delay. Recorded agent audio/energy is separate evidence for caller-heard claims.
Repeat Diagnosis
A repeat is a property of visits, not state IDs. The debugger must show:
menu#1 → cap_leave_message__collect#1 → menu_announce#2 → menu#2as separate frames.- The transition reason and guard/verdict for each edge.
- Caller/Gemini/Whisper observations before the return edge.
- A repeat summary: first visit, repeated visit, intervening state path, and earliest evidence divergence.
- A typed conclusion only when supported: caller retry, no-match/retry policy, timeout/no-input, tool failure, or
REPEAT_CAUSE_UNAVAILABLE.
The implemented comparison uses semantic projections rather than raw event
objects. It compares entry context, boundary values, acted-on input,
source-labelled caller observations, slot writes, playback, and exit decision;
it ignores volatile visit/event/playback IDs, sequence numbers, and timestamps.
Entry context remains repeat topology, not a claimed behavioral defect. The
engine chooses the earliest candidate-side difference by recorded timeline
position. If an unavailable required lane occurs earlier, it returns an
earliest_known_divergence but withholds first_divergence and refuses to move
the synchronized cursor.
When telemetry for a turn arrives after synchronous transitions have re-entered
the same state, the event carries both the causal visit and the visit active at
observation time. Causal ownership wins; otherwise the debugger can fabricate a
repeat by assigning the earlier input to the later occurrence.
Work Packets
These packets are deliberately small enough for lower-capability implementation
agents. A packet owns its listed files, lands its red test first, and ends with
one focused commit. Agents must not combine an evidence-contract change with a
caller-facing flow change. No packet requires a real OTA call unless its gate
explicitly says HUMAN GO.
CP0 — Canonical Manifest Loader and Cross-Client Parity
Objective: browser, CLI, Codex, and Claude build byte-identical artifacts from one session bundle.
Primary files: new riff/postcall/session_evidence.py,
riff/state_program_debugger_service.py, scripts/state_program_debugger.py,
tests/test_state_program_debugger_service.py, and
tests/test_state_program_debugger_cli.py.
Steps:
- Introduce
SessionEvidenceBundle.open(session_dir)as the only canonical, manifest-mediated loader. It validates paths, checksums, artifact kinds, availability, channel maps, and clock maps. - Move
_audio_refs()into that loader. Stop canonical code from globbing WAVs and inventing barerecordedreferences. - Make the CLI and HTTP service call the same loader and the same
build_state_program_artifact()function. Put legacy filename fallback behind an explicitimport_legacy_session()adapter. - Add
artifact_revision, builder version, and source checksums to the derived artifact. Redefine completeness by domains: manifest, ledger, state, variables, Gemini transcript, Whisper transcript, playback, audio, and clock.
Gate: a golden session produces byte-identical JSON, completeness domains,
audio refs, and determinism hash through CLI, HTTP, and direct Python. The
current case where CLI says AUDIO_CLOCK_UNAVAILABLE while HTTP says aligned
must fail the red test and pass after the change.
CP1 — One Transport-Neutral Evidence Finalizer
Objective: browser and phone sessions freeze the same complete evidence bundle, in the same order, without relying on teardown races.
Primary files: new riff/postcall/finalize_call_evidence.py,
riff/audio/session_recorder.py, riff/live/session.py,
riff/phone/telnyx_transport.py, and focused browser/phone finalization tests.
Steps:
- Implement idempotent
finalize_call_evidence()with one session-scoped bus spool. Do not slice a global fixed-size event buffer at teardown. - Stop and flush audio, freeze the ledger and trace, persist Gemini transcript rows, register typed absences, and then write the final manifest atomically.
- Enqueue Whisper after raw evidence is frozen. Record
pending,complete,failed, orunavailable; a background result updates the manifest through an atomic revision and rebuilds the derived artifact. - Mark empty WAV headers unavailable. Retain manifest
finalized: falseuntil all synchronous evidence is durable. - Resolve and freeze flow revision, engine revision, transcript model versions,
and audio producer versions at call start. A missing revision is a named
completeness defect, not
nullwith no explanation. - Retain
agent_media_frames.jsonlwith playback ID, PCM fingerprint, source and reserved call times, actual send time, queue/clear generation, collision or overwrite result, and clipped sample count.
Gate: browser and phone fixtures produce the same layout and completeness schema; double finalization is byte-stable; a crash at each phase leaves a recoverable partial manifest rather than a falsely complete call.
CP2 — Automatic Aligned Per-Channel Transcription
Objective: every permitted finalized call receives a timed transcript on the call clock without an operator running a script.
Primary files: new riff/postcall/transcript_hydration.py, a thin
scripts/hydrate_transcript.py wrapper, the finalizer, the manifest schema,
and tests/test_hydrate_transcript.py.
Canonical transcript row: stable ID, role, source
(gemini, whisper, static_text, dtmf, or selected), text or typed
redaction, source artifact/channel, model, language, media start/end, call-clock
start/end, clock status, confidence/no-speech evidence, turn index, state/visit
attribution, and source references.
Steps:
- For new local calls, split
audio_replay.wavinto caller-left and outbound-agent-right in process; do not downmix and do not require an externalffmpegbinary. - Transcribe the channels independently, apply the retained clock map, and merge only after both roles have timestamps.
- Keep
audio_agent.wavas a separate generated/intended observation. Its transcript may be useful for suppression comparison, but stays untimed when its clock is unverified. - Persist
whisper_transcript.jsonl, list it in the manifest with its model and source checksum, and make the artifact loader consume it. - Preserve Gemini and Whisper rows side by side. Never overwrite one with the other and never treat Whisper as the selected FSM input unless the recorded resolver actually selected it.
Gate: a synthetic stereo fixture with non-overlapping phrases at known offsets yields correct roles and call times within 100 ms. A barge-in fixture proves generated menu options can exist without appearing on the played lane. No accepted transcript segment may extend beyond its source duration.
CP3 — Transcript Validation and Disagreement Adjudication
Objective: silence hallucinations and downmix artifacts are visible as rejected hypotheses, not promoted as spoken words or duplicate defects.
Primary files: riff/postcall/transcript_hydration.py,
riff/postcall/state_program_artifact.py, and focused validation fixtures.
Steps:
- Reject invalid bounds (
start < 0,end <= start, or beyond duration plus tolerance) and non-finite values. - Require overlapping channel energy for accepted speech, retain
no_speech_proband model confidence, and classify low-evidence rows asrejected_silenceoruncertainrather than deleting them silently. - Group Gemini, Whisper, static text, DTMF, and selected-input observations by overlap. Emit word diff, source confidence, agreement status, and the source the FSM acted on.
- Detect repeated audio from playback IDs, content fingerprints, and outbound channel spans. Text-only repetition is corroboration, never sufficient proof.
Gate: the LAW whole-mix false Goodbye. Goodbye. does not create a repeat
finding because the outbound mono/lifecycle contains one occurrence. The
property action/menu repeats remain findings because they exist on
played/replay evidence. Silence fixtures containing repeated hallucinated words
do not enter the selected conversation.
CP4 — One Causal Timeline and Explicit State Ownership
Objective: audio, transcript, LLM, FSM, variables, and tools become one step-able event stream.
Primary files: riff/postcall/state_program_artifact.py, live-session
telemetry emitters, riff/state_program_debugger.py, and artifact/engine tests.
Steps:
- Produce one
timeline_eventsarray ordered by(clock_ms, causal_suborder, stable_event_id). Untimed evidence goes in an explicit untimed lane and never defaults to call zero. - Persist capture-state ID, decision-state ID, turn, and visit when speech is received. Deterministic hydration may fill a missing visit by interval, but must retain that the attribution was derived.
- Join static playback fingerprints to immutable resolved text. Retain model request/instruction hash, output transcript, first token, first audio, suppression reason, playback lifecycle, tool start/result, guard verdict, slot writes, and transition commit as linked spans.
- Make conversation observations cursor events rather than a side list.
- Use recorded state entry/exit clocks for visit bands. Display visit ordinal
(
menu#1,menu#2) separately from total occurrence count. - Add
barge_in_candidate,barge_in_accepted,barge_in_rejected, andbarge_in_suppressedevents with energy, threshold, active playback ID, state/visit, decision reason, and queue-clear generation. Render outbound frame overwrite/clipping evidence on the same span.
Gate: stepping and reversing across a caller utterance selects the same audio sample, transcript group, owning visit, resolver verdict, slot write, and transition in browser, CLI, and direct API. A post-transition transcription cannot leak into the prior announce frame.
CP5 — State Boundary Variable Integrity
Objective: Entry, Current, and Exit variables are trustworthy enough to support mutation and regression promotion.
Primary files: slot/boundary emitters, state_program_artifact.py, and
state-program artifact tests.
Steps:
- Reproduce each
VARIABLE_ENTRY_FOLD_MISMATCHandVARIABLE_EXIT_FOLD_MISMATCHfrom the three study calls with a minimal fixture. - Align snapshot and fold fingerprint schemes; capture every supported writer, including seeding, nested/dict mutation, tool writes, sets, and clears.
- If a mutation cannot be instrumented, emit a typed gap naming the writer and
affected variables. Never label an unobserved variable
unchangedorignored.
Gate: fixtures and one browser/phone integration call have zero unexplained fold mismatches. Overlay editing changes only simulated Current/Exit values and never recorded Entry/Exit evidence.
CP6 — Synchronized UI and Durable Problem Markers
Objective: a human and an AI can open the same point, hear it, and inspect the same causal evidence without raw-log work.
Primary files: web/state-program-debugger.html,
web/state-program-debugger.js, debugger HTTP routes, and Playwright tests.
Steps:
- Render fixed lanes for state visits, caller Gemini/Whisper observations, selected input/DTMF, agent resolved text, generated audio, played audio, model spans, tools, slots, guards, and transitions.
- Make Play/Pause/Scrub/Step/Reverse drive one cursor and auto-focus the active row. Raw audition remains visibly separate and cannot move an aligned cursor as if it were calibrated.
- Add append-only durable annotations with point/range, session and artifact revision, event/visit refs, category, note, author, status, timestamps, and evidence window. The share URL carries annotation ID.
- Add
rewind_to_first_divergenceandcompare_visits; highlight the return edge and first changed input/variable/playback rather than merely coloring every repeated state.
Implemented checkpoint: the visit rail exposes ordinals and the immediately
previous same-state occurrence. Selecting a repeat renders one compact
comparison band, and the jump command moves the canonical cursor to the
candidate-side event. Equal or unavailable evidence fails closed with typed
errors and no cursor motion. inspect_pause_gaps is a deterministic companion
projection that classifies caller wait, agent response delay, and back-to-back
output with zero model calls; per-call model-token usage remains a typed capture
gap until the live transport persists it.
Gate: an operator marks the property status defect, opens the marker through CLI/API, replays a bounded audio range, and sees the same state and evidence hash. Desktop and mobile Playwright screenshots show no overlap or clipped labels.
CP7A — First Behavior Regression: Property Status Loop
Objective: use the completed debugger to fix the observed call, then promote the evidence into a cold-rebuild regression.
Primary files: riff/business_pack/compile.py, locked-choice retry ownership
in the FSM engine, tests/business_pack/test_replicant_compile.py, and a
recorded/synthetic replay fixture. Do not hand-edit generated
flows/property_management.yaml as the source fix.
Required defects to prove independently:
cap_existing_request__statusentered and exited in roughly 4 ms onwhen: always, with no playback lifecycle between them, so the status summary was skipped.__locked_choice_retry_cap_existing_request__choose_actionsurvived return to the capability and immediately satisfiedchoice_retry_exhausted._replicant_selected_ticket_idsurvived menu return, so laterselect_ticketvisits auto-advanced without a new selection.
Steps:
- Add a compiler/runtime gate: a caller-facing state with resolved speech may not take an immediate edge until an observable speech-completion boundary. For dynamic status text, generate an announce boundary whose success/failure edges both preserve explicit playback evidence.
- Make locked-choice retry counters state-visit scoped or clear them on every matched/exhausted exit.
- Clear selected-ticket ID/title/status and capability-local choice/action state when returning to the main menu.
- Recompile the property flow from the pack and run the cold-rebuild test.
Gate: one status request speaks the recorded status once, then offers the next operation once. Re-entering existing requests waits for a fresh ticket choice and a fresh action. The debugger reports no immediate retry-exhausted loop and no unexplained duplicate playback.
CP7B — Second Behavior Regression: LAW Barge-In and Case Status
Objective: prove why an early Status was ignored and prevent a capability
question from being mistaken for a completed case-status message.
Evidence from v3:SsLr...: the caller says Status during the full menu,
but no accepted interruption event exists; the repeated post-menu choice routes.
After the caller asks whether cases can be listed, the agent states its
limitation, repeats the office-message question, and immediately returns to the
main menu.
Steps:
- Use CP4 barge-in decision telemetry to distinguish insufficient energy,
disabled barge-in, echo rejection, and accepted interruption. Do not infer a
reason from the absence of
audio_interrupted. - Separate case identifier, caller message/status question, and capability
question in the LAW capability schema. A question such as
can you list them?must not satisfy the message-complete guard by itself. - After a limitation response, offer a concrete recovery (
tell me the case name/number,leave a message, orreturn to menu) and wait for a new caller decision. - Fix the source generator/compiler and run a cold rebuild; do not patch only the generated LAW YAML.
Gate: the first menu-time Status either interrupts and routes or appears
as a rejected candidate with the exact reason. A case-list capability question
does not auto-complete the collection or return to the menu without a caller
choice.
CP8 — Immutable State Capsule and Branch Simulation
Objective: let Codex, Claude, and the UI change an input, variable, tool fixture, or candidate audio and replay forward without altering recorded truth.
Primary files: new capsule schema/runner, debugger engine/API, CLI, UI, and
simulation tests. Existing scripts/replay_branch.py remains explicitly legacy
until replaced.
Steps:
- Export
riff.state_test_capsule/v1: exact state definition and legal edges, recorded engine/flow revisions, Entry snapshot, selected and alternate transcript observations, DTMF/control input, deterministic tool fixtures, audio outcome, and evidence refs. - Run the capsule in a bounded subprocess with network disabled, side-effecting tools rejected, fixture-only reads, and explicit event/time budgets. Missing recorded code must fail; never fall back silently to HEAD.
- Implement
fork,set_overlay,set_input,set_audio_fixture,run,compare_branches, andpromote_regressionover the same structured API used by browser and AI clients. - Store simulated output separately with purple provenance, fork root, overlays, candidate revision, exit/variable/audio deltas, and invariant results. Recorded artifacts remain immutable.
Gate: mutate the property status state from the misheard input to status,
run forward with a fixture ticket, hear candidate audio, compare recorded and
candidate exits/variables, reverse to the fork, and export a deterministic test.
UI, CLI, Codex, and Claude return the same branch hash.
CP9 — Carrier-Heard Confirmation (HUMAN GO)
Objective: distinguish RIFF sent/queued evidence from what the phone carrier recorded on both legs.
Primary files: Telnyx recording lifecycle, finalizer/download worker, manifest channel/alignment schema, and synthetic/canary tests.
Steps: start dual-channel recording from answer, retain recording webhooks, download without silence trimming, verify direction with non-overlapping tones, align to the outbound PCM fiducial, and record offset, drift, confidence, and channel map. Low-confidence alignment remains unavailable.
Gate: caller-only, agent-only, overlap, barge-in, interrupted playback, and early hangup are audible on the expected carrier channels. Enabling this on a real line still requires explicit human authorization; all prior packets run on local or synthetic evidence.
Handoff Contract for Implementation Agents
Every assigned packet must include:
- the single packet ID and objective;
- exact allowed files and explicit files that are out of scope;
- the failing test to add first;
- fixture/session IDs and expected evidence codes;
- commands to run and the required pass criteria;
- whether a service restart, browser cache bust, or human GO is required;
- resulting commit, changed files, evidence artifact, and next dependency.
Lower-cost models are appropriate for CP0 loader consolidation, schema fixtures, HTML rendering, and narrow UI/test packets. Use stronger models for CP1 teardown concurrency, CP3 transcript adjudication, CP7 FSM ownership, and CP8 sandboxed simulation. Parallel agents must use isolated worktrees and must not edit the same schema or generated flow concurrently.
Acceptance Gates
- A finalized call cannot have audio without explicit ledger, transcript, and clock availability results.
- CLI, HTTP, browser, Codex, and Claude consume the same manifest-backed bundle and return the same artifact bytes and determinism hash.
audio_replay.wavduration equals manifest call duration within 20 ms; its channel map and clock provenance are retained. Raw tracks remain unverified unless separately calibrated.- Automatic Whisper transcribes replay channels independently. No accepted row has invalid bounds, exceeds source duration, or lacks role, source, model, clock status, and evidence reference.
- Every speech-energy span longer than one second is covered by an accepted or uncertain transcript row, or by a typed reason. Silence hallucinations do not become selected conversation.
- Gemini, Whisper, static text, DTMF, and selected input remain distinct. Disagreement shows what the FSM actually used.
- Play, pause, seek, step, and reverse keep audio, transcript, state visit, variable projection, playback, model/tool span, guard, and transition on the same call-time cursor.
- Every state repeat renders as a distinct visit ordinal with the return edge, total occurrence count, and earliest causal divergence shown separately.
- Entry and Exit folds match recorded boundary snapshots, or name the exact uninstrumented writer and affected variables.
- Barge-in candidates record accepted/rejected/suppressed outcomes and queue clear generation. Replay derivation records collisions, overwrites, clips, drift, and confidence.
- A durable issue marker opened by browser, CLI, or AI resolves to the same bounded audio range, event/visit references, and artifact revision.
- A controlled double-play produces two matching lifecycle/audio spans with owning visits; a downmix-only duplicate transcript produces no false defect.
- The property status regression speaks status once, resets retry/selection state, and does not enter an immediate menu loop.
- The LAW status regression explains the early barge-in decision and does not auto-complete a capability question as an office message.
- A simulated branch is immutable, network-isolated, separately labelled, reproducible by hash, and cannot overwrite or masquerade as recorded audio or events.
- Every call freezes flow and model revisions or reports a typed reproducibility gap.
- Carrier-heard claims require the separately authorized Telnyx dual-channel artifact; local replay remains labelled sent/queued evidence.
Rollout
- Land CP0 first. No additional UI or AI client may introduce another session loader while artifact parity is unresolved.
- Land CP1 and CP2 behind the retained-call feature flag. Run browser and phone fixtures; keep asynchronous Whisper failures typed and non-fatal.
- Land CP3 through CP5 and make domain completeness visible. Keep the gate warning-only for one release while collecting real gap counts.
- Land CP6, including durable markers, after the canonical timeline is stable. A human must replay the property repeat without consulting raw logs.
- Use that marker to implement CP7A and CP7B, cold-rebuild both generated flows, and promote the two calls into regression fixtures.
- Land CP8 only after recorded evidence gates pass. The Fork button remains disabled until capsule validation and isolation tests are green.
- Implement CP9 with synthetic/canary evidence, but do not enable real carrier
recording without explicit
HUMAN GO.
This sequence makes the reported double-play observable and attributable. It does not itself fix the duplicate; behavior changes follow only after replay evidence identifies the responsible request, client playback, state transition, or model narration.
Operator Validation
After CP0, run one browser simulator call and one phone-path fixture. For each:
.venv/bin/python scripts/state_program_debugger.py build data/sessions/<session-id> --out /tmp/<session-id>.json
.venv/bin/python scripts/state_program_debugger.py command /tmp/<session-id>.json --json '{"action":"list_visits"}'
Open /state-program-debugger.html, enter the session ID, and verify the same visit list, completeness status, and repeat markers. A call without required evidence is a capture failure to fix before using it to judge flow behavior.
To persist the exact defect point for browser, Codex, and Claude:
.venv/bin/python scripts/state_program_debugger.py session-command \
data/sessions/<session-id> \
--json '{"action":"record_issue","issue":{"issue_id":"issue:repeat-1","status":"confirmed","time_ms":123400,"event_id":"bus:42","visit_id":"menu#3","reporter":"codex","category":"duplicate_agent_playback","note":"Same outbound prompt twice."}}'
This appends an annotation; it never rewrites recorded audio, events, or state.
Then run automatic transcript hydration and verify that the browser timeline,
inspect_timeline, and the CLI expose identical caller/agent transcript groups,
audio refs, visit ownership, and artifact hash. The study session
v3:B-ly... is the required negative/positive control: its real outbound
repeats must remain visible, while raw-track timestamps past 206.507 s and
downmix-only duplicate words must not enter the synchronized lane.