Pipecat produced replies. Our local audio setup still took too long.
Nine offline audio runs completed. The eight warm runs reached first generated audio at a median of 2.9 seconds, before playback or telephony.
The test bench / 9 field reports
We try a specific task with a tool or repo, and keep the evidence. Here’s what worked, what didn’t, and what a small test can actually tell you.
The version, task and conditions.
Measurements you can inspect.
What we didn’t establish.
Nine offline audio runs completed. The eight warm runs reached first generated audio at a median of 2.9 seconds, before playback or telephony.
Sixty answers from one AnythingLLM configuration revealed repeatable failures, including a user message that overrode the real cancellation policy.
A signed-in extension found an existing HeyGen output. That proves a useful task worked; it does not establish cheaper browser automation.
Six matched pairs gave the existing setup five wins. Here is how the blind review worked and why we are skipping this particular extra checklist.
Sam accepted 11 of 15 generic backgrounds. The paid comparison is unfinished, and the access terms keep this out of our automated pipeline.
A bonus field-notes series on image generation, design prompts, browser automation, website-grounded answers and voice latency. Small tests, specific decisions, and the gaps still open.
A disposable PostgreSQL database, synthetic records, and two read-only setups. The rows stayed unchanged, but a batch’s outer success flag did not tell the whole story.
One Shossip guide page, a clearer starting point, and a coverage note that moved from 2.59:1 to 7.35:1 contrast. A design case study with the before and after.
A paired coding experiment produced a smaller diff and the same passing tests. The cost result is why we are keeping the verdict open.