Generating a WAV file and making that speech audible to a caller are separate tasks. The first is speech synthesis; the second is audio routing. A TTS command can succeed while the call transmits silence because the generated audio never reaches the softphone’s microphone input.
Define the call environment and model first
This workflow targets a Linux desktop with PulseAudio or PipeWire’s PulseAudio compatibility service and a softphone that accepts a selectable microphone. It does not assume that an Android cellular call accepts desktop audio. A physical phone needs a supported hardware or software audio interface; playing a file on a nearby computer is not an injection method.
Use an installed, tested Coqui-compatible environment and model. Coqui’s linked documentation describes TTS 0.22.0; newer maintained distributions can have different Python, PyTorch, and model requirements. Record package and model versions instead of treating an old installation command as universal. Check the chosen model’s license before publication or commercial use.
tts --help
tts --list_modelsChoose a supported single-speaker model for the first test. Multispeaker or multilingual models can require additional speaker and language arguments. No voice cloning is necessary for this audio-routing exercise.
Generate and listen to a short phrase
tts --text "Hello. This is a synthetic speech audio test." \
--model_name REPLACE_WITH_SELECTED_MODEL_NAME \
--out_path ./speech.wavListen to the resulting file locally before involving the call. Check pronunciation, clipping, long pauses, and whether the file was actually written. Model download and startup can dominate the first run. For interactive use, measure a warm synthesis as well as a cold start, and keep the model loaded in a supported long-running application if needed.
Use a generic synthetic voice and identify the audio as synthetic to the person participating in the test. Keep private call content out of model prompts and public logs.
Create a dedicated virtual audio sink
pactl info
pactl load-module module-null-sink sink_name=tts_call \
sink_properties=device.description=TTS_Call
pactl list short sinks
pactl list short sourcesThe load command returns a module ID; record it for cleanup. A PulseAudio-compatible null sink normally exposes a monitor source such as tts_call.monitor. That monitor captures audio sent to the sink and can be used as the softphone’s microphone. Confirm the actual name in the source list.
paplay --device=tts_call ./speech.wavSelect the monitor of TTS_Call in the softphone’s microphone settings. If the application does not expose it, use an audio routing utility such as pavucontrol to select it for the active recording stream. A native PipeWire graph tool is another option when compatibility-module behavior differs.
Keep received call audio out of the transmitted path
Send the softphone’s speaker output to your normal headphones, not to TTS_Call. Routing received audio into the same monitored sink sends the caller’s speech back to them. Keep operating-system notification sounds and unrelated applications out of that sink too.
Use the softphone’s microphone meter and a local recording test to verify that only the generated phrase appears. Start with a dedicated TTS input; mixing a physical microphone is a separate routing step. Adjust levels conservatively and test whether automatic gain control or noise suppression distorts the synthetic voice.
Test the call and measure the delay
Place a test call to a consenting participant or an approved test endpoint. Generate one short phrase and ask whether it was complete, intelligible, and free of echo. Measure text submission to remote playback, including synthesis, buffering, and the call’s own latency. A codec mismatch at a SIP gateway belongs to that gateway’s supported configuration, not to an arbitrary WAV sample-rate change.
For a file-based workflow, serialize playback so phrases do not overlap. Define an interrupt or stop control before adding automation. Longer output can arrive too late for a natural conversation even if it sounds excellent locally.
Restore normal audio after the test
Return the softphone microphone to the intended physical device before unloading the virtual sink. Unload only the module ID created for this test:
pactl unload-module REPLACE_WITH_RECORDED_MODULE_IDSave the tested model, sink, source, application selection, and measured latency. The setup is complete when a remote participant hears the intended phrase and the local audio path returns cleanly afterward. Keeping synthesis and routing tests separate makes later failures much easier to locate.