Changes apply to the next call. Preview first — it synthesizes one line with exactly these settings.
A few seconds of clean, non-looping speech as a 16-bit WAV, plus its exact transcript — a mismatch quietly degrades the clone rather than failing.
Fully local pipeline on one RTX 3090: Silero VAD → Smart Turn v3 → Whisper large-v3 → PhoneLLM Alpha 1 → Breeze TTS 2.
Silero VAD (CPU)Smart Turn v3 (ONNX, CPU)faster-whisper-large-v3phonellm-alpha-1 Q4_0Breeze TTS 2 streamingWebRTC