What is loading
A local neural voice is more than a small sound file. TTS Lines downloads model data from a public host when you first generate or preview a local voice. English and many other languages use one model; Hebrew uses a separate pronunciation and speech pair.
The files may be cached, but a cache can be cleared or evicted. A later visit is not guaranteed to start instantly. The model also needs memory and processing time after the download ends.
If it appears stuck
Start with a short line and wait for the voice engine status. Check that the connection can reach the model host. If a low-memory phone closes the tab or repeatedly reloads it, try a desktop browser. Avoid opening several generation tabs at once.
When the engine is ready, generate one line. If that works, continue with the rest of the script. Saving a copy of your text outside the browser protects it if you need to reload.
A fair speed comparison
Measure from a cold first visit separately from a return visit with cached models. Device, browser, connection, language, and script length all matter. This article gives no universal load-time promise because we have not published a matched device benchmark.