An API audit found the real root cause: no language signal on short words, plus a model (v3) the voice isn't even verified for. What's left: with that fixed, does plain text already sound right, or does it still need alias respelling? Everything already decided is summarized below, not re-litigated.
Every clip below uses the corrected setup: model
eleven_flash_v2_5 (voice-verified, unlike v3), language_code:
"id" (the missing signal the audit found), and stability-tuned
voice_settings (no more hallucinated pauses or question-intonation on short
input). The only difference between the two players in each row is
whether a pronunciation dictionary was attached.
ARM N (no dict) — plain,
correctly-spelled Indonesian text, zero pronunciation rules. This is the
hypothesis test: if cepat comes out “ch” (not
“s”) and the e-vowels land right with no respelling at all, the
language signal was the missing piece all along — not the alias layer.
ARM D (alias dict) — identical setup + one dictionary
(created once for this run) with the derived-convention respelling for each
word. Measures whether respelling still adds anything once the model is
correctly configured.
Tap a verdict on each row — plain is right / needs respelling / both wrong.
One tap each — a free-text note is optional.
eleven_flash_v2_5 +
language_code: "id" (+ one persistent alias dictionary, only
for the words D-TTS-7 says still need it) as THE pronunciation strategy?
This replaces every prior TTS setup in the app — no more v3, no
more per-run temp dictionaries. If D-TTS-7 is
alias-for-irregulars, say which words in a note; that becomes the
dictionary's rule set.