NarrateNow
← Back to blog
Engineering·4 min read

How We Built a Text-to-Speech System That Doesn't Sound Robotic

The history of TTS is a history of systems that technically work but feel wrong. Here's what we learned building something that actually sounds like a person.

The history of text-to-speech is littered with systems that technically work but feel wrong. You know the voice immediately when you hear it: the flat cadence, the misplaced emphasis, the mechanical rhythm that signals "this is a machine." Even as synthesis quality has improved dramatically, that gap between technically correct and genuinely listenable has been hard to escape.

When we built NarrateNow, we spent a significant amount of time on the question of voice quality. Not because we thought we could build better speech synthesis from scratch, which we neither could nor tried to do, but because selecting the right foundation and tuning how we use it turns out to matter a lot.

The problem with generic TTS

Most TTS APIs are general-purpose. They're designed to read anything: navigation prompts, error messages, product descriptions, legal disclaimers. The voice needs to handle all of it, so it's tuned for neutrality.

Blog content is different. It has a register that is conversational but considered, informal but not careless. It has rhythm, rhetorical structure, emphasis patterns. When you strip all of that and feed it to a general-purpose voice, you get something that's technically accurate but tonally wrong.

The first version of our audio output confirmed this. The transcription was correct. The voice quality was adequate. But it felt like being read to by a particularly thorough bureaucrat.

What actually affects perceived quality

Through testing with real content, we found a few factors that matter disproportionately:

Sentence boundary handling. The pause between sentences is where a lot of robotic TTS fails. Too short, and the audio feels rushed and breathless. Too long, and it sounds halting. The ideal pause length also varies with sentence structure, since a short punchy sentence wants a different breath than a long compound one.

Proper noun pronunciation. Blog content is full of proper nouns: brand names, technical terms, names of people and places. Generic TTS systems often mispronounce these in ways that immediately break the illusion. We built a preprocessing layer that handles common cases and lets us add overrides for domain-specific terms.

Heading and list treatment. Markdown structure implies prosodic structure. A heading should sound like a heading: slightly more deliberate, with a natural pause before and after. A bulleted list has a different rhythm than flowing prose. We parse the document structure and add prosody markup accordingly.

Contraction and informal speech patterns. Good prose often includes contractions and casual constructions that formal TTS systems handle awkwardly. We found that choosing a voice trained on conversational speech made a significant difference here.

The preprocessing pipeline

Before text reaches the synthesis engine, it passes through a preprocessing step that does several things:

First, it strips or converts any markdown syntax. Asterisks become emphasis markers, not literal characters. Backtick code spans get different treatment than prose.

Second, it handles common abbreviations and numeric formats. "2.5x" should be read as "two-point-five times," not "two-point-five-ex." "Jan." should expand to "January." "e.g." should become "for example." There are hundreds of these cases, and getting them wrong is distracting.

Third, it adds SSML markup for prosodic control, covering pauses, emphasis, and rate adjustments. This is where the structural parsing matters most. Headings get a slight rate reduction and a preceding pause. Lists get consistent inter-item pacing.

Caching as a quality lever

One underappreciated aspect of audio quality is consistency. If you generate audio fresh on every request, small variations in synthesis can make the same content sound different between sessions. Listeners notice this in ways they can't always articulate.

By caching generated audio at the page level, we get a secondary benefit beyond performance: every listener hears exactly the same audio for a given page. The quality is fixed at generation time, and it's consistently good rather than inconsistently variable.

When a page is updated, we regenerate the audio. This also means the audio always matches the current version of the content, which turns out to matter for SEO and trust. A page where the written and audio versions say different things is a problem we wanted to avoid entirely.

Where we're headed

Voice quality is improving faster than almost any other AI capability. The gap between synthesis and human recording is narrowing, and the content types we focus on, such as blog posts, articles, and documentation, are where that gap matters least. Listeners aren't comparing your audio player to a studio recording. They're comparing it to not having audio at all.

We're already past the threshold that matters: audio good enough that listeners focus on the content, not the voice. The occasional edge case aside, that's where NarrateNow sits today. The voice handles itself. The content stays yours.