Narrate your manuscript with Text-to-Speech

Arqenne can read your manuscript aloud — paragraph by paragraph, in any of its built-in voices — entirely on your Mac. There's no cloud round-trip, no per-minute fee, and no token limit. Once the voice models have downloaded, narration works completely offline, and your text and audio never leave your machine.

This guide covers turning a manuscript into spoken audio: switching the editor into Audio mode, choosing voices, generating and tuning narration, and listening back to a full chapter.


Before you start

You need:

  1. Arqenne installed and open, with a manuscript loaded in the editor.
  2. The voice models downloaded. These download automatically the first time you launch Arqenne — there's no manual install step. Standard Text-to-Speech is small (~90 MB) and ready quickly; Enhanced Text-to-Speech is larger (~3.5 GB) and may take a few minutes the first time. After that, everything runs offline.
  3. An Apple Silicon Mac. Standard Text-to-Speech is light enough to run on any supported Mac; Enhanced Text-to-Speech (and voice cloning) use the GPU and require Apple Silicon.

If you just opened the app, give the engine a moment. The voice engine loads on startup, and generation only works once it's ready. See Troubleshooting if a paragraph won't render right after launch.

For more on what gets downloaded and where it lives, see Downloading & managing local models.


Standard vs Enhanced Text-to-Speech

Arqenne ships two on-device voice engines. You choose between them per voice — pick a fast voice while you're drafting, then switch the same paragraph to a studio-quality voice when you're finalizing.

Standard Text-to-SpeechEnhanced Text-to-Speech
Best forDrafting and proofing by earStudio-quality narration, final audio
SpeedFast, low-latencyHeavier and slower to render
QualityClear and naturalRicher, more expressive
Voice cloningNoYes (zero-shot, on-device)
Runs onAny supported MacApple Silicon, uses the GPU
Model size~90 MB~3.5 GB

Both engines run an on-device neural model, and both produce the same 24 kHz WAV audio. The difference is the trade-off: Standard is quick enough to iterate paragraph by paragraph; Enhanced takes longer but sounds polished and can clone a voice from a short reference recording.

A practical workflow: draft and proof with a Standard voice, then re-render the paragraphs you're keeping with an Enhanced voice for the final pass. See Tips for a good narration.

For voice cloning specifically, see Voice cloning with Enhanced TTS.


Narrate a manuscript, step by step

  1. Open a manuscript in the editor.
  2. Switch to Audio mode. The editor has a Text / Audio toggle. Text mode is for writing; Audio mode is for narration.
  3. Each paragraph becomes its own lane. In Audio mode, every paragraph is shown with its text on top and audio controls underneath. You work one paragraph at a time.
  4. Choose a voice. Each lane has its own voice selector. Pick the voice you want for that paragraph.
  5. Click Generate Audio. Arqenne renders that paragraph locally on your Mac. Longer paragraphs take a little longer.
  6. Listen back. When rendering finishes, a waveform and playback controls appear in the lane. Press play to hear it.

That's the whole loop: pick a voice, generate, listen. Repeat per paragraph, or use Play All to hear everything in order.


Choosing voices

Arqenne includes 29 built-in voices, shown by name and spanning American and British accents, female and male — for example Thalia, Athena, Morpheus, Odysseus, and Circe.

PlanVoices
Free1 voice (Morpheus)
PaidAll 29 voices

The free tier includes one voice so you can try narration end to end. Unlocking the full set requires a paid plan — see pricing.

Different voices for different paragraphs

Because each lane has its own voice selector, you can assign different voices to different paragraphs. This is useful for dialogue-heavy work: give the narration one voice, and a character's lines another. Set the voice on a lane before you generate it; the voice is baked into that paragraph's audio.


Tuning a paragraph

Each voice has two adjustments that shape how a paragraph sounds:

  • Speed / pacing — how fast the voice reads, roughly 0.5x to 2x. Slow it down for a measured, dramatic read; speed it up for a brisk proofing pass.
  • Expressiveness — how flat or animated the delivery is. A lower setting is steady and even; a higher setting is more emotive.

Set these before generating a paragraph, then listen back and adjust. Because Standard voices render quickly, it's cheap to try a few settings and keep the one that reads best.

Stale detection — regenerate only what changed

When you edit a paragraph's text after generating its audio, Arqenne flags that paragraph as stale. The old audio is kept so you can still hear it, but the lane tells you it no longer matches the text.

This means you only regenerate what actually changed. Tweak one sentence in a long chapter, and just that one paragraph is marked stale — the rest stay current. Regenerate the stale lane and you're back in sync.


Listen to the whole thing

The Play All control plays every rendered paragraph in document order, as one continuous read-through. As it plays, Arqenne highlights the paragraph currently being read, so you can follow along and spot anything that needs work. It shows progress as it moves through the chapter, and you can stop at any point.

Play All only plays paragraphs that have been rendered. If a lane hasn't been generated yet — or is marked stale — generate it first so it's included in the read-through.


Exporting your narration

When your narration sounds right, you can export the finished audio out of Arqenne — either as individual paragraph files (stems) or as a single combined master track for the whole manuscript.

Export formats depend on your plan: Starter exports MP3, Creator adds WAV, and Studio adds M4B — the audiobook container used by Apple Books and most audiobook platforms. The free tier lets you narrate and listen inside the app; exporting requires a paid plan. See pricing for the full comparison.


Tips for a good narration

  • Proof by ear. Hearing your prose read aloud surfaces clunky phrasing, repeated words, and run-on sentences that your eye glides past on the page.
  • Iterate with Standard, finalize with Enhanced. Standard voices are fast, so use them while you're still editing. Once a paragraph is settled, switch its voice to an Enhanced one and regenerate for the final, studio-quality version.
  • Regenerate only stale paragraphs. After an edit, look for the stale flag and re-render just those lanes instead of the whole chapter.
  • Match voices to roles. For dialogue, a distinct voice per character makes a read-through far easier to follow than a single narrator.

Troubleshooting

"This voice isn't available on your plan"

The free tier includes one voice (Morpheus). To use any of the other 28, either pick the included voice or upgrade — see pricing.

A paragraph won't generate right after launch

The voice engine loads when Arqenne starts up. On a fresh launch this takes a few seconds (longer the very first time, while models download). Wait a moment and try Generate Audio again.

Generation is taking a long time

A few things make rendering slower: long paragraphs (more text to synthesize), and Enhanced voices, which produce higher-quality audio and are heavier than Standard voices. If you're iterating quickly, draft with a Standard voice and save Enhanced for the final pass.

A paragraph still sounds like the old text

You probably edited it after generating. Check for the stale flag on the lane and regenerate that paragraph so the audio matches your current text.


See also


Privacy

Narration is generated entirely on your Mac. Your manuscript text and the audio it produces never leave your machine — there's no cloud service in the loop, no per-minute charge, and no token limit. Once the voice models have downloaded, Text-to-Speech works fully offline.

See plans and voice limits → · Get Arqenne →