MojoPad

Read Aloud

MojoPad reads to you. Any page, web clip, PDF, or imported book can be read out loud — on your Mac, with the sentence and word you're hearing highlighted as it goes, so you can follow along or just listen. It's the twin of Dictation: that turns your voice into text; this turns your text into voice.

Start a read

Open any page and press ⌥⌘R, click the speaker in the format bar, or choose Edit ▸ Speech ▸ Read Page Aloud. Reading begins from the top, and a small reading bar appears with the controls. To start partway down instead, put your cursor where you want to begin and choose Edit ▸ Speech ▸ Read Aloud from Here (⇧⌥⌘R) — or right-click and pick it. That's the quick way to skip a book's front matter and start at chapter one.

It reads whichever pane you're in. With a page open beside your writing in Split View, select words over there and press the speaker, or right-click in that pane and choose Read Aloud from Here: the reading follows the pane you acted in, not the page you happen to be writing in. Listening to a source on one side while drafting on the other is the point of it.

The reading bar

While a read is running, the bar shows play / pause, a stop square, and your place (sentence 12 / 340). Change the speed on the fly — 0.8× to 2× — and it takes effect from the next sentence. Pick a different voice from the same bar and the read switches to it. Reading also stops on its own the moment you leave the page, so a read always belongs to the page it started on.

Two kinds of voice

System voices — every voice built into your Mac (System Settings ▸ Accessibility ▸ Spoken Content has more to download). They're instant and fully offline. Enhanced and Premium system voices are grouped at the top of the picker.

MojoPad Voice — a small neural voice that runs entirely on your Mac for a warmer, more natural read (choose from Heart, Bella, Nicole, Michael, Fenrir, Emma, or George). The first time you pick one, MojoPad asks to download the voice model once (roughly 90–330 MB, depending on your hardware). After that it lives on your Mac and nothing you read aloud ever leaves the machine — the download is the only moment it touches the network.

Writing in more than one language

Writing a paper in English that quotes Japanese sources? Or working on a Mac set to Japanese, in a document that keeps slipping into English? Read Aloud hands each part of a sentence to whichever of your voices can say it — your main voice reads the English, your second voice reads the Japanese, inside the one sentence, without you touching anything as it goes.

It matters because a voice can only read the writing it was made for. Handed Japanese, an English voice does not mispronounce it — your Mac falls back to reading out the words “Japanese character”, once per character, so 研究 is announced twice rather than read. That is a limit of the voice, not of the words, and it is why the sentence is split rather than simply handed over.

Set both voices in Settings ▸ Reading Aloud, before you start — or from the reading bar while a read is running, which takes effect from the next sentence. There are two:

  • Main voice. Reads everything your second voice doesn't.
  • Second voice. Reads anything the main voice can't pronounce, and it offers whatever the first one can't say: choose an English voice and it lists the Japanese ones; choose a Japanese voice and it lists the English ones — and the MojoPad Voices with them, so the English words inside your Japanese can be read by the better voice rather than merely a different one. Choose automatically keeps the same person where macOS ships them in both languages — pick Eddy and you keep Eddy. Or say Don't switch voices and have everything read by the one voice as it was before.

It works in both directions, which is the part worth knowing if your Mac runs in Japanese. The voice you already hear everywhere else on that machine is a Japanese one, and handed English it reads the words with Japanese phonetics. You do not have to nominate an English voice, or swap your two voices round when you change language: the question asked is never “is this Japanese?” but “can the voice in hand say this?”, so the English is handed to an English voice on its own.

The MojoPad Voices speak English, and only English — the model they use has no other language in it, so this isn't a setting waiting to be found. Choose one and Japanese in the same sentence is read by your second voice: the on-device voice and your Mac's own take turns inside the one sentence, and you hear a sentence, not a seam. If your Mac has no Japanese voice installed, there's nothing to take the turn — add one in System Settings ▸ Accessibility ▸ Spoken Content, and the diagnostic report will tell you whether you have one.

Your Mac has more voices than it has installed, and better ones. Both lists show every voice on the machine, and macOS ships a great many it doesn't install: other languages, and other accents of English — Australian, Irish, Indian, South African. Add them in System Settings ▸ Accessibility ▸ Spoken Content ▸ System Voice ▸ Manage Voices…, then reopen Settings and they are in both lists. The ones marked Premium or Enhanced are much better than the plain versions and are a download rather than a purchase — and the plain versions are what most people have heard, which is why Mac voices have the reputation they do. If you want a particular accent reading your pages, that is where it comes from rather than from the MojoPad Voices.

A word of one character counts. 本, 人, 国 and 私 are each a whole word, and each gets the Japanese voice, the same as a longer term does.

Everything else goes on counting in whole sentences — the place-marker, the progress, the page a PDF turns to — so nothing else about reading changes. With no Japanese voice installed the second picker does not appear at all, and neither does any of this.

It follows along everywhere

Read a note and the spoken sentence and word light up in place — the page scrolls to keep them centered, like a teleprompter. The same follow-along works in web captures, in PDFs (the reader turns pages as it reads), in ePub books (it reads the chapter you're on), and on canvas cards — the speaking card wears a ring. Folded-away content stays silent; what you've collapsed isn't read.

Keeping a reading: audiobooks

The thirty pages you have to read before Thursday, on the walk to the station. Or the draft of your own chapter, heard back on the drive home, where the sentence that does not work is obvious in a way it never is on screen. Or a paper in two languages that no audiobook app can read, because it does not have your second voice. Edit ▸ Speech ▸ Save as Audiobook… takes whatever MojoPad would read aloud and keeps it: a file with chapter marks, in a player you already have.

It is the same reading. The same sentences, the same voices, the same turn-taking inside a sentence that mixes writing systems — the difference is only that it is kept instead of heard once. Set your voices first in Settings ▸ Reading Aloud; whatever they are, that is what the file will sound like.

Your voices are matched for loudness. Voices are not built to a common volume and they are a long way apart — on an ordinary Mac the quietest system voice sits eight decibels below an ordinary one, and the voice chosen automatically for Japanese happens to be the quietest of all. Reading aloud you never notice, because you set the volume once and listen. In a saved book you would: every change of language would be a change of volume, in a file you cannot adjust as it plays. So each voice is measured before the reading starts and brought to the same level, and the join between them is a change of voice and nothing else.

Where the chapters come from

Chapter marks are the whole difference between an audiobook and a long slab you scrub through, so MojoPad takes them from whatever the source already knows rather than guessing:

  • An ebook brings its own chapters, named as its author named them.
  • A PDF that carries a table of contents is divided by it, and nested sections are found however deep they are written. A PDF with no contents — most scans — is divided by page, so you can still jump to roughly where you were.
  • A page you wrote is divided at its headings. It divides at the biggest heading you actually used, so a page written entirely in second-level headings is divided by those rather than treated as one lump. Headings below that one are read as part of the section they belong to, which is where they belong. A page with no headings at all is one chapter — there is nothing to divide it by. If you want an hour of reading to be navigable, give it headings before you save it; they cost nothing and they are what the chapter list is made of.

A short breath is left between chapters, and the mark sits on the first word rather than on the pause, so jumping to a chapter starts at the chapter.

It takes as long as reading takes

A book is hours, not seconds — it is a real reading, at reading speed, and there is no way round that. So it does not hold the app hostage while it happens. Choose where the file goes and MojoPad gets on with it in the background while you carry on working; a small card in the corner shows which book, how far in, and how much sound exists so far, with Pause and Stop on it.

Quit part way through and nothing is lost. The reading made so far is on disk, and the card comes back when you next open MojoPad with a Continue on it. That is the difference between a feature you can start on a Friday and one you have to sit and watch.

The diagnostic report has a Saved readings section: which readings are in progress, how far each has got, and — the part you cannot hear from the outside — which voices actually read it. If a book came out in one voice when you expected two, that is where it says so.

A word about whose book it is

This is for listening to your own material: your notes, your drafts, the papers and the books you are working through. A recording of somebody else's book is still somebody else's book, and making one to hand around is not something the feature is for.