Writing for the Ear

We give a text to a machine, and it reads it back to us. The voice is so natural that calling it a machine reading already feels strange. The intonation is right, the breathing is right, the pauses fall in the right places. What the technology has reached is astonishing.

But something happens as we listen. Our mind slips. Halfway through a sentence we lose the thread, we want to go back to the start, we ask ourselves what was just said. The voice is flawless, yet the experience of listening is still poor.

This is where I want to begin. Because the strangeness here points to something deeper than we suppose. If the voice is good, if our hearing is sound, if the machine reads correctly, why does comprehension still slip away? The answer is this: the text we were listening to was never written to be heard.

Let us first set out what is actually before us. You give a machine a text, and it reads it to you. Let us call this voicing what is written. It is what the phone does, what the computer does. It takes the text, turns letters into sound, reads from beginning to end. It changes nothing. Whatever is placed before it, that is what it voices, exactly as it stands. This is the ordinary case. The usual, the expected, the work everyone knows.

Now let us look at where the problem lies. That text was written for the eye.

Think for a moment about writing. Writing exists to be read, which is to say, to be laid out before the eye. The eye, too, has a privilege, one we are most of the time not even aware of: the eye can go back. If you have not understood a sentence, your eye drifts on its own to the line before and reads it again. You halt at a tangled sentence, break it into its parts, press on it to ask which clause this one was bound to. You set your own pace. When you tire you stop, when your attention wanders you take it back. You move across the text as though it were your own property.

The ear is not like this. The ear is condemned to time. Sound flows once and is gone. There is no going back. You cannot summon again the place you missed, and the stream will not wait for you. The sentence you did not grasp is already gone while you are still trying to grasp it, and another has come down on top of it. The ear must take what it is given in that moment, at that speed, in that order. It has no other chance.

This is where the crack opens. A text written for the eye is built in reliance on the eye's privileges. You can write a long sentence for the eye, because the eye divides it. You can wedge a parenthesis into the middle, because the eye steps out and comes back in. You can build a sentence whose end ties back to its beginning, because the eye returns to the start. You can line up abstract nouns one after another, because the eye stops and thinks.

When you hand all of this to the ear as it stands, you have leaned on a faculty the ear does not have. The long sentence that comes easily to the eye falls apart in the ear. The parenthesis the eye slips in and out of with ease snaps the thread in the ear. The construction that loops to the beginning lands on nothing in the ear, because there is no beginning to return to. The machine reads that text flawlessly. But the text itself does not suit the ear. The flaw is not in the voice. The flaw is in the text.

Here I must head off a misunderstanding at once. Someone will say that carrying writing over into sound is nothing new. True. But things that look like the same work belong to wholly different worlds. To confuse them with ours is the greatest trap. So let me first say what we are not.

One: the audiobook. It reads a novel, a book, aloud from beginning to end. Once it was always a person who read; now sometimes a machine reads too. But the concern of the audiobook is not comprehension. Its concern is this: to deliver a book by other means to the person who has no time, no wish, no eyes free for reading. While driving, while walking, as the eyes close for sleep. This is an old world, its roots in the radio play, and further back, in the grandmother telling her tale. Within itself it is not one thing but a mixed world: some audiobooks are a plain reading, others almost a play, a full production, and the audiobook has its staged forms, like Storytel. But what they all share, and this is just what concerns us, is that the audiobook does not change the text. It stays faithful to what is written. It reads the novel word for word. So the audiobook does not even carry our concern, the concern of fitting a text for the eye to the ear. Because that is not its concern. It carries the writing. We rewrite the writing. A different work.

Two: the podcast, and those machine conversations of recent times, the tools that take a text and turn it into a two-person exchange. These are further still. These do not even read the text. They take it and turn it into an altogether different work: a chat, an interview, a performance. The listener no longer receives the text but a dramatized version of it. A curtain comes between. This does not make comprehension easier. It waters the text down, setting a representation in place of the original.

Why did I bring up these two worlds? Not because they are kin, but on the contrary, because they are not. To draw the line. We are neither the audiobook nor the podcast. The audiobook stays faithful to the text and carries it, with no concern for comprehension. The podcast leaves the text behind and crosses into another work. To set these on the same scale as us is to mix apples and oranges. What we do is neither of the two. It is a third, separate thing. Now I can say what it is.

Recall the problem: a text for the eye does not fit the ear. The solution to this problem can be sought in two places. Either in the voicing, or in the text.

What does it mean to seek the solution in the voicing? To labor at correcting the machine's voice, its tone, the way it stresses a word. But this is not our work. How a word is pronounced, which syllable is drawn out, where the stress falls, all of this is the work of the voicing engine. Whether the machine reads a word flatly or throws the stress the wrong way, we cannot correct it word by word; it is left to the engine's own doing. Chasing after diction takes us nowhere, because diction is not our domain.

The solution lies in the other place, in the text. In an idea that is very simple yet changes everything: before you give the text to the machine, rewrite it for the ear.

Break the long sentence. Turn the construction that sends you backward into a sequence that flows forward. Place the information in the order the ear will expect, the familiar first, the new after. Pour the frozen, nominalized sentence back into verbs, because the ear follows action, holds on to the verb. Replace the word that will blur in the listening with a clear one from the start.

Notice this: the machine is still reading what is written. The ordinary case has not been disturbed. The machine is again voicing what is placed before it. The only thing that has changed is what is placed before it. We have written the text for the ear.

This is exactly where we step outside the ordinary case. While everyone else voices what is written, we rebuild what is written for the ear before it is voiced. We do not touch the voicing; we touch the text. This is the name of the work: writing for the ear. To take a text written for the eye and rewrite it as a text for the ear.

This is why we are not the audiobook, which does not change the text. We are not the podcast, which turns the text into another work. We rewrite the text as the same text, but as a text for the ear. The same content, the same language, a different mode.

One is tempted to call this translation. Not without reason. We carry something from one form into another, we keep one thing and change another, we make decisions as we go, we gain some things and lose others. Translation is exactly such a work.

But translation in the classical sense happens between languages. From Turkish into another tongue. Here there is nothing of the sort. The language is the same, the content is the same. The only thing that changes is the mode: from a text for the eye to a text for the ear.

This is why we can neither quite call it translation nor quite call it not. It is translation, because it carries over, it decides, it opens a question of fidelity, it holds loss and gain within it. It is not translation, because there are not two languages here, but two states of one text.

The truest thing is to state this tension as it stands. It is called translation, yet it is not. It is translation, yet it is not. Something different depending on where you stand. This blurriness is not a shortcoming but the nature of the work. Writing for the ear is a work that stands at the border of translation, resembling it but never fitting inside it. If we cannot quite name it, that is because the work is new and a thing of its own.

So far we have laid out what we are doing. Now let us come to how it is done. These are not rules, but a few handholds to keep hold of.

Let the verb carry. The ear follows action. Not "the bringing of the problem to a resolution," but "to solve the problem." Loosen the frozen, nominalized structure, let the sentence run on its verb.

Go from the familiar to the new. Let each sentence begin with the familiar thing the one before it left behind, and save the new thing for the end. This is how the ear builds its connection. If you do the reverse, if you put the new at the front, the ear is left with nothing to hold on to.

Break the sentence. The long sentence the eye takes in one breath falls apart in the ear. A short sentence, a flowing sentence.

Build the rhythm through punctuation. Here punctuation is not grammar but breath. The period a halt, the comma a breath, the ellipsis a fading out. Through punctuation you tell the machine where to stop and where to flow. This is our domain, because this is not diction but rhythm. The stress inside a word is the engine's work. But where the sentence stops and where it runs on, that is the work of the text, which is to say our work.

Change the ambiguous word at its source. The word the eye keeps apart but the ear will confuse, change it from the start. If a word stays unclear when the machine says it, reach for a plainer one, choose the clear word that fits the context. If it is not obvious whether you mean bare or bear, build the sentence so that no ambiguity can arise. Do not hand the problem off to the listener. Solve it at the source.

Listing these does not finish the work. Nor will it. Because these are not a list of rules but a beginning.

The real teacher is the material itself. Every text, every sentence, every language trains the hand anew. Think of a sculptor. Working in marble, a master. Crossing over to bronze, no longer the same master, an apprentice once more. Because marble is the art of subtraction: you take away the excess, you bring out the form already inside the stone. Bronze is the art of casting: you fill the void, you raise the form from nothing. The same hand works in the opposite direction.

The passage from writing to the ear is just like this. Writing resembles the art of subtraction: the eye scans the excess, discards it, goes back. Writing for the ear resembles casting: you fill the flow, you carry it forward with no return. You learn the touch all over again.

This is why writing for the ear does not close with a book of rules. The rules are only a threshold. Where you cross that threshold, the craft begins. And craft is learned only by doing, by working at the material again and again. What lies before us is not a closed subject. A door left ajar.

Version 1.0 · First published: 28 June 2026

Permanent archival record, all versions: 10.17613/xt1cr-nk646 · Knowledge Commons