From Text to Voice

A text between two worlds – the eye reads, the ear listens, and something is lost in between.

April 2026 · 45 academic sources · 12 research files

A text is written. It sits on the page. The eye scans, backtracks, pauses, rereads. You can start from the middle of a paragraph, jump from the end to the beginning, linger on a single sentence three times and move on. That is what reading is – time belongs to you.

Then someone converts that text into sound. Time is no longer yours. The voice flows, cannot be stopped, cannot easily be rewound. Once a sentence passes, it has passed. If it was not understood, the loss is silent: the listener does not even realize they misunderstood. Nicole Ayasse and colleagues showed this in 2021: when a complex sentence is heard, the brain switches from full analysis to sampling words and guessing at meaning.Ayasse 2021 The error happens without notice.

This report tells the story of that crossing – what written text loses when it enters the audio channel, what it gains, and where synthetic voice stands in the equation. But it needs to begin with a warning: most of the data here was produced for the "average listener." At the end of the report I will turn this limitation on itself – because listening is anything but an average experience.

Two forms of the same language

Speech and writing look like two channels of the same language, but Wallace Chafe showed in 1982 that this is an illusion. There are two fundamental differences. The first is fragmentation and integration. A speaker can convey only one new concept per intonation unit – cognitive capacity imposes the limit. The result is short units strung together with "and... and... and..." A writer is free from this constraint: embedded clauses, nested structures, multiple concepts in a single sentence. The second difference is involvement and detachment: the speaker sees the listener, adjusts in real time, carries meaning through gesture and expression; the writer does not know the reader, and the text must stand on its own.Chafe 1982

So far this looks straightforward – speech is fragmented and involved, writing is integrated and detached. But M.A.K. Halliday made a striking correction in 1985: speech and writing are complex in different ways. Writing carries high lexical density – three to six content words per clause on average, compared to one and a half to two in speech. Speech carries high grammatical intricacy – clauses link to each other in unexpected ways. The common assumption that "writing is more complex than speech" is widespread but wrong: the two are complex along different axes.Halliday 1985

Douglas Biber took this a step further in 1988, analyzing sixty-seven linguistic features through multi-dimensional analysis. The most discriminating dimension is "involved versus informational": a phone conversation scores plus thirty-five, an academic paper scores minus fifteen. But no single dimension cleanly separates speech from writing – the relationship is not binary but a spectrum. A prepared lecture approaches written text; a personal letter approaches speech. Elinor Ochs had shown in 1979 that the determining factor is not the channel but the degree of planning: an impromptu message is unplanned, a conference talk is planned – both are "speech" but they do not occupy the same space.Biber 1988, Ochs 1979

This spectrum view also softens Walter Ong's 1982 "Great Divide" thesis. Ong had identified nine cognitive features of primary orality cultures – additive, aggregative, redundant, conservative. But Scribner and Cole, studying the Vai people in 1981, found that the cognitive effects of literacy were far smaller than expected. The distance between speech and writing is real, but the size of the distance changes with culture, context and the individual.Ong 1982, Scribner & Cole 1981

Not three but four

The traditional distinction recognizes three categories: natural speech, writing and reading aloud. In natural speech, producer and receiver face each other, interaction is immediate, the body is present. In writing, they occupy different times and places, there is no interaction, no body. In reading aloud, a human or a machine voices the text, the listener receives it, but interaction is still absent.

Today there is a fourth category, and it is the real subject of this report: text adapted for voice. A human writes, a machine speaks, a human listens – but that human is usually doing something else while listening: walking, washing dishes, driving. No interaction, no body, and on top of that the listener's attention is divided. This fourth category is neither speech nor writing nor reading aloud – it sits at the intersection of the three but fits into none of them.

The practical consequence is this: some of Chafe's, Halliday's and Biber's findings apply directly – lexical density differences, the concept of autonomous text, the need for information structuring. But dimensions like involvement, real-time adjustment and body language do not apply. The right question is not "speech or writing" but: how do we reconstruct a written text so that it reaches a wandering mind through a synthetic voice?

When text enters the audio channel

Donald Rubin systematically defined the concept of "listenability" in 2000 and 2012. Readability formulas – Flesch-Kincaid, Gunning Fog – were designed for written text and do not apply to listening text. Rubin's experiment is clear: when the same content was prepared in two different language styles – one oral-based (active voice, personal pronouns, sequential sentences, concrete examples), the other literate-based (passive voice, abstract concepts, embedded clauses) – the oral-based text was understood significantly better. Processing written-mode language while listening is especially costly.Rubin 2000, 2012

The most concrete finding at the sentence level came from Kadayat and Eika: twenty-one participants listened to five texts of varying length through a screen reader, and the highest comprehension with the lowest cognitive load was achieved at sixteen to twenty words per sentence. Sentences exceeding twenty words reduced comprehension and increased load. Below ten words, context fragmented – sentences that were too short also caused problems.Kadayat & Eika 2020 But this "golden range" holds for the average listener; individual differences are large. I will return to this at the end of the report.

Sentence structure also matters. Constable and colleagues showed in 2004 with fMRI that object-relative clauses ("the woman who the man loved") required far higher brain activation than subject-relative clauses ("the man who loved the woman") – left frontal lobe, temporal region and angular gyrus all worked harder.Constable 2004 Peelle and colleagues added in 2010 that speech rate and syntactic complexity can each be tolerated individually but together they overwhelm capacity. Simplifying syntax is the primary intervention.Peelle 2010

Information ordering engages a real-time mechanism in the brain. Kaiser and Trueswell showed in 2004 with eye tracking that listeners begin searching for a new referent before the word is even heard. When given information comes first and new information comes last, this predictive mechanism works; reversed ordering breaks it.Kaiser & Trueswell 2004 This is a default setting in the brain – it allows for individual flexibility, but the general principle is robust.

The transformation of visual elements is a separate issue. Zong and colleagues worked with thirteen blind and low-vision participants in 2022 and found that a layered, navigable structure is preferred over raw table data. The order they sought turns out to be the one sighted readers use as well: first an overview that gives the whole, then the descent to individual values. The same study warns against imposing a fixed order: which arrangement serves depends on what the listener is trying to do at that moment. A table is silent information; voice must convert it into sequential narration.Zong 2022

Headings serve a different function in audio. Lemarié and colleagues defined seven functions of structural markers in text in 2012; a heading carries several of them at once, such as introducing the topic and announcing that a section has begun. But the real lesson comes from the same team's study of speech: Lorch, Chen and Lemarié delivered the same text in two forms, once printed and once read by a synthetic voice, and the headings that did their job on paper collapsed in listening. Most of what a visual heading carries lies not in its words but in its type size, the white space around it, its weight, and the speech software conveys none of that. Preview sentences, by contrast, worked in both forms. The rule follows: in audio a heading is not enough, it has to be put into words. Likewise, announcing a heading and signaling a transition are not the same thing. "Now we move to the second section" is formulaic and empty; "this brings us to a different question" is organic and connected. Topic transitions and new-information signals are markers that recall attention.Lemarié 2012, Lorch 2012

The conversion of numbers is equally unavoidable. Seven decades of radio journalism make this clear: "1,987,452" becomes "roughly two million" in audio. As Jonathan Kern wrote in his NPR editorial guide, "print and audio live in the same world but the terrain is different." The radio version of a print story is not about turning quotes into sound bites – it is a thorough reconstruction.Kern 2008

Human voice, machine voice

Does listening through synthetic voice carry an additional cost? Govender and King measured this between 2018 and 2019 using pupillometry and dual-task paradigms: in all conditions, synthetic speech generated higher cognitive load than natural speech. In noisy environments the gap widened further – from low-quality HMM to high-quality DNN, every synthesizer fell behind natural voice.Govender & King 2018a, 2018b, 2019

But most of this data predates modern neural TTS. Ibelings and colleagues tested a DNN-based synthesizer on everyday sentences in 2022: the difference was 1.2 decibels in speech reception threshold – within the same range as the variation between different natural speakers. In verbal response times, a measure of listening effort, no significant difference was found.Ibelings 2022 Craig and Schroeder found the same pattern in learning outcomes between 2017 and 2019: a modern TTS engine achieved reliability and perceived learning statistically indistinguishable from human voice in transfer learning tasks.Craig & Schroeder 2017, 2019

Contradictory findings also exist – a 2024 eye-tracking study with thirty students found that human recordings produced significantly higher learning performance and attention engagement than synthesized voices.Jing 2024 This conflicts with Craig and Schroeder. The difference likely lies in context: for short information transfer, synthetic voice is sufficient; in long-form learning, prosodic advantage becomes more pronounced.

Prosody is indeed a critical variable. Pagnotta and colleagues showed in 2025 that short descriptions listened to with positive prosody yielded more correct answers, and human voice produced higher accuracy than TTS. But a 2025 study published in Cogent Psychology added a different dimension: in immediate recall, natural human voice produced better memory retention than all other conditions; one week later, the difference between conditions disappeared. The prosodic advantage fades when spread over time.Pagnotta 2025

Virginia Clinton-Lisell's 2022 meta-analysis draws the big picture: forty-six studies, four thousand six hundred eighty-seven participants. Overall comprehension difference between reading and listening: g = 0.07 – statistically nonsignificant. But moderators matter: self-paced reading has an advantage (g = 0.13), inferential comprehension strongly favors reading (g = 0.36), lexical comprehension shows no difference (g = -0.01). And a critical finding: in languages with transparent orthographies – Turkish is close to this – the reading-listening gap approaches zero (g = 0.001).Clinton-Lisell 2022

Yang and colleagues showed in 2026 that neural-level differentiation emerges after a brief training session of about twelve minutes – the brain synchronizes rapidly with synthetic voice, but conscious detection is far more difficult. Segedin and colleagues had shown in 2019 that listeners apply the same phonetic adaptation mechanisms used for human speakers to TTS speech as well. The ear adapts – whether the source is human or machine.Yang 2026, Segedin 2019

When imitation ends

For twenty years, synthetic voice tried to imitate the human voice. Now a different question is being asked: does it have to?

Diel and Lewis showed in 2024 that synthetic voices have successfully escaped the uncanny valley. What triggers the uncanny valley is not being synthetic but deviating from the typical human voice. Pathological and distorted voices produced more discomfort than purely synthetic ones. This suggests that synthetic voice is already accepted as its own category – the listener positions it not as "bad human voice" but as "something different."Diel & Lewis 2024

Nussbaum, Fruhholz and Schweinberger went further in 2025, writing in Trends in Cognitive Sciences that the concept of "voice naturalness" is scientifically undefined. There are two types of naturalness: deviation-based (how far it strays from human voice) and human-likeness-based (how closely it resembles a human). This ambiguity opens the possibility that synthetic voice may define its own naturalness.Nussbaum 2025

Im and colleagues found the opposite of expectations in 2023: in information-based tasks, users preferred synthetic voice. In functional tasks, synthetic voice was perceived as more fluent, received higher competence ratings and generated positive attitudes. Torre and colleagues had identified the "congruency effect" in 2018: if a voice is perceived as trustworthy and subsequent experience is consistent, trust is reinforced. The unchanging tone of synthetic voice facilitates this consistency – the same quality, the same rhythm, the same reliability every time.Im 2023, Torre 2018

Le Maguer and Cowan put the provocation in the title at ACM in 2021: "Synthesizing a Human-Like Voice Is the Easy Way." The hard but right path is for synthetic voice to find its own way. The "Cabinet of Voices" study at DIS 2025 is the most concrete example: the concept of "more-than-human vocalities." Participants imagined synthetic voices producing sounds like a whale – not imitation but exploration.Le Maguer & Cowan 2021, Cabinet of Voices 2025

Lamas gave this transformation a name in 2025, writing in Explorations in Media Ecology: "synthetic orality." Synthetic voice is neither primary orality (it has no body), nor secondary orality (it is not like radio), nor writing – it is a fourth category. From McLuhan's perspective, every medium creates its own language: the printing press created the novel, television created the commercial. Synthetic voice is creating its own genre – NotebookLM's AI podcast is an early example.Lamas 2025

But this paradigm shift is not yet complete. Much is missing: there is no systematic "synthetic voice paradigm" thesis – the study that connects the pieces has not yet been written. There are no dedicated metrics for synthetic voice – what to measure instead of "human-likeness" remains unclear. Long-term listener adaptation studies are insufficient. And the vast majority of existing research is English-language and Western-centric – cultural diversity is lacking.

The particular case of Turkish

Turkish is an agglutinative language – multiple suffixes attach to a single root, and word forms are theoretically infinite. This creates two concrete problems for synthetic voice. The first is morphological: standard vocabulary-based approaches fall short because a form like "gelememislerdensmis" (roughly: "apparently they were among those who could not come") cannot be pre-loaded into a lookup table. Kosaner and colleagues developed a phoneme-based rule-driven TTS system for Turkish in 2024, demonstrating that morphological analysis is essential for correct pronunciation.Kösaner 2024 Liu and colleagues found in 2024 that morphology-aware language model pre-training improves quality in low-resource agglutinative languages – an approach directly applicable to Turkish.Liu 2024

The second problem is ambiguity: the same spelling can yield different readings. Hakkani-Tür and colleagues showed in 2002 that Turkish morphological ambiguity can be resolved by statistical methods; Külekçi applied this disambiguation to pronunciation in 2006 and showed that choosing the correct pronunciation is critical for Turkish synthetic speech quality. In text adaptation, this means identifying and resolving structures that could produce ambiguous pronunciation.

But Turkish also has an advantage. In Clinton-Lisell's meta-analysis, languages with transparent orthographies – languages that are read as they are written – show a comprehension gap between reading and listening that approaches zero. Turkish is not perfectly transparent, but it is close: Latin alphabet, largely phonemic spelling. This suggests that text adaptation in Turkish may produce less loss than in English.Clinton-Lisell 2022

Where the mind wanders

How far the mind wanders depends on where it is measured. Risko and colleagues found forty-three percent in 2012 among people watching a lecture video alone in the laboratory; when the same team measured in a classroom, the rate fell to thirty-nine percent. What matters more is that the number is not fixed: thirty-five percent in the first half of a lecture, fifty-two in the second. The longer the listening lasts, the more attention leaks away.Risko 2012 The audio-only condition produces the highest mind wandering, Kopp and D'Mello demonstrated this in 2016. The mechanism is straightforward: processing audio demands relatively few resources, the remaining capacity is freed, and attention drifts.

But this is not entirely a bad thing. Free attention can build different connections, can see angles invisible on the page. The problem is not the wandering itself but the inability to return. Faber and colleagues' 2018 finding becomes relevant here: when no event boundary occurs – when the narrative flows flat and unchanged – wandering becomes excessive. The solution is not to interrupt the flow but to embed recall markers within it: topic transitions, new-information signals, an unexpected metaphor. The listener's mind will wander – the text must be able to call it back.

This report's biggest blind spot

Everything I have described so far rests on an assumption: that there is such a thing as an "average listener." Kadayat and Eika have their twenty-one participants, Clinton-Lisell has her forty-six studies, Govender has his pupillometry data – in all of them, means are computed, confidence intervals calculated, general principles extracted. This is what the scientific method requires and it is not wrong. But it is incomplete.

Incomplete because listening is a profoundly individual experience. Two people hear the same sentence: one grasps it, the other loses it – and the one who loses it does not notice. Kadayat and Eika's "golden range of sixteen to twenty words" is the average of twenty-one people. But among those twenty-one, there were probably some who maintained coherence at the fortieth word and others who lost it at the tenth. The average represents neither.

This is not only a matter of sentence length. Individual variance is large across every dimension of cognitive profile: abstraction capacity, information-density tolerance, flow preferences, aesthetic standards, sensitivity to voice. One listener needs pedagogical bridges ("let me explain this with an example"), another perceives the same bridge as filler and loses attention. One listener is reassured by formulaic transition signals ("we now move to part two"), another finds them irritating. For one listener the consistency of synthetic voice is an advantage – the same quality every time; for another it is monotony.

The dates of the academic research are also a limitation. A significant portion of the sources in this report predate 2020 – synthetic voice technology has changed fundamentally since then. The "high-quality DNN synthesizer" of 2018 is not even baseline level in 2026. Govender's cognitive load findings should be retested with current engines. Craig and Schroeder's "no difference" finding might become "synthetic voice advantage" with today's models. The data is still valuable, but questioning its currency is mandatory.

In short: this report is grounded in scientific data, but that data paints a picture tailored to no one. Think of a map – it shows the general shape of the country but not your front door. To find your house you need to go down to street level. In the listening experience, street level is personalization.

Street level on the map

Personalization is not an abstract concept – it needs to show concretely what changes. I have a case at hand: thirty thousand lines of conversation records analyzed, a profile extracted across six cognitive-linguistic dimensions, this report's seven general principles calibrated to that profile. Below I share some of the calibration results in anonymized form – to make the concrete difference visible.

Sentence length: This report's general principle says sixteen to twenty words. But the profile data shows that this listener's natural processing range is twenty-five to forty-five words – across thirty thousand lines of recorded conversation, the phrase "I don't understand" or "simplify" appears zero times. When the golden range is applied to this listener, thought fragments: a structure that carries context in a single sentence, split into two, loses its connection. The general principle is correct but harmful for this listener.

Abstraction capacity: The general principle says "support abstract concepts with concrete examples." This listener does not merely understand abstract concepts – they generate them, naturally using self-coined abstractions in conversation. Pedagogical bridges ("let me illustrate this with a simple example") are filler for this listener: if a concept needs explanation, explain it directly, do not announce that you are about to explain.

Information density: The general principle says "thin out information, one concept per unit." Across one hundred sixty files from this listener, there is not a single sign that information density was excessive. Complaints always point to insufficiency, to shallowness. When information is diluted for this listener, loss occurs – it is not the density that drops but the meaning.

Transition preference: The general principle says "use transition signals: we now move to the second topic." This listener's thought pattern follows a spiral deepening – topics emerge from one another, and formulaic boundary markers interrupt the flow. Instead of "let us now move to part two," what is needed are transitions that arise organically from the last thought of the previous section.

Synthetic voice preference: General research says synthetic voice increases cognitive load. This listener prefers consistency over peak performance – even when listening to music, the choice is source fidelity over effects. The infinite consistency of synthetic voice is not a disadvantage for this listener but an active advantage.

Five examples, five different deviations. Every general principle is correct – but all five are wrongly applied to this listener. The map is right, the street is wrong. And this is a single listener – every listener's deviation pattern will be different.

How artificial intelligence does this

Could a human perform the calibration described above? Yes – but it does not scale. Reading thirty thousand lines of conversation records, extracting a profile across six dimensions, calibrating seven principles, deriving five new ones – this would require weeks of work from a linguist, and every new listener would mean starting over. Artificial intelligence does it differently: large language models are a natural infrastructure for text analysis and can automate every stage of personalization.

The process works in roughly four stages. First stage, profile extraction: the listener's conversation or correspondence records are analyzed. Sentence lengths, word choices, abstraction levels, topic-transition patterns, satisfaction and dissatisfaction markers – all are extracted. This is a classic text-mining task, but large language models can go beyond surface-level statistics because they understand context: the absence of "I don't understand" and the absence of "simplify" are not the same thing – the model grasps this difference.

Second stage, calibration: academic principles are compared against the profile. For each principle, the question is: can it be applied to this listener, can it not, and how should it be modified? The model takes the academic finding as a starting point for the "average" and computes deviations using profile data. Kadayat and Eika's sixteen-to-twenty-word range starts as the default, but if the profile shows twenty-five to forty-five words, the range is updated.

Third stage, text adaptation: the written text is rewritten according to the calibrated principles. Sentence structure, information ordering, transition signals, heading formats, metaphor density, information density – all are adjusted to fit the profile. This stage is where large language models are strongest: text generation and transformation. The same source text can be converted into different forms for different profiles – one with short sentences and pedagogical bridges, another with long flowing sentences and organic transitions.

Fourth stage, continuous updating: the profile is not static. The listener's preferences, cognitive capacity and interests change over time. Each new interaction updates the profile. The model compares previous calibration results with new data and corrects deviations. This differs from traditional A/B testing – it is not a choice between two options but continuous fine-tuning.

Technically, this process requires several components: a large-context-window language model (tens of thousands of lines must be read for profile extraction), structured output generation (profile dimensions, calibration tables), long-form text generation (the adapted text) and persistent memory (profile updates). Current models can do all of this – the limitation is not in the technology but in the implementation. The vast majority of text-adaptation platforms still rely on single-dimensional metrics like "readability level"; personalized calibration is not yet standard.

There is also an ethical dimension. Analyzing thirty thousand lines of conversation records produces a deep cognitive profile – from sentence-processing capacity to aesthetic standards, from abstraction level to emotional response patterns. In the wrong hands, this information can become a tool for manipulation. The line between personalization and manipulation lies not in intent but in transparency: the listener should know what their profile is, how it is used and which decisions it influences. Raghavan and Schneier argued in IEEE Spectrum in 2025 for deliberately robotic voicing of synthetic speech – the idea that synthetic voice carrying its own identity is not a devaluation but a function. The same logic applies to personalization: the fact that "this was customized for you" should be visible, an open service rather than a hidden adjustment.Raghavan & Schneier 2025

This report told a story of translation – translation within the same language. When written text enters the audio channel, sentence structure changes, information ordering changes, tables become narration, headings become signals, numbers get rounded. Synthetic voice adds its own dimension to this transformation: cognitive load rises but closes the gap, prosody is absent but consistency becomes an advantage, imitation ends but its own paradigm has not yet fully emerged.

And the point where the report questions itself: all of these findings are for the "average listener." But listening is anything but an average experience. Personalization is not a luxury, it is a condition of accuracy – without calibrating general principles to the individual profile, any claim of "correct adaptation" remains incomplete. Artificial intelligence makes this calibration scalable: profile extraction, principle calibration, text adaptation, continuous updating. The technology is ready – the implementation is not yet standard.

Converting text to speech is an act of translation – within the same language.

Bibliography

Ayasse, N.D.; Hodson, A.J.; Wingfield, A. (2021). "Principle of Least Effort and Sentence Comprehension." Frontiers in Psychology. PMID: 33796047

Biber, D. (1988). Variation across Speech and Writing. Cambridge University Press.

Cambre, J. et al. (2020). "Choice of Voices." CHI 2020.

Chafe, W. (1982, 1985). "Integration and Involvement in Speaking, Writing, and Oral Literature." Spoken and Written Language.

Clinton-Lisell, V. (2022). "Listening Ears or Reading Eyes: A Meta-Analysis." Review of Educational Research.

Constable, R.T. et al. (2004). "Sentence Complexity and Input Modality Effects: fMRI." NeuroImage. PMID: 15109993

Craig, S.D.; Schroeder, N.L. (2017). "Reconsidering the Voice Effect." Computers & Education, 114.

Craig, S.D.; Schroeder, N.L. (2019). "Text-to-Speech Software and Learning." JECR, 57(6).

Diel, A.; Lewis, M. (2024). "Deviation from typical organic voices best explains a vocal uncanny valley." Computers in Human Behavior Reports, 14, 100430.

Dylman, A.S.; Glarén Diaz, D.; Blysa, A.; Jansson, B. (2025). "The effect of prosody on listening comprehension: Immediate and delayed recall." Cogent Psychology, 12(1), 2576785.

Faber, M. et al. (2018). "Driven to Distraction: A Lack of Change Gives Rise to Mind Wandering." Cognition, 173.

Govender, A.; King, S. (2018a). "Using Pupillometry to Measure the Cognitive Load of Synthetic Speech." Interspeech 2018, 2838-2842.

Govender, A.; King, S. (2018b). "Measuring the Cognitive Load of Synthetic Speech Using a Dual Task Paradigm." Interspeech 2018, 2843-2847.

Govender, A.; Wagner, A.E.; King, S. (2019). "Using Pupil Dilation to Measure Cognitive Load When Listening to Text-to-Speech in Quiet and in Noise." Interspeech 2019, 1551-1555.

Hakkani-Tür, D.Z.; Oflazer, K.; Tür, G. (2002). "Statistical Morphological Disambiguation for Agglutinative Languages." Computers and the Humanities, 36(4), 381-410.

Halliday, M.A.K. (1985). Spoken and Written Language. Deakin University Press.

Ibelings, S.; Brand, T.; Holube, I. (2022). "Speech Recognition and Listening Effort of Meaningful Sentences Using Synthetic Speech." Trends in Hearing 26. PMC9549212.

Im, H.; Sung, B.; Lee, G.; Kok, K.Q.X. (2023). "Let Voice Assistants Sound Like a Machine." Computers in Human Behavior.

Jing, B.; Wu, C.; Pi, Z.; Zhou, Y.; Zhang, Y.; Liu, H. (2024). "Cute Computer-Synthesized Voice Hinders Learning Performance in Instructional Videos." The Journal of Experimental Education (online first).

Kadayat, B.B.; Eika, E. (2020). "Impact of Sentence Length on the Readability of Web for Screen Reader Users." HCII 2020, Lecture Notes in Computer Science, 261-271.

Kaiser, E.; Trueswell, J.C. (2004). "Role of Discourse Context in Processing Flexible Word-Order Language." Cognition. PMID: 15582623

Kern, J. (2008). Sound Reporting. University of Chicago Press.

Kopp, K.; D'Mello, S. (2016). "The Impact of Modality on Mind Wandering during Comprehension." Applied Cognitive Psychology, 30(1), 29-40.

Koşaner, Ö.; Özgen, M.; Birant, Ç.C. (2024). "Developing a Rule-Based Text-to-Speech System for Turkish." International Journal of Humanities Social Science and Management, 4(4), 109-121.

Külekçi, M.O. (2006). Statistical morphological disambiguation with application to disambiguation of pronunciations in Turkish. PhD thesis, Sabancı University.

Lamas, N.J. (2025). "AI, Language Processing and Synthetic Orality." Explorations in Media Ecology, 24(3).

Le Maguer, S.; Cowan, B.R. (2021). "Synthesizing a Human-Like Voice Is the Easy Way." ACM CUI 2021.

Lemarié, J.; Lorch, R.F., Jr.; Péry-Woodley, M.-P. (2012). "Understanding How Headings Influence Text Processing." Discours, 10.

Liu, R. et al. (2024). "TTS for Low-Resource Agglutinative Language with Morphology-Aware Pre-Training." IEEE.

Lorch, R.F., Jr.; Chen, H.-T.; Lemarié, J. (2012). "Communicating headings and preview sentences in text and speech." Journal of Experimental Psychology: Applied, 18(3), 265-276.

Nussbaum, C.; Frühholz, S.; Schweinberger, S.R. (2025). "Understanding Voice Naturalness." Trends in Cognitive Sciences, 29(5).

Ochs, E. (1979). "Planned and Unplanned Discourse." T. Givón (Ed.), Syntax and Semantics, Vol. 12: Discourse and Syntax in (51-80). Academic Press.

Ong, W.J. (1982). Orality and Literacy. Methuen.

Pagnotta, M. et al. (2025). "Prosody! When intonation helps." Educational Psychology.

Peelle, J.E. et al. (2010). "Neural Processing During Older Adults' Comprehension." Cerebral Cortex, 20(4).

Raghavan, B.; Schneier, B. (2025). "AI Voices Should Sound Robotic Again." IEEE Spectrum.

Risko, E.F.; Anderson, N.; Sarwal, A.; Engelhardt, M.; Kingstone, A. (2012). "Everyday Attention: Variation in Mind Wandering and Memory in a Lecture." Applied Cognitive Psychology, 26(2), 234-242.

Rogowsky, B.A.; Calhoun, B.M.; Tallal, P. (2016). "Does Modality Matter?" SAGE Open.

Rubin, D.L. (2012). "Listenability as a Tool for Advancing Health Literacy." J. Health Communication, 17(S3). PMID: 23030569

Rubin, D.L.; Hafer, T.; Arata, K. (2000). "Reading and Listening to Oral-Based vs Literate-Based Discourse." Communication Education, 49(2).

Scribner, S.; Cole, M. (1981). The Psychology of Literacy. Harvard University Press.

Segedin, B.F.; Cohn, M.; Zellou, G. (2019). "Perceptual Adaptation to Device and Human Voices." Interspeech 2019.

Torre, I.; Goslin, J.; White, L.; Zanatto, D. (2018). "Trust in Artificial Voices: A 'Congruency Effect' of First Impressions and Behavioural Experience." Technology, Mind, and Society (TechMindSociety '18).

Yang, J. et al. (2026). "Short-Term Perceptual Training Modulates Neural Responses to Deepfake Speech." eNeuro.

Zong, J.; Lee, C.; Lundgard, A.; Jang, J.; Hajas, D.; Satyanarayan, A. (2022). "Rich Screen Reader Experiences for Accessible Data Visualization." Computer Graphics Forum, 41(3), 15-27.

OTHER WRITINGS · The Invisible Gap

Hallucination: Whose Problem? · On Hatred · Research Constitution · Exactly · Tea Table · Picture of Happiness · Sturgeon's Law · Understanding

This text was originally written in Turkish. We tried to make it as natural as possible in English – staying true to the voice rather than the words. If you spot something that could sound better, write to [email protected] – you'd be enriching our translation.

Transcreation Notes

"backtracks" – geri doner
The Turkish geri donmek (to return, to go back) is the most ordinary verb imaginable – it carries no literary weight. The earlier draft used "goes back," which is equally flat. We chose "backtracks" because it implies retracing steps with purpose, not just retreating – it better captures the deliberate, investigative quality of the eye scanning a page. The ucgen review flagged the original as too passive.
"linger on" – ciğnemek (chew on)
The Turkish original uses cignemek – literally "to chew." It is a physical metaphor: the mind chews on a sentence the way a mouth chews food, extracting meaning through repetition. English "chew on" exists but sounds colloquial in an academic-adjacent text. We settled on "linger on" – which preserves the temporal quality (staying with something longer than expected) but sacrifices the bodily, almost gustatory image. The mouth became the mind.
"a text between two worlds" – iki dunya arasinda bir metin
The Turkish subtitle iki dunya arasinda is a common phrase that usually means "caught between two worlds" with a sense of belonging to neither. Applied to a text about eye-reading versus ear-listening, it implies the text itself is homeless – not fully at home in either medium. English "between two worlds" works but lacks the arada kalmak (being stuck in between) undertone. We accepted the loss because adding "caught" or "stranded" would have over-dramatised a subtitle.
"synthetic orality" – sentetik sozluluk
Lamas coined "synthetic orality" in English; the Turkish translation sentetik sozluluk is our coinage. Sozluluk does not exist as a standard Turkish word – we derived it from soz (word/speech) + -luk (the quality of being). This is a case where the English came first and the Turkish transcreation had to invent a term. Reversing into English was straightforward; the creative debt ran the other direction.
"the ear has no such luxury" – kulak icin bu luks yok
The Turkish luks (luxury) when applied to a sensory organ sounds sharper than the English equivalent – it implies that going back and re-reading is not just convenient but extravagant, a privilege the ear is denied by physics. English "luxury" can sound hyperbolic in this context; Turkish readers hear it as precise. We kept it because no other word – "privilege," "option," "freedom" – carried the same sense of something valuable that one channel possesses and the other simply does not.