Breath moves cursor.
Thought arrives aloud.
Paragraphs become landscape.
The voice makes clay.
Draft zero, still warm.
I compose forwards.
A Way of Composing
I was telling someone the other day about just how completely dictation and speech-to-text software has taken over my life. How it’s transformed not only my relationship with writing, but also with my computer.
One of the things about having a blog and writing regularly on it is that sometimes, you’ll mention something in passing, then it’ll come up again and again. The same is true of keeping a journal, it means you can follow the thread, a thought, or a concept across years of your life.
Using dictation and speech-to-text software, apparently, is one of those threads.
I first mentioned using speech-to-text dictation on here way back in 2018, in the sixth-ever episode of Permanently Moved in fact:
The other thing I wanted to talk about this week is the dictation feature in Google Docs. Now that my microphone is permanently set up at my desk, I’ve kind of gotten into the habit of using it. It’s probably not news to anyone who already uses the feature a lot, but I’ve realized that I can just sit down, start speaking, and suddenly I’ve got 500 to 750 words on the page. I mean, it still needs a lot of editing—because, well, it’s dictation—but it’s fairly accurate. More importantly, the words are on the page, and I didn’t have to sit down and type them. There’s no inner critic. You can just get words out of your head and onto the page.
Funnily enough, I have another episode of Permanently Moved named Abracadabra in the works! Probably the first episode of 2027, haha.
Anyways.
Back in 2018, I spent an intense period trying to use dictation tools as much as I could because I knew it would be valuable to me. But at the time I got extremely frustrated with how bad the hit rate was on the word recognition. Plus Google Docs voice dictation, to this day, still doesn’t like it when you pause to think between thoughts. It just turns off, which was/is deeply frustrating. So I abandoned that mode of human-computer interaction for a while.
I started using it again after the Pixel Recorder app came out. I remember clearly that the first episode of Permanently Moved I made using speech-to-text for the draft was Episode 2318 in 2023. Clear memories of walking down the road speaking to myself feeling weird and self conscious about it. The proof of concept was achieved though, I could record a voice note, get a rough transcription from the app, clean it up with an AI prompt, and start my edit. It wa sa revelation.
A year on, I mentioned voice transcription in my review of Day One, the journaling software from Automattic, in April 2024:
I’ve been experimenting with recording myself whilst i’m out on my daily walk. “Take a memo” style. With the advances in AI, speech to text has reached the point where its *good enough*. And then if you run the good enough transcription though GPTChat you can get something *quite usable*. Many of my podcasts, and longer bits of writing have begun life using this workflow in the last few months.
Going out for a 20min walk and coming home with 2k+ words of notes in a format that only needs a light edit to be useable is a real game changer
A few months later, it had become so embedded in my workflow that I noted it again in Episode 2420, documenting my use of AI tools at the time:
As I’ve become more comfortable with transcription and speech-to-text in general, the concept of speaking to my computer has opened up. I now use my M2 MacBook’s livetext tool all the time. I’ve even have it mapped it to my Globe function key. I’ll hit it and speak the rest of an email or paragraph I’m writing aloud. Despite having learned to put one word after another over the last 10 years, and making episodes about typing at the speed of thought, speech is just so much faster than fingers.
And I mentioned it again in the penultimate episode of 301, And it’s been over a year since then, and I honestly can’t imagine generating tokens draft text any other way now.
Whilst making Permanently Moved every week, keeping a diary, and writing a blog basically trained to touch-type at a fairly respectable speed, during the same period speech-to-text basically became a solved problem. Every new hardware device I’ve bought, and in some cases with incremental software updates, has brought better improvements to the accuracy of the transcription tools.
I currently have three main routes for using speech-to-text in my life:
- Pixel Recorder App: The primary one is on my Pixel 9 Pro, with transcripts exported to Google Docs. I then copy and paste these into whatever LLM window happens to be open, using a custom clean up prompt I have saved in my clipboard snippets in Alfred (let me know if you would like me to share it). I use this almost every day, sometimes multiple times a day.
- Fluid Voice: An open-source speech-to-text app that combines multiple models running locally to achieve some really impressive results. It’s MLX/Mac native, and based on Nvidia’s Parakeet transcription model with an additional tidy-up model which formats the transcription, adds full stops and paragraphs, etc. If you’ve spoken a list aloud mid stream, it will format it as numbered or bullet-point it for example. I use this multiple times a week.
- Built-in macOS Dictation: Also I still use the built-in macOS dictation button with some regularity, I’m using it right now. MacOS dictation is instant, and I don’t always have Fluid Voice and the LLM models loaded into memory. I use this almost every day, mostly to reply in text on group chats or Discord or whatever.
Lastly, and with more frequency I’ve been using voice interaction with Claude and ChattyGPT, particularly when I am vibe coding. The ability just to give conversational spoken instructions and or explain something to an AI and have it instantiate something in code in front of your eyes is literally abracadabra. I find it truly amazing.
Whilst I recognise that these AI voice inputs aren’t necessarily robust accessibility tools; natural language voice input, or speaking to my computer, is becoming a very natural UI interaction for me. I wonder if it’s because I’ve now done so much voice dictation that whatever self-conscious part of me that once existed about speaking to my computer like Scotty being cringe saving the whales, no longer exists. Besides, Star Trek: The Next Generation showed us the way so we might as well walk the road.
Interestingly, never, have I ever, felt the desire to ask Alexa, Siri, or ‘OK Google’ anything on my phone. This might change in future, but not right now. But I’ll pin this observation and note is a potential future thread.
Thinking in Paragraphs
The biggest shift I’ve noticed over the years is how speaking aloud reorganises thought.
When typing, I can cruise along at roughly the same speed as my thoughts, or at least slow my thoughts down to one word after another at the same speed as my fingers at my. Voice dictation however is fundamentally different, as long as you arn’t mumbling you can’ speak as fast as you want into the dictation machine. Self-consciousness aside, walking around town with wired IEMs, speaking aloud in fits and starts while passers-by assume I’m on a phone call, Is just something I do now. Maybe I’m the local crazy man? Either way the ideas that emerge during dictation feel distinct from the way they would if I were to have typed them out.
Initially, my dictation transcripts would be full of verbal meta-comments:
“Note to self: move this point above the previous one.”
I would realise, as I was speaking, that the point I was making I something I should have made earlier one. But the more I have practised with the tools, the more the aperture of thought seems to have widened, and if i have something to say, then the order that I’m going to say it is better arranged. I guess I would say that I have begun to think more and more in entire paragraphs rather than discreet units of thought one after another, or isolated sentences.
As you expound on a topic aloud, (whilst walking around like a crazy man) you can feel the destination of your the point you are making approaching. It’s a bit like the difference between looking at a mountain peak on a map, and actually walking the landscape. As you put one word/step after another aloud, the structure of the next section will come into view slowly as you reach it.
It also has another quality that’s hard to explain, but I can attempt to explain it by saying how it’s different from typing:
Speaking text aloud, in the knowledge that you are being recorded/transcribed, feels like it’s behind the cursor, your breath is pushing the words out, and forward, propelling the cursor from left to right.
Typing meanwhile, feels like thoughts are on the other side of the cursor. The words are already there in my head, somewhere out there in front of the cursor. Typing words out for me feels less propulsive. Like writing / typing is more like a scratch-off card situation where the words are already present somewhere beneath the surface of the subconscious and the screen. You just need various combinations of keypresses reveal them. Like you are reporting on what you are thinking, on something that is already there.
With dictation, it feels much more like you are uncovering what you think by speaking it, rather than discovering what you already think by typing it out?
Draft Zero
I’ve written about this before, but I think of dictation transcripts as “draft zero.” It’s all pre draft, and pre edit. Text is a raw material like clay. Anything generated from an LLM as a similar quality. Raw drafts of any kind, typed, dictated or AI output, are all extremely cheap. And so are the the ideas behind them. As ever, the costly and fun bit, is in the edit.
Dictation reduces the cost of producing the raw material known as ‘words’. Not as far as LLMs have, but never the less it massively reduces it. For some of the writing I am most proud of, I’ve spoken a piece three separate times over several days to generate the *right* raw material for a workable draft.
Another side effect from all this dictation over the last few years I’ve noticed is that it has also changed the way I speak in my day-to-day life. Whilst I’ve always had the tendecy to monologue stamp on my ticket to the A-train, I’ve notice recently that my thoughts now tend to emerge in a more structured way (thinking in paragraphs again). Conversations, particularly technical or professional ones, have become less like a tennis rally back and forth, and more like a speech generator.
Like I’m Chat-JayTP answering questions/prompts, an off-the-cuff unspooling of thought until the next turn is to be taken.
It’s a skill I long admired in Gordon White. We spoke about it once, and I said that I thought his multi-hour live Rune Soup member Q&As in the late 2010’s served as a bootcamp for developing his long, uninterrupted, cogent, and off the cuff rhetorical style.
The last thing I want to note about the consequences of using speach-to-text software is the tone of the text that you produce. Dictation has affected my prose style.
One of the results of having done an undergraduate degree in philosophy when out in the world of work, was the beatings I received in 121’s with my manager in my first office job about overwritten emails. So since then I’ve always tried to write clearly.
So as a a writer, as an adult, my target register has for a long while been to try to hit a clear, accessible tone in a register that a smart teenager or undergrad would grasp. And speech turns out to be an excellent source for that register. Whilst my most recent podcast episode, Internet Operator (see below), is probably the “wordiest” thing that I’ve ever put out on the Internet, I nevertheless tried really hard to make it clear. Whatever success I had with that is up to you, but I feel like some of it comes from its speech-to-text transcription DNA.
Looking back, I really do wonder how different secondary school or university would have been if speech-to-text had been this ubiquitous and accurate when I was younger. It would I think have completely transformed my relationship with homework, school, and writing.
But after eight years of increasingly talking to my computers, I don’t really think of dictation as transcription software anymore.
It is a way of composing of generating text.
For anyone who hasn’t tried using transcription as their primary engine for generating text, I highly encourage it, give it a real go, if only for the fascinating cognitive shifts you might observe.
On The Blog
Start Select Reset Zine #016 | INTERNET OPERATOR
Issue #016 of Start Select Reset is now out, and on the way to via snail mail to my supporters (£5/month+)

Last issue I said that the end-to-end production run of both the podcast and the zine took 5x longer, and more effort, than I thought it would. But this time around came together a lot quicker. Perhaps 3x. As I get more practised, I expect that multiplier will keep coming down.
With the previous issue, a great deal of time was spent in Affinity, getting comfortable with the layout and basically learning what I was doing while laying out 36 pages of text.
Start Select Reset 📑
Subscribing to SSRZ supports my online work and creative projects.
As a thank you, I send you my zine four times a year, just like it’s 1994.
Photo 365

Terminal Access
This piece from 2024 up on Dazed is worth revisiting I feel: Why does nothing feel real anymore?
Similarly, as we crawl out of our simulation pods, AKA our phone screens, glimpsing at the dystopian horror that surrounds us, there’s a pervading sense that nothing feels real anymore, with an ever-accelerating news cycle that takes major historical events splintering them entirely into unrecognisable bytes of information – out-of-context video clips, self-referential memes and AI remixes – on our social media, leaving us disorientated amid the chaos.
Dipping the Stacks
The Rise and Fall of the Artificial State
History, Lepore demonstrates, is driven not by machines but by individuals, societies, cultures and ideas. The architects of the Artificial State had no theory of governance, nor did they have any interest in the rule of law, or any capacity for restraint. They drew their ideas not from science, but from science fiction.
By utilizing a shared transformer architecture, SHELLS maintains global surface consistency while requiring only 12% of the GPU memory of volume-based approaches. Experimental results show that SHELLS reduces median registration error by 21% – 29% and achieves a 3.5× speedup, predicting 18k-vertex meshes in 0.08 seconds. Notably, our model is trained exclusively on synthetic data yet generalizes effectively to real-world captures
NVIDIA details DLSS 5: three models, per-object controls and single-GPU support – VideoCardz.com
DLSS 5 adds a generative stage after the game produces a conventionally rendered frame. The original renderer still defines the geometry, camera, lighting setup and composition. DLSS 5 then processes that frame to improve materials, shadows, reflections and other visual details.
The Night France Saved the British Grid
What follows is the anatomy of a near-miss that the official record describes, accurately but incompletely, as a day when the grid kept working. It did. In the same sense that a driver who runs a red light and isn’t hit keeps driving.
The U.S. Is Pulling Back on Energy Efficiency
The turn away from efficiency has been oddly timed, said Christine Egan, chief executive of CLASP, a nonprofit group that works on efficiency policies worldwide. Energy prices are going up, demand is going up, and you’d think we would want our products to be as efficient as possible.
Reading
I’m in that part of the year where the crunch is on, I find my self with a bunch of books on the go. I’m currently reading:
- Postal Intelligence: The Tassis Family and Communications Revolution in Early Modern Europe by Rachel Midura
- The Mountain of Silence: A Search for Orthodox Spirituality by Kyriacos C. Markides
- Real Presences by George Steiner
- The Experience of God : Being, Consciousness, Bliss by David Bentley Hart
Big christian mysticism and meaning vibes, alongside a very detailed (but enjoyable) read on the history of the postal service in renaissance Italy. hahah.
Music
The Limits of the Frame – Cole Pulice (Single)
Pulice is one of my favourite contemporary musicians full stop. This new single is according to his Twitter “completely made with live signal processed saxophone, tho it could easily be mistaken for pure electronics”. I listened to it (stunning) i was curious to know more and found this on the band camp single release:
Cole sought to draw new resonance from their beloved saxophone simply by not playing it. Jokingly calling the work ‘anti-saxophone” in its avoidance of using breath or the mouthpiece, Cole fabulated sound by processing resonant vibrations from the body of the saxophone and the percussive clacking of its keys.
Which recontextualizes the title in a big way. If you like ambient glitch and the blade runner soundtrack this is a single for you.
Remember Kids:
Getting to zero is an ethico-political compass for hospicing modernity.
Vanessa Machado De Oliveira – Hospicing Modernity

Leave a Reply