From a folder of raw interviews to broadcast-ready audio: the path is shorter than it’s ever been, and almost all of it happens in the transcript. Work with what people actually said, natively, on your Mac.
Ten steps for the craft, plus one optional step for people who script their work. Almost all of it lives in the transcript.
Download Fuchsia Studio from the Mac App Store and open it. That’s the whole setup: no sign-up, no login, no cloud project to create. You land in a native Mac document window and open a story — or start one. Studio is free for fourteen days — and the fourteen days start when you first place a clip, not when you download — then $129 or equivalent, one time; it runs on macOS 26.
Drag in your interviews and field recordings — hours of them, at whatever sample rates your sessions produced. Studio transcribes everything and separates the speakers, and both jobs run entirely on your Mac: transcription uses Apple’s on-device speech engine, and speaker separation happens locally too. Your audio doesn’t go anywhere to be understood.
Speakers arrive as unnamed voices — Speaker 1, Speaker 2 — and you can rename them as you recognize who’s who. (If you later opt in to the AI features, Studio can propose names — but only when the name is actually spoken on tape, and always with that moment attached for you to play. Whether the name belongs to this voice is your call, not Studio’s. More on that in step 8.)
Open Tape view (⌘2) and your source lays out as a readable page — the interview as text, verbatim, with every word still anchored to its moment in the audio. Press ⌘F and type a keyword to land on a half-remembered line across hours of material.
And when a moment strikes you, select it: Studio surfaces up to eight passages elsewhere in your tape that echo it, even when they’re worded differently. Both ways of finding things — the search field and the related moments — run on your Mac, with no key, no account, and no network call. To be precise about what each does: the search field finds the words people said, and words with similar meaning; the echoes start from a moment you select, not from a typed theme.
This is the gesture the whole app is built around. Drag across the words that matter in the transcript, and your selection lifts into a clip. You build the piece from what people said, instead of scrubbing a waveform hunting for a line you half-remember.
Switch to Assembly (⌘3) and put your clips in order. This is your piece as structure — one running order you sculpt scene by scene, not a multitrack timeline. When a cut lands mid-word, Studio finds the clean edge for you — so you’re not nudging handles a frame at a time. Whenever you want to check a clip against its source, View in Context (⌘J) jumps straight to its place in the original interview. And to read the whole story as one continuous document, there’s Script view (⌘4).
Open any clip to refine it. Scrub under the loupe with varispeed tape-rock — it behaves like rocking reels on a deck. Trim to the word. Cuts snap to the natural silences in the speech, and you can override the snap whenever your ear disagrees.
Where a join needs a breath, add authored air: a beat of quiet to adjust the pace of your story. Air is a stored, visible value on the seam which you can adjust or clear. By default, consecutive clips of the same voice join tight and a change of voice earns a breath, so dragging clips in creates a typical audio storytelling flow. Studio decides which is which from the separated voice on the tape, and where a voice hasn’t been separated it falls back to the recording the clip came from.
Write your narration at the exact point in the piece where it will play, and record takes right there. Every take becomes tape — transcribed like the rest of your material, so your own voice is as searchable and cuttable as your interviews.
One thing Studio deliberately does not do: make a voice. There’s no synthesis, no cloning, no machine narration. Narration is always your own recorded performance, through your own microphone.
Everything above runs with no AI at all. If you want it, connect your own AI provider — your key, your account — and Studio adds the reads: the Thread, a grounded collaborator that finds moments, proposes cuts, and answers “what is this section really about?”; interpretations of your material that land in Reading view (⌘1); and those speaker-name proposals — a name is suggested only when it is actually spoken on tape, the moment plays for you, and you confirm that it belongs to this voice. Names you’ve typed yourself are never overwritten.
The honest shape of this: it’s opt-in, it runs on your own provider account rather than on your Mac, and Studio sends only text — never your raw audio. Connecting doesn’t sweep the tape you already have; it’s the new tape you add that gets read automatically, and you can skip any single import or pause a whole library. And in every case the AI proposes while you approve: it quotes the tape verbatim or it refuses — it never invents a line, it doesn’t decide who’s speaking (the tape does), and it never drafts your narration.
Real-world tape gets hurt, so Studio ships two repairs. De-hum notches out mains hum and its harmonics (50 or 60 Hz). De-clip rebuilds the peaks a clipped recording lost, reading them back from the waveform either side. Neither ever invents audio that wasn’t recorded — they restore what the recording itself implies. Like the rest of the craft, they run on your Mac.
When the piece is cut, bring it to broadcast loudness with a finishing stage built for the spoken word, then export it your way — WAV, AIFF, AAC, Apple Lossless, FLAC or MP3. Or hand the whole story to your DAW as a session for Ableton Live or Adobe Audition, or a Final Cut Pro XML that Logic imports. Sample rate takes care of itself: bring it all in at any rate, and Studio always finishes at 48 kHz, conforming other rates as it renders and naming each source it resampled in the export report. The file you deliver is a separate choice — the Export sheet offers the rates your chosen format can actually deliver.
The DAW session carries the raw audio: your clips, nicely labelled, in place, cut from your original tape, without any repairs, compression or loudness baked in.
If your work is pipelines, Studio Pro adds a command line and an MCP connection, so you can drive the same studio from your own AI assistant: batch a week of tape overnight, then render a finished, loudness-targeted file with no app window open.
Studio Pro isn’t available yet. Get in touch to discuss pricing and availability — it will come direct from Fuchsia rather than through the App Store. Pro adds automation; every step above is in the base app.
That’s the path: tape in, read it, select the words, shape the order, give the seams air, speak your narration, finish to loudness, and hand one finished file to the world. The core craft — import to finish — happens on your own machine, and you stay the author of every cut.
Present tense means it ships today. Anything on the roadmap, we say so.
Plain answers about the edges — what’s in the app today, and what isn’t.
Studio gives you two ways to find things. The search field finds moments: type what you remember and Studio surfaces where it was said — the words themselves, and words with similar meaning, so you don’t need the exact phrase. Every result carries a play control. You can also find moments related to something unearthed in search, even when the words are different. For word-by-word searching inside a single tape, the transcript view has its own find (⌘F). All of it runs entirely on your Mac — no key, no account, no network connection.
What Studio doesn’t have is a box where you type a theme — “grief”, “leaving home” — and get back the concept. Finding by idea starts from a moment you select, not from free text. And naming the themes and through-lines in your tape is the job of the optional AI reads, which run on your own provider account once you choose to connect one.
It separates them; it doesn’t recognize them. When you add tape, Studio automatically tells the voices apart — that happens on your Mac, no account — and each speaker arrives unnamed: Speaker 1, Speaker 2, and so on. You can name a voice yourself at any point, and a name you type is never overwritten.
If you’ve connected an AI provider (optional, on your own key), Studio can also propose a name — but only when that name is actually spoken on tape. The proposal arrives with the moment attached: you play it, hear it said, and confirm that it belongs to this voice. No confirmation, no name.
No. There’s no voice synthesis in Studio — nothing that generates, imitates, or clones a voice. Narration is your own performance: you write it at the point in the piece where it plays, record takes right there through your microphone, and every take becomes tape, transcribed like the rest of your material. What your listeners hear is what you recorded.
No. It proposes; you decide. And it’s off until you turn it on — the AI features are opt-in, connected through your own provider account, and Studio sends them text only, never your audio.
Once connected, the Thread behaves like a well-prepared collaborator: it finds moments, proposes cuts, and answers questions like “what is this section really about?” When it speaks about your tape, it quotes it verbatim — or it refuses. It doesn’t draft your narration, invent a line, or put words in the wrong mouth. And nothing it proposes changes your story until you approve it. You stay the author.
Sessions and stems, yes. Studio exports a native session for Ableton Live and Adobe Audition, and a Final Cut Pro XML for Logic Pro — every clip in place on the timeline and labelled from your transcript (whoever’s speaking, or the line itself where the voice isn’t named yet), with your authored air preserved on the seams. Logic opens it through File ▸ Import ▸ Final Cut Pro XML — turn on Logic’s Complete Features first, and use Import rather than double-clicking the file, since Final Cut owns that file type. Final Cut Pro opens it too, and there double-clicking is the way in: your clips arrive in place and labelled, as sound without picture. Audition and Ableton open their own session files directly. For anything else there are universal stems: the same labelled audio, plus a plain description of the layout.
Be clear about what lands: the raw audio, cut from your original tape — mono, unprocessed, on one track named for your story, with narration on its own. Your repairs, compression, loudness and stereo placement are not baked in, because the point is to hand the mix to your DAW, not to pre-empt it. If you want the finished sound, export the finished audio instead.
Two ways to find things — and both run on your Mac, with the network off.
No — none of the three. Both the search field and the related-moments feature run on an index built on your Mac, and neither makes a network call. Search works with the network off.
Not from search itself. Search finds moments — the words people said and words with similar meaning — and echoes a moment you select; it doesn’t hand you a named theme. Naming the themes and through-lines is interpretive work, and that’s the job of the optional AI reads — which run on your own provider account, only once you choose to connect one.
The category we expect you to read most closely, so the answers are scoped precisely. The short version: the craft runs on your Mac, your audio never leaves it, and the AI is something you turn on — on your own account — not something that happens to you.
No. Your audio stays on your Mac — through transcription, speaker separation, search, editing, assembly, and finishing. That holds even if you later connect an AI provider: the Thread and the reads work from text, and Studio sends text only, never your recordings.
On your Mac, with no account: import at any sample rate, transcription (Apple’s on-device speech engine), speaker separation, both search paths and the index behind them, clip editing and assembly, narration recording, the two tape repairs (de-hum and de-clip), and finishing to a loudness-targeted file. Once your Mac has what transcription needs, all of that runs with the network off.
That one condition, in full: the speech model is supplied by macOS, not shipped by us, so if this Mac doesn’t already have it, macOS downloads it — once. Studio asks for it in the background the moment you open the app, so it isn’t something you wait on to start working; a tape you import before it arrives simply has no words until it does. It is a download, not a send: nothing of yours goes out with it. If your Mac already has the model, nothing is requested at all.
Not on your Mac: the optional AI reads — the Thread, the interpretations that land in Reading view (⌘1), and speaker-name proposals. Those run on a hosted AI provider you connect yourself, on your own key and account, and they receive text only. If you never connect a provider, that whole layer simply stays off.
No. There is no upload, no cloud project, and no account. Studio transcribes, separates speakers, indexes, and edits entirely on your Mac — and once your Mac has the speech model macOS transcribes with, you can do the whole craft with the network off.
Just text. Everything in the AI layer — the Thread, the interpretations in Reading view, the speaker-name proposals — works from text drawn from your tape, so what that layer takes off your Mac is text, sent to the provider you connected under your own account. Your audio files are never sent. That layer is also the only thing that carries your material off this Mac at all: the app’s other uses of the network — macOS fetching its speech model, and asking the App Store what you’ve bought — send nothing of yours.
Yours. AI is opt-in, and turning it on means connecting your own provider account with your own key. There is no Fuchsia account behind it — your agreement is with the provider you chose, and any AI usage costs are between you and them.
No. Connecting a provider doesn’t sweep the tape you already have — nothing from your existing library is sent on its own. What gets read automatically is new tape you add from then on — and even there you stay in charge: you can skip the read for any individual import, and pause reading for an entire library.
One thing to know precisely: if you later ask the Thread about older material, or open a read on it, the text needed to answer is sent at that moment — because you asked, not because Studio swept.
Yes. If you never connect a provider, nothing is sent at all. If you do connect one, the automatic read is yours to control: skip it for any individual import, or pause it for a whole library. Beyond the automatic read, text is sent only when you yourself ask the Thread or open a read on material — those requests happen when you make them, not on a schedule.
Studio doesn’t police this for you, and we think that’s the right call — we can’t know what permissions you hold, and tape you didn’t record yourself often comes with a waiver that covers you. You decide what gets sent. The controls are yours: don’t connect a provider at all and nothing is ever sent; skip the automatic read on any single import; or pause it for a whole library. And in every configuration, only text is ever sent — never your audio.
On your machine. Both retrieval paths — the search field across your tape, and the related moments that echo a passage you select — run against an index built on your Mac. No key, no account, no network call, for the index or the searches.
We designed the boundary to be exact so you can make that judgment with full knowledge. Out of the box, Studio asks for no account, and your audio never leaves your Mac in any configuration: your tape, transcripts, search index, and edits all live there, and nothing derived from them goes anywhere.
Some things do reach the network while carrying none of your material, and we would rather name them than leave you with an assurance there are none. If this Mac doesn’t already have the speech model macOS transcribes with, macOS downloads it — one time, inbound; if it already has it, nothing is requested. The app asks the App Store what you’ve bought. And if you have saved an AI provider key, pressing Test beside it in Settings sends one minimal request to that provider to check the key works — it carries the key, and nothing drawn from your tape. None of them sends your tape, your transcripts, or anything else drawn from them. These are the routes we can name and that you can act on — not a promise that there is nothing else: macOS itself does things on any Mac that aren’t ours to describe, such as checking that an app is properly signed.
If you choose to connect an AI provider, the AI layer works by sending text — never audio — to that provider, on your key, under your agreement with them. That text moves on two paths, and you should know both. First, new tape you add is read automatically, and you can skip any import or pause a whole library. Second, when you yourself ask the Thread a question or open a read — including on tape that was in your library before you connected — the text needed to answer is sent at that moment.
Whether a given provider is appropriate for a given source is your call, as it should be. Studio’s job is to make sure you always know exactly what would leave, and to keep everything else on your machine.
Optional, opt-in, and grounded in your tape — here’s exactly how the reads behave.
They’re Studio’s optional layer of interpretation. Connect your own AI provider and Studio can read new tape as you add it, landing its interpretations in Reading view (⌘1) as material you can draw on — and it gives you the Thread, a collaborator you can ask to find moments in your source, propose cuts, and answer “what is this section really about?” The role is deliberately bounded: the AI reads and proposes. It never writes your story, and it never changes anything without you.
No. Studio is a complete tool without them. Importing, transcription, speaker separation, search, editing, assembly, and finishing all run on your Mac, with no account and no key — and, once your Mac has the speech model macOS transcribes with, with the network off as well. If you never connect an AI provider, you’re not using a lesser version of the app — you’re using the whole craft.
Three things: find moments in your tape, propose cuts, and answer questions about what your material is saying. It works from your transcripts, and it holds to one working rule — when it refers to your material, it quotes the tape verbatim, or it refuses. It proposes; it never decides. Nothing it suggests touches your assembly until you approve it.
You approve every change. Anything the Thread proposes — a cut, a placement — arrives as a proposal showing you plainly what it refers to, and nothing moves until you say yes. You stay the author; the AI is a reader in the room.
No. This is the Thread’s working rule, not a best effort: it quotes the tape verbatim or it refuses. It never invents a line. It doesn’t decide who’s speaking — the tape does. And it never changes your story without your say-so.
Which voice a quote belongs to is a separate question, and the AI does not answer it: a quote is attributed by the voice registry, never by the model. So if the speaker separation put a line under the wrong voice, the AI will carry that through rather than catch it. In documentary work a misattributed quote is a serious failure — if you ever see one, we’d like to hear about it.
Every answer and every proposal is tied back to the actual material it cites, so you check it against the tape itself — not against the AI’s say-so. And honestly: it can still point at the wrong moment. That’s exactly why nothing it proposes takes effect until you’ve seen what it resolved to and approved it. Trust here comes from the method, not from a promise that the model is infallible.
No. The Thread finds, proposes, and answers questions — it doesn’t draft. Narration in Studio is your own work: you write it at the point in the piece where it plays and record it in your own voice, and every take becomes tape like the rest of your material.
Speakers arrive unnamed — Speaker 1, Speaker 2 — and stay that way unless you name them or opt in to the reads. With AI connected, Studio proposes a name only when that name is actually spoken on tape: the moment plays, and you confirm before the name is applied. Could a proposal be wrong? Yes — tape is messy, and that’s exactly why the moment plays and you decide. A name you’ve typed yourself is never overwritten.
It’s off. Out of the box Studio has no AI connected and nothing derived from your tape goes anywhere — there’s nothing to opt out of. (Studio still asks the App Store about your trial or purchase, and macOS fetches Apple’s speech model if this Mac hasn’t got it; neither carries a word of your material.) The reads switch on only when you connect your own provider account. From then on, new tape you add is read automatically so the interpretations are ready when you are; you can skip the read for any individual import, and pause a whole library at any time. Connecting doesn’t reach back and sweep the tape you already have.
One: Claude. The reads and the Thread run on Anthropic’s Claude models, on your own API key and your own account, and they only ever receive text — never your audio. One thing to know before you start, because it catches people out: the key is an Anthropic Console key, billed pay-per-token on a Console account. A Claude Pro or Max subscription is not API access and can’t be used here. There’s no local-model or self-hosted path in this release: the reads are a cloud workload by design, which is part of why they’re strictly opt-in.
Studio itself is free for fourteen days, then a one-time purchase of $129 or equivalent. The reads run on your own provider account, under your own API key, so your provider bills you directly at their rates — what the reads cost depends on how much tape you have them read, and that’s between you and your provider. If you never connect a provider, the whole craft still works, and there’s no usage to be billed for.
No. The reads and the Thread are a cloud workload: they run on your provider’s models, over the network, on your own key — there’s no local-model mode. Everything else in Studio is on-device: import, transcription, speaker separation, search, editing, assembly, and finishing all run on your Mac, and all of it works with no connection once your Mac has the speech model macOS transcribes with.
The base app and Pro both contain every feature you can reach from the UI. Pro adds ways to drive it from outside the app.
A command line, and an MCP server you can connect to your own AI assistant — two headless ways into the same studio. It’s for people who want to script and batch their work, or run Studio from their own AI assistant: batch a week of tape overnight, build a pipeline you can run every story, render the finished audio without opening a window.
Yes, really. Studio Pro includes an MCP server — MCP is the open standard AI assistants use to work with tools outside the chat window. Connect it and your own assistant can work in your Studio projects: find moments in the tape, propose cuts, answer questions about the material, and render the finished audio. When it’s done, open the piece in Studio and carry on by hand.
No. The craft is all in the base app: import, transcription, speaker separation, search, clips, assembly, narration, tape repair, finishing and export. Your library is managed for you — the app makes one and looks after it, so you’re not answering a storage question before you’ve made anything, and the work inside it stays yours to open and export.
Contact us to discuss pricing and availability for Fuchsia Studio Pro. You’ll obtain it directly from us as it can’t be sold through the Mac App Store: Pro adds command-line tools, and a tool bundled inside a Mac App Store app has to run inside that app’s sandbox, which means it can’t be run from a terminal at all.
One programme, one file — here’s exactly what comes out of Studio.
When your story is ready, Studio finishes it on your Mac with a render stage built for the spoken word. This surveys your show as a whole, and brings it to broadcast loudness — measured to ITU-R BS.1770 — while preserving the dynamics you cut, using a bespoke linear normalizer. Your whole piece hits your chosen target with a single, even gain. After the bake, Studio shows you the measured loudness and the true peak, read back off the finished file, indicating clearly when it had to hold that gain back to stay under the true-peak ceiling, landing a touch under target.
The finished audio: a single loudness-targeted file of your whole piece. Export works like a print dialog — File ▸ Export… (⌘E) — where you choose the format and the loudness target, confirm, and Studio bakes the file. Before it renders, a readiness check looks for two things: if a clip’s source is offline it warns you and says how many clips will be skipped, and if a tape carries more than two channels it refuses and names the tape. Neither becomes a hole you find after delivery.
Six: WAV and AIFF (lossless, for handing to an editor or a DAW), Apple Lossless and FLAC (lossless, compressed), and AAC (.m4a) and MP3 — the compressed files podcast hosts ingest directly, each with a choice of bitrate. Sample rate is a separate row, and it offers the rates your chosen format can actually deliver. The loudness target is chosen per export, right in the sheet: Streaming at −14 LUFS is the default, with Podcast (−16), Broadcast (−23), Download (−9), and Off — rendered and measured, but not normalized. Between the six, both the handoff and the publish are covered.
Yes. The finished audio is a standard file with nothing proprietary about it — the 24-bit WAV drops straight into Logic, Ableton, Audition, or wherever you finish. If Studio’s finished file is your final deliverable, ship it; if the piece goes on to a final mix elsewhere, hand that session the WAV — or export a DAW session instead, and get the parts rather than the whole.
Yes — a session for Ableton Live or Adobe Audition, a Final Cut Pro XML that Logic imports, or universal stems. In the three session formats every clip lands in place and labelled from your transcript — whoever’s speaking, or the line itself where the voice isn’t named yet — with your authored air intact; the stems are that same labelled audio plus a plain description of the layout. What lands is the raw audio, unprocessed, cut from your original tape: all your tape on one track, with your narration and any sound beds on their own. Speech stems are mono, because a second channel might be a second microphone and only you can say which one is the voice; bed stems are stereo, because both channels of music are the material. Your repairs, compression and loudness stay in Studio, so the mix is yours to make. If you want the finished sound instead, export the finished audio.
Sample-rate worries are a thing of the past. Cut a 44.1 kHz archive clip against 48 kHz field tape and never think about it — just bring it all in. There are two rates here, and only one of them is yours to choose. Studio always finishes at 48 kHz: sources at other rates are conformed to it as it renders, and the export report names each source it resampled — say, 44.1 → 48 kHz — so nothing changes silently. The file you deliver is a separate choice: the Export sheet’s Sample Rate row offers the rates your chosen format can actually deliver, and Studio converts the finished 48 kHz render to whichever you pick.
No — export is one-way out: the finished audio, a session for Ableton Live or Adobe Audition, or a Final Cut Pro XML that Logic imports — every clip in place and named, your authored air intact, cut from your original tape. Nothing imports back in, and it doesn't need to: the story stays fully editable on your Mac, so when the piece changes you re-export rather than round-tripping.
Yes. Name a chapter in Assembly (⌘3) as you build the running order, rename it as the shape changes, and leave any chapter out of the chapter list when it isn’t one for listeners. Export to AAC or Apple Lossless (.m4a) and the names travel inside the file as a chapter track. Whatever format you choose, Studio also offers the chapter timestamps when the export finishes — as text you can paste into a show description. A chapter that holds no audio is left out.
Yes — File ▸ Export Captions… writes a caption file for your finished story, as WebVTT (.vtt) or SubRip (.srt). It carries the words of the finished piece at the finished piece’s own times: the cues are read from the same layout the render used, so there’s no second pass to drift from it. WebVTT can carry the speaker’s name alongside the words; SubRip has no speaker construct, so there it’s the words alone.