Transcript Cleaner
Paste a transcript or .srt/.vtt file. Remove timestamps, speaker labels and filler words.
0 characters · 0 words · 0 lines
0 characters · 0 words · 0 lines
Everything runs in your browser. Your text never leaves this page.
What is a transcript cleaner?
A transcript cleaner takes the raw text you exported from a video or audio recording (or a pasted .srt or .vtt subtitle file) and turns it into something readable. Recordings from YouTube, Zoom, podcasts, online courses and interviews export with timestamps on every line, speaker tags, filler words like "um" and "uh" and choppy line breaks. This tool sweeps all of that away so you are left with the actual conversation. Everything runs in your browser, with no upload, so even private interviews and client calls stay on your device.
Remove timestamps and speaker names
The two options that matter most for transcripts are on by default. Remove timestamps deletes times such as 00:00, 1:02:33 and bracketed forms like [00:12]. Remove speaker labels drops a name and colon at the start of a line, so "Speaker 1:" or "INTERVIEWER:" disappear while the words they said remain.
Convert .srt and .vtt subtitle files to plain text
Exported captions come as SubRip (.srt) or WebVTT
(.vtt) files, padded with scaffolding around every line. Paste the
file contents and the cleaner removes the WEBVTT header, the cue
numbers (1, 2, 3…) and the timecode lines
like 00:00:01,000 --> 00:00:04,000, leaving only the spoken
words. A caption file becomes a clean transcript you can read, quote or
repurpose, without uploading anything.
Remove filler words (um, uh, er)
Spoken transcripts are full of verbal fillers. Turn on Remove filler words and the cleaner deletes um, uh, er, ah, hmm, mm and uh-huh, along with the stray comma they often leave behind. It is deliberately conservative: it only targets these non-words, so genuine words that sometimes act as crutches ("like", "actually", "you know") are left exactly where they are. You get a tighter transcript without the tool quietly rewriting real sentences.
What the transcript cleaner removes
Each option is independent, so you can keep what you need and switch off the rest. Here is what the main transcript options strip out:
| Option | Removes | Example in → out |
|---|---|---|
| Clean .srt / .vtt subtitles | WEBVTT header, cue numbers, timecode lines | 00:00:01,000 --> 00:00:04,000 → removed |
| Remove timestamps | Inline and line-start times | [00:12] Hello → Hello |
| Remove speaker labels | Name + colon at line start | John: Hi there → Hi there |
| Remove filler words | um, uh, er, ah, hmm, mm, uh-huh | So um, I think → So I think |
Clean YouTube and Zoom transcripts
YouTube auto-captions and Zoom exports tend to be long lists of tiny lines. Normalizing line breaks stitches those fragments back together, and trimming and collapsing spaces removes the gaps left behind. The result reads like a document instead of a caption file, ready to quote, summarize or publish. With a clean word count in hand, you can also check how long it takes to read aloud.
From transcript to chapters
A clean transcript is the perfect starting point for video chapters. Once you have tidy text, open the YouTube Timestamp Generator to map your sections to times and copy a chapter list straight into your video description.
Good to know
Frequently asked questions
How do I remove timestamps from a transcript?
Keep Remove timestamps ticked and paste your transcript. It strips bracketed times like [00:12] and bare times at the start of a line like 00:12 or 1:02:33, while leaving normal numbers in your sentences untouched.
How do I remove "John:" and other speaker labels?
Turn on Remove speaker labels. It deletes a name followed by a colon at the very start of a line (John:, Speaker 1:, INTERVIEWER:), so a colon used in the middle of a sentence stays in place.
Can I convert an SRT or VTT subtitle file to plain text?
Yes. Paste the contents of a .srt or .vtt file and keep "Clean .srt / .vtt subtitles" on. It strips the WEBVTT header, the cue numbers (1, 2, 3…) and the timecode lines like 00:00:01,000 --> 00:00:04,000, leaving only the spoken text as a clean transcript.
Does it remove filler words like um and uh?
Yes, with "Remove filler words" on it deletes spoken fillers such as um, uh, er, ah, hmm, mm and uh-huh. It is deliberately conservative: it only targets these non-words, so real words like "like", "actually" or "you know" are never touched.
Do I need to upload my file?
No. Unlike many subtitle and transcript tools, nothing is uploaded. The cleaning runs entirely in your browser, so your transcript never leaves your device, which is useful for private interviews, calls and client work.
Does it work with YouTube auto-generated transcripts?
Yes. Auto captions usually paste as many short lines with timestamps. Removing timestamps and normalizing line breaks merges them back into readable text you can edit or repurpose.
Can I turn a cleaned transcript into chapters?
Yes. Once the transcript is clean, head to the YouTube Timestamp Generator to build a copy-ready chapter list from your sections.
Will it keep my paragraphs?
By default it normalizes line breaks rather than deleting them. If your transcript is one line per caption, you can add turn line breaks into paragraphs to group the text, or leave it as is.