Turn a plain text file — or pasted text — into subtitle lines, with generated time codes. The time codes can then be aligned to the actual audio with a forced aligner or via speech to text.

Paste or type text into the edit box, or click Import… to load a text file. Tick Import from multiple text files (one file is one subtitle) to switch to a file list where each file becomes one subtitle line — useful for batch workflows; .txt and .rtf files can also be dragged onto the list, and time codes embedded in file names are picked up automatically.
Without fixed duration, each subtitle’s duration is calculated from its text length using your optimal characters/second setting, clamped between the minimum and maximum display duration. Lines are laid out sequentially starting at zero.
The preview grid updates as you change options and shows the resulting subtitles with their time codes.
Sequential time codes from text length are only a starting point. If you have the video, two buttons at the bottom can time the text against the actual speech:
A forced aligner matches the text you already have against the audio, without transcribing it first — this is fast and keeps your text exactly as written. Long videos are aligned in windowed chunks, so any length works.
Setup dialog:
Missing models are downloaded automatically when you press OK. Progress is shown per window and line while aligning. If the script is longer than the speech in the video, the trailing lines cannot be aligned and a warning tells you how many lines were matched.
Runs normal speech to text on the video, then matches your script against the transcript. Useful when the forced aligner does not support your language. Lines that cannot be matched get interpolated time codes.
Pressing OK replaces the currently loaded subtitle with the imported lines. Save your current work first if you need it.