Insiab caption intelligence
Captions that are already right when you open the timeline.
- Two layers, not one
- Speech recognition gets the words. A language model decides where each caption should break.
- Timings never invented
- The model regroups word ids. Every timestamp is computed from the audio, so it cannot drift.
- Arabic and English together
- Mixed speech is normalised term by term, and direction is resolved per caption.