0 / 60 seg.

So the solution is actually to take inspiration from another domain: speech recognition.