0 / 60 seg.

So to do that, can we actually teach the computer to imitate the way someone talks by only showing it video footage of the person?