0 / 60 seg.

And it sees multiple people in a scene, remembers where individual people are, and looks from person to person, remembering people.