Methods Video-stroboscopy was used to reveal the laryngeal cavity and dynamic changes of the vocal fold, and sound spectrograph was used to examine audio frequency, pitch and width.
So essentially how AI voice generation works is that if you have enough audio of someone talking, it will learn like the actual voice frequencies, the actual like audio frequencies of someone's voice and the way they talk.