Automatic Speech Recognition (ASR) converts spoken audio inputs from a caller into text or commands that a voice operating system can process. Distractors like text-to-speech perform the opposite conversion (text to audio), while validating inputs and passing data to computer telephony integration (CTI) are separate telephony functions.