Of the whole video, only the audio is used
There's no need to convert anything beforehand. The file is uploaded as is, the internal audio track is processed, and the picture is never analysed: it adds nothing useful to the transcription and would only slow the process down.
This applies to all common containers: the MP4 from a phone, the MOV from an iPhone, the MKV from a screen recording, the WEBM from a downloaded meeting. If the file contains audio, it's transcribed; if it doesn't, you're told before any minutes are used from your quota.