Speech, Vision & Image Generation Concepts
10 free practice questions with explanations
10 free questions · instant explanations · no sign-up
PassNova has 10 free Microsoft AI-901 (Azure AI Fundamentals) practice questions on Speech, Vision & Image Generation Concepts, each with a clear explanation. Practise them in the browser with instant feedback — 100% free, no sign-up, on any device. Updated for 2026.
Speech, Vision & Image Generation Concepts: example questions & answers
10 worked examples with answers and explanations below. Practise them in the browser with instant feedback on every answer.
A mapping app reads turn-by-turn directions aloud to the driver from text instructions. Which speech capability is the app using?
- ASpeech recognition (speech-to-text)
- BBeam search decoding (hypothesis search)
- CSpeech synthesis (text-to-speech)✓
- DAcoustic modelling (phoneme prediction)
Answer: Speech synthesis, also called text-to-speech, converts written text into spoken audio. The learning path lists navigation instructions in mapping and GPS applications among its notification and alert scenarios for synthesis. Speech recognition works in the opposite direction, transcribing spoken words into text, while acoustic modelling and beam search decoding are internal stages of the recognition pipeline rather than capabilities an app selects.
Which feature-extraction technique splits audio into overlapping 20-30 millisecond frames, applies a Fourier transform, maps the frequencies to a scale matching human hearing and typically produces 13 coefficients per frame?
- AMFCC feature extraction✓
- BNeural vocoding
- CG2P conversion
- DBeam search decoding
Answer: Mel-Frequency Cepstral Coefficients (MFCCs) are the most common feature extraction technique in speech recognition, producing one compact feature vector per frame as the input to acoustic modelling. Grapheme-to-phoneme conversion belongs to speech synthesis and maps written letters to sounds. Beam search decoding selects the best word sequence later in the recognition pipeline, and a neural vocoder converts mel-spectrograms into audio during synthesis.
A speech recognition system's acoustic model cannot distinguish 'their' from 'there' because both share identical phonemes. Which pipeline stage applies vocabulary, grammar and common word patterns to resolve the ambiguity?
- APost-processing
- BLanguage modelling✓
- CAcoustic modelling
- DBeam search decoding
Answer: Language models resolve phoneme-level ambiguity by applying knowledge of vocabulary, grammar and statistical word patterns, for example knowing that 'The weather is nice' is far more common than 'The whether is nice'. Acoustic modelling only predicts a probability distribution over phonemes for each frame, which is where the ambiguity arises. Decoding searches for the best hypothesis using both models' scores, and post-processing handles capitalisation, punctuation and number formatting on the chosen text.
A text-to-speech voice pronounces every phoneme correctly, yet listeners describe it as flat and robotic. According to the learning path, which stage of the synthesis pipeline is most likely at fault?
- ALinguistic analysis
- BProsody generation✓
- CWaveform generation
- DText normalisation
Answer: Prosody generation determines the rhythm, stress and intonation of speech by predicting pitch, duration and energy for each phoneme. Robotic-sounding output usually results from flat, monotone prosody rather than from imperfect phoneme pronunciation. Text normalisation expands abbreviations and numbers into spoken forms, linguistic analysis maps text to phonemes, which in this case were already correct, and waveform generation by a neural vocoder simply renders the prosody targets it is given.
A grocery store's smart checkout must scan several items placed together and report both what each item is and where it appears, as rectangular bounding boxes. Which computer vision technique does this describe?
- AObject detection✓
- BImage classification
- CSemantic segmentation
- DLaplace filtering
Answer: Object detection models examine multiple regions of an image to find individual objects, returning each detected class with the coordinates of its rectangular bounding box. Image classification predicts a single label for the main subject of the whole image, so it cannot locate several items at once. Semantic segmentation goes further by classifying individual pixels, and Laplace filtering is an edge-highlighting image-processing operation rather than a recognition model.
An image-processing step slides a 3x3 kernel of weights across an image, calculating a weighted sum for each patch of pixels and writing the result to a new array, with the effect of highlighting the edges of shapes. What is this operation called?
- AConvolutional filtering✓
- BSemantic segmentation
- CSoftmax classification
- DPooling to downsize feature maps
Answer: Convolutional filtering convolves a filter kernel across an image, multiplying each pixel in a patch by the matching kernel weight and summing the results. The edge-highlighting kernel in the learning path is a Laplace filter, values outside the 0 to 255 range are adjusted to fit and the uncalculated outer edge is padded, usually with 0. Semantic segmentation classifies pixels by object, pooling shrinks feature maps inside a CNN, and a softmax function turns network outputs into class probabilities.
A convolutional neural network trained on apple, banana and orange images returns [0.2, 0.5, 0.3] for a new image. What does this output represent?
- AA loss value per class, produced by comparing to labels
- BA flattened feature array, produced before the dense layers
- CA feature map, produced by the convolutional filter layers
- DA probability for each class, produced by a softmax function✓
Answer: The output layer of a CNN uses a softmax or similar function to produce a probability value for each possible class. A result of [0.2, 0.5, 0.3] therefore means banana (class 1) is the most likely label for the image. Feature maps are the arrays generated by the filter layers and the flattened array is the single-dimensional input to the fully connected network, both of which come earlier, while loss is a single training value calculated by comparing predicted probabilities with the actual label such as [0.0, 1.0, 0.0].
How does a vision transformer (ViT) differ from a language transformer in what it encodes as embedding vectors?
- AIt applies a Laplace kernel to every pixel and embeds the resulting edge maps as vectors
- BIt flattens feature maps produced by a CNN and embeds them through a softmax output layer
- CIt extracts patches of pixel values and embeds visual features such as colour, shape and texture✓
- DIt extracts text tokens from associated captions, then embeds their linguistic attributes
Answer: A vision transformer extracts patches of pixel values from an image and generates a linear vector from each patch. The same attention technique used in language models then relates the patches to one another, so the embedded values reflect visual features such as colour, shape, contrast and texture, and features that commonly appear together, like a hat and a head, end up with similar vector directions. Encoding caption tokens is what the language encoder does, while Laplace kernels and flattened CNN feature maps belong to convolutional approaches.
Most modern image-generation models begin with a random set of pixel values and iteratively remove noise to create structure, comparing the image with the prompt after each iteration. What is this technique called?
- ADiffusion✓
- BConvolution
- CAttention
- DSegmentation
Answer: Diffusion is the technique used by most modern image-generation models, in which a prompt identifies related visual features and an image is refined iteratively from random noise until it depicts the requested scene. Convolution is the filtering operation used in image processing and CNNs, and attention is the mechanism transformers use to weigh context between tokens or patches. Segmentation classifies the pixels of an existing image rather than creating a new one.
A video-generation model uses the same approach as image generation to identify the visual features associated with language tokens. Which additional factors does the learning path say video generation must also take into account?
- AAudio sampling rate and mel-spectrogram resolution
- BFilter kernel weights and feature map pooling
- CBounding box coordinates and per-pixel object classification
- DPhysical behaviour of objects and temporal progression✓
Answer: Video generation extends the diffusion approach by also modelling the physical behaviour of real-world objects and the temporal progression of the scene. That is what keeps a generated dog's feet on the ground and makes the clip a logical sequence of activity. Bounding boxes and per-pixel classification are outputs of object detection and semantic segmentation, sampling rates and mel-spectrograms belong to speech processing, and kernel weights and pooling are CNN training concepts.