Generative AI Models & Deployment
12 free practice questions with explanations
12 free questions · instant explanations · no sign-up
PassNova has 12 free Microsoft AI-901 (Azure AI Fundamentals) practice questions on Generative AI Models & Deployment, each with a clear explanation. Practise them in the browser with instant feedback — 100% free, no sign-up, on any device. Updated for 2026.
Generative AI Models & Deployment: example questions & answers
12 worked examples with answers and explanations below. Practise them in the browser with instant feedback on every answer.
In the context of generative AI models, what is a token?
- AA numeric vector that encodes the meaning of a word
- BThe smallest unit of text or data that a model can process✓
- CThe rate limit assigned to a model deployment each minute
- DThe natural-language prompt that starts a completion
Answer: A token is the smallest unit of text or data that a generative AI model can process, and models break input into tokens such as words, sub-words, characters or punctuation. The vocabularies of the latest LLMs consist of hundreds of thousands of tokens. A vector that encodes meaning is an embedding, the per-minute rate limit is the TPM allocation, and the text that starts a completion is the prompt.
In a simplified transformer architecture, which component creates the embeddings by applying attention to each token in the context of the tokens around it?
- AThe decoder block
- BThe positional encoding
- CThe encoder block✓
- DThe masked attention layer
Answer: The encoder block of a transformer creates the embeddings by applying attention, examining each token in turn to determine how it is influenced by the tokens around it, with multi-head attention evaluating multiple elements in parallel. The decoder block uses those embeddings to predict the next most probable token in a sequence started by a prompt. Positional encoding only indicates where a token appears in the sequence, and masked attention is the decoder-side technique that ignores the tokens after the current one during training.
While a transformer decoder is trained on text for which the full sequence is already known, the tokens that come after the current token are ignored. What is this technique called?
- AMasked attention✓
- BCosine similarity
- CMulti-head attention
- DPositional encoding
Answer: Masked attention is the technique used to train the decoder: because the context for predicting a token can only include the tokens that precede it, the tokens after the current one are ignored, and the predicted token is compared with the known next token so the learned weights can be adjusted. Multi-head attention is the parallel evaluation of vector elements that makes encoding efficient, positional encoding records where a token sits in the sequence, and cosine similarity measures how close two embeddings are.
A developer wants to check whether the embeddings for the tokens dog and puppy are semantically close to each other. How can this be measured?
- ABy comparing the positional encodings assigned to them
- BBy counting how often each token appears in the training text
- CBy calculating the cosine similarity of their vectors✓
- DBy comparing their unique integer token identifiers
Answer: Embeddings are vectors whose dimensions are calculated from how tokens relate linguistically, so tokens used in similar contexts point in similar directions, and semantic closeness is measured by calculating the cosine similarity of the vectors. Token identifiers are arbitrary integers assigned when the text is tokenised, frequency counts do not capture meaning, and positional encoding only records sequence position.
A startup is building an agent that must run locally on a mobile device and only needs to handle a narrow, specialised topic area. Which type of model is recommended for this?
- AA reasoning model such as DeepSeek R1
- BA frontier model like Claude Opus 4.5
- CA small language model (SLM)✓
- DA large language model (LLM)
Answer: Small language models tend to work well in scenarios that focus on specific topic areas or that require easily deployed small models for local applications and agents on devices. Large language models are powerful and generalise well but are more costly to train and use, and frontier and reasoning models such as Claude Opus 4.5 and DeepSeek R1 are hosted, high-capability models rather than lightweight on-device options.
In Foundry, what does the term model family mean?
- AA group of models deployed into the same Foundry project so that they can all be invoked through a single project endpoint
- BA group of related models that share an underlying architecture or lineage but differ in size, capability, specialisation or version✓
- CA group of models sold directly by Azure, hosted by Microsoft under Microsoft Product Terms with enterprise-grade service level agreements
- DA group of models that appear together in the catalog because they support the same inference task, such as chat or coding
Answer: A model family is a group of related models that share the same underlying architecture or lineage but differ in size, capability, specialisation or version, with GPT-5.x being one example. Models sold directly by Azure are a catalog category based on who hosts them, a shared inference task is a catalog filter rather than a family, and sharing a project endpoint is a deployment arrangement.
A company needs a model for a high-volume customer-support chat application that must respond quickly and at scale, and every developer in its Foundry environment must be able to use the model without registering. Which model should it choose?
- AClaude Opus 4.5
- BGPT-4.1✓
- CGPT-5.x
- DDeepSeek R1
Answer: GPT-4.1 is available to all Foundry users and is optimised for speed, efficiency and low-latency inference, which makes it ideal for real-time chat, customer support and interactive applications and better than reasoning-heavy models for high-volume production workloads. The GPT-5 model family currently requires registration, and DeepSeek R1 and Claude Opus 4.5 are positioned as reasoning and frontier models rather than low-latency chat options.
A developer wants a model family optimised for multi-step reasoning and orchestrating multi-tool agents, with adjustable thinking levels that let her trade speed for accuracy. Which family fits?
- APhi-4
- BGPT-5.x✓
- CGPT-4.1
- DMistral Large 3
Answer: The GPT-5.x family is optimised for multi-step reasoning, structured logic, planning and agentic workflows, and it supports adjustable thinking levels that let developers trade speed for accuracy when needed. GPT-4.1 is tuned for speed and low latency rather than deep reasoning, Mistral Large 3 is a general-purpose model that balances cost and throughput, and Phi-4 is a small language model.
Which deployment parameter in Foundry determines where and how inference is processed, with options such as standard, global batch and regional provisioned throughput?
- ATokens per minute (TPM)
- BModel version
- CRequests per minute (RPM)
- DDeployment type✓
Answer: Deployment type is the parameter that determines where and how inference is processed in Foundry, with standard, global batch and regional provisioned throughput as examples tied to throughput and data-processing requirements. Model version selects which build of the model is served, the TPM rate limit sets how many tokens the deployment may consume per minute, and RPM is a rate-limit boundary that follows from the TPM allocation.
A team deploys an image model in Foundry and is surprised that its quota is not expressed as tokens per minute. What is the most likely explanation?
- ASpecialised and image models often operate under capacity units instead of TPM✓
- BImage models are only available through global batch deployments, which have no quota
- CImage models are rate limited purely by requests per minute rather than by tokens
- DImage models are always sold by partners, so the partner's own quota system applies
Answer: Rate limits differ by model family: high-end reasoning models may have high TPM ceilings, while specialised or image models often operate under capacity units instead of TPM, which is why an image deployment does not show a tokens-per-minute figure. Global batch is one of several deployment types rather than an image-only route, partner hosting is unrelated to how quota is expressed, and RPM is a boundary derived from TPM rather than a substitute for it.
A deployed model starts returning rate-limit errors during peak traffic because its deployment-level quota is being exceeded. What is the recommended remedy?
- ALower the max tokens setting or reduce concurrent requests in code✓
- BAssign a lower tokens per minute allocation to the deployment
- CSwitch the deployment type from standard to global batch
- DRedeploy the same model using a newer model version from the model catalog
Answer: Larger prompts and higher max output token settings consume more TPM, so exceeding the deployment-level quota triggers throttling and rate-limit errors. When throttling appears, the recommended fix is to lower max tokens or reduce concurrent requests in code. Lowering the TPM allocation would reduce capacity further, a newer model version does not change the quota, and global batch is a deployment type for a different processing pattern rather than a remedy for peak-time throttling.
A retailer has a well-defined requirement to extract sentiment from product reviews. Rather than choosing a model from the Foundry model catalog, what alternative should it consider?
- AA small language model deployed locally on the retailer's own hardware so that no catalog model is required
- BA foundation model from the catalog, customised through fine-tuning on the retailer's own labelled review data
- CA Foundry tool powered by prebuilt models that provide predictable performance and built-in compliance✓
- DThe highest-ranked model on the Foundry quality leaderboard for the summarisation task
Answer: When a use case is well-defined, a Foundry tool can be chosen instead of a model from the catalog, because Foundry tools are powered by prebuilt models that provide predictable performance, built-in compliance and fast time-to-value without custom modelling. Fine-tuning a foundation model, deploying an SLM locally or picking a leaderboard winner all still involve selecting and operating a model yourself.