Text & Speech Solutions with Foundry Tools
18 free practice questions with explanations
18 free questions · instant explanations · no sign-up
PassNova has 18 free Microsoft AI-901 (Azure AI Fundamentals) practice questions on Text & Speech Solutions with Foundry Tools, each with a clear explanation. Practise them in the browser with instant feedback — 100% free, no sign-up, on any device. Updated for 2026.
Text & Speech Solutions with Foundry Tools: example questions & answers
18 worked examples with answers and explanations below. Practise them in the browser with instant feedback on every answer.
In the Foundry portal, which approach to text analysis returns structured, deterministic results for a specific task rather than responding to a natural language prompt?
- AA general language model deployment
- BAzure Language's purpose-built analyzers✓
- CA vision-enabled GPT model deployment
- DThe OpenAI Responses API in Foundry
Answer: Azure Language in Foundry Tools uses purpose-built analyzers that apply statistical techniques to return structured, deterministic output. That predictability suits automated pipelines where consistent results matter. A general-purpose language model deployment answers natural language prompts and its output varies with how the request is phrased, the Responses API is simply the interface used to send prompts to such a model, and a vision-enabled model is for image input rather than deterministic text analysis.
A developer asks a general-purpose model in the Foundry chat playground to identify the known entities in a review and return a link to Wikipedia for each one. Which text analysis task is this?
- AEntity linking✓
- BOpinion mining
- CPII detection
- DSummarization
Answer: Entity linking identifies known entities in text together with a link to Wikipedia. Opinion mining judges whether text is positive or negative, summarization condenses text to its most important information, and PII detection finds personal details such as names and phone numbers rather than linking entities to a reference source.
In the entity recognition example shown in the unit, the phrase '3 hours' is detected as which entity type and subtype?
- ADateTime, Time
- BDateTime, Duration✓
- CQuantity, Dimension
- DQuantity, Number
Answer: A duration such as '3 hours' is classified as the DateTime entity type with the Duration subtype. Quantity entities cover values such as '25%' (Percentage), '40' (Number) and '10 miles' (Dimension), while the Time subtype of DateTime is used for a clock time such as '8:00 AM'.
Which Azure Language capability identifies personal details such as names, phone numbers and addresses in text, covers personal health information (PHI), and can optionally redact what it finds?
- AOpinion mining
- BPII detection✓
- CSummarization
- DEntity linking
Answer: Personally identifiable information (PII) detection finds personally sensitive details in text, including personal health information, and can optionally redact them before the text is stored or shared. Entity linking connects known entities to Wikipedia, opinion mining judges positive or negative sentiment, and summarization condenses a passage, none of which redact sensitive data.
Which pip command installs the client library that the unit uses to call Azure Language features such as detect_language and recognize_pii_entities?
- Apip install azure-ai-projects
- Bpip install azure-ai-textanalytics✓
- Cpip install azure-cognitiveservices-speech
- Dpip install azure-ai-contentunderstanding
Answer: The Azure Language Python SDK is installed with pip install azure-ai-textanalytics, and it provides the TextAnalyticsClient whose methods include detect_language and recognize_pii_entities. The azure-cognitiveservices-speech package is the Azure Speech SDK, azure-ai-contentunderstanding is the Content Understanding SDK, and azure-ai-projects is the Foundry project client rather than the text analytics library.
In the Azure Language SDK sample, which expression prints the ISO 639-1 code of the language detected for a text string?
- Aresult.primary_language.iso6391_name✓
- Bresult.primary_language.language_code
- Cresult.primary_language.name
- Dresult.detected_language.iso_code
Answer: The detect_language() method returns a result whose primary_language object exposes the ISO 639-1 code through the iso6391_name attribute. The name attribute of primary_language gives the language name such as Spanish and confidence_score gives the 0 to 1 confidence, whereas language_code and detected_language.iso_code are not attributes shown in the sample.
A developer wants their Python app to return both a redacted copy of a customer message and a list of the personal details it contains, each with a category and confidence score. Which TextAnalyticsClient method should they call?
- Adetect_language()
- Bextract_key_phrases()
- Crecognize_entities()
- Drecognize_pii_entities()✓
Answer: The recognize_pii_entities() method identifies personal details in text and returns both the redacted version of the text and a list of the entities found, including each entity's category and confidence score. The detect_language() method returns the language, its ISO code and a confidence instead, and general entity recognition or key phrase extraction returns entities or phrases rather than a redacted copy of the text.
A team's Python app sends prompts to a gpt-4.1-mini deployment through the OpenAI library and gets slightly different wording on each run. They now need a consistent language code, confidence score and redacted text for every call. What should they change?
- AMove the endpoint and key into a .env file
- BDeploy gpt-4.1 in place of gpt-4.1-mini
- CRename the deployment to match the model
- DSwitch to the Azure Language SDK✓
Answer: A general-purpose model generates text probabilistically, so two calls with the same prompt can return slightly different wording or formatting. The Azure Language SDK returns structured, deterministic values such as a language code, confidence score and redacted text. Renaming the deployment, moving credentials into a .env file or switching to the larger gpt-4.1 model would not stop the output varying between calls.
How does the unit describe MCP (Model Context Protocol)?
- AA Foundry playground for testing deployed language models with sample prompts
- BA managed Azure service that hosts generative AI models on behalf of agents
- CAn open standard for connecting AI agents to external tools and data sources✓
- DA Python client library for calling Azure Language features from application code
Answer: The Model Context Protocol is an open standard that defines how AI agents connect to external tools and data sources, working like a universal adapter between an agent and an MCP server. It is a protocol rather than a hosting service, a client library or a playground; the Azure Language MCP server is a managed service that exposes Azure Language through that protocol, but MCP itself is the standard.
A company wants a Foundry agent to route support tickets by detected language and redact PII, without writing code that calls the Azure Language REST API or manages authentication tokens. What should they add to the agent?
- AA second Azure Language resource as a connection
- BThe Azure Language MCP server as a tool✓
- CA vision-enabled GPT model as a tool
- DThe TextAnalyticsClient as a hosted agent
Answer: The Azure Language MCP server is a managed service that exposes Azure Language capabilities such as language detection and PII detection through MCP, so an agent can call them with the same protocol it uses for any other MCP server and no direct REST calls or token handling are needed. A Foundry resource already includes access to Language tools, so no separate Azure Language resource is required; TextAnalyticsClient is an SDK class for client code, and a vision-enabled model handles image input rather than deterministic text analysis.
When you add the Azure Language MCP server as a tool in the Foundry playground, what do you configure the connection with?
- AYour model deployment name
- BYour Azure subscription ID
- CYour Foundry resource name✓
- DYour project API key
Answer: To connect to the Azure Language MCP server, you configure the connection with your Foundry resource name after searching the tools for Azure Language in Foundry Tools. A subscription ID, a model deployment name or an API key is not what the connection asks for, because the Foundry resource already provides unified access to the Language tools.
According to the unit, which two kinds of model does speech-to-text software typically use?
- AA multimodal model that converts audio into text, and a reasoning model that decides how to respond
- BA tokenization model that splits speech into words, and a translation model that maps words to a language
- CA neural voice model that converts text into phonemes, and a prosody model that assigns pitch and rate
- DAn acoustic model that converts audio into phonemes, and a language model that maps phonemes to words✓
Answer: Speech recognition typically combines an acoustic model, which converts audio into phonemes representing specific sounds, with a language model that maps those phonemes to words. Neural voices and prosodic units belong to speech synthesis, tokenization and translation are text processes, and a multimodal model with a reasoning step describes a speech-to-speech pipeline rather than the two models inside speech-to-text.
In the Azure Speech Python sample, which class is instantiated with use_default_microphone=True to supply audio to a SpeechRecognizer?
- Aspeechsdk.audio.AudioConfig✓
- Bspeechsdk.audio.AudioOutputConfig
- Cspeechsdk.AudioInputStream
- Dspeechsdk.SpeechConfig
Answer: The recognizer's microphone input is defined with speechsdk.audio.AudioConfig(use_default_microphone=True), which is then passed as audio_config to speechsdk.SpeechRecognizer. AudioOutputConfig with use_default_speaker=True is the speaker output used by a SpeechSynthesizer, SpeechConfig holds the subscription key and endpoint, and an audio input stream is not what the sample constructs.
A Python app connects handlers to a SpeechRecognizer so that partial text appears while the user is still speaking and final text appears once each phrase is complete. Which two events does it connect to?
- Alistening and transcribed
- Bsynthesizing and synthesized
- Cstarted and completed
- Drecognizing and recognized✓
Answer: The SpeechRecognizer exposes a recognizing event that fires with interim text while speech is still being processed and a recognized event that fires with the final text, each connected with .connect(handler) and read from evt.result.text. Synthesizing relates to text-to-speech, and started, completed, listening and transcribed are not the events the sample connects.
A legal firm has thousands of recorded interviews stored in Azure storage and wants them transcribed without anyone speaking live. Which Azure Speech option should it use, and how is the audio referenced?
- ASpeech synthesis, reading each file from a shared folder
- BVoice Live, uploading each file in the Foundry playground
- CReal-time transcription, streaming each file from a microphone
- DBatch transcription, pointing to the files with a SAS URI✓
Answer: Batch transcription is for audio recordings stored on a file share, a remote server or Azure storage; you point to the audio files with a shared access signature (SAS) URI and asynchronously receive the transcription results. Real-time transcription is for someone speaking live through a microphone or stream, Voice Live is for spoken conversations with an agent, and speech synthesis produces audio from text rather than transcribing it.
Which SpeechConfig property does the text-to-speech sample set to choose the voice, giving it the value 'en-US-Ava:DragonHDLatestNeural'?
- Aspeech_synthesis_output_format
- Bspeech_recognition_language
- Cspeech_synthesis_voice_name✓
- Dspeech_synthesis_language
Answer: The sample sets speech_config.speech_synthesis_voice_name='en-US-Ava:DragonHDLatestNeural' so that the neural multilingual voice is used, and that voice can speak different languages based on the input text. The recognition language, synthesis language and output format properties control other aspects of the SDK rather than which voice speaks.
Which statement describes how Azure Speech Voice Live differs from building a voice agent out of separate services?
- AIt runs the speech-to-text and text-to-speech stages locally on the client device
- BIt requires a separate Azure Speech resource for every stage of the pipeline
- CIt replaces the generative AI model with its own acoustic models for reasoning
- DIt combines speech-to-text with reasoning and text-to-speech in one managed service✓
Answer: Instead of building and connecting many separate pieces, the Voice Live API combines speech-to-text, AI reasoning and text-to-speech into one service that Azure fully manages, so you do not set up or maintain backend systems yourself. Voice Live still uses a generative AI model that you choose alongside its own acoustic models, runs in Azure rather than on the device, and needs no separate resource per stage.
Besides azure-ai-voicelive, which packages does the unit say you also need to install in order to run a Voice Live application in Python?
- Asounddevice, python-dotenv and azure-core
- Bpyaudio, openai and azure-ai-projects
- Cpyaudio, azure-cognitiveservices-speech and openai
- Dpyaudio, python-dotenv and azure-identity✓
Answer: A Voice Live application built with the azure-ai-voicelive package also needs pyaudio, python-dotenv and azure-identity installed. The openai and azure-ai-projects packages, sounddevice, azure-core and the azure-cognitiveservices-speech SDK are not the companion packages the unit lists for Voice Live.