Microsoft AI-901 (Azure AI Fundamentals)

Prompts, Chat Apps & Agents in Foundry

22 free practice questions with explanations

22 free questions · instant explanations · no sign-up

PassNova has 22 free Microsoft AI-901 (Azure AI Fundamentals) practice questions on Prompts, Chat Apps & Agents in Foundry, each with a clear explanation. Practise them in the browser with instant feedback — 100% free, no sign-up, on any device. Updated for 2026.

Sample questions

Prompts, Chat Apps & Agents in Foundry: example questions & answers

22 worked examples with answers and explanations below. Practise them in the browser with instant feedback on every answer.

  1. A chat application sends the text "You're a helpful assistant that responds in a cheerful, friendly manner." to the model before any question from the user. Which type of prompt is this?

    • AAn agent tool call
    • BA model completion
    • CA user prompt
    • DA system prompt✓

    Answer: A system prompt sets the behaviour and tone of the model and any constraints it should adhere to, and it is usually set by the application that uses the model. A user prompt elicits a response to a specific question or instruction, a completion is the model's response to a prompt, and a tool call is an action an agent takes rather than an instruction given to the model.

  2. A developer keeps getting unstructured, rambling answers from a model. Which of the unit's tips for better prompts most directly addresses the format of the response?

    • AAvoid giving examples to the model
    • BAsk for structure such as bullet points✓
    • CAlways set the temperature to 1
    • DKeep every prompt under ten words

    Answer: Asking for structure, such as bullet points, tables or numbered lists, is one of the recommended tips for better prompts and directly shapes the format of the response. The other tips are to be clear and specific, add context and use examples, so avoiding examples runs against the guidance. Keeping prompts artificially short or fixing the temperature at 1 are not among the tips and do not control response format.

  3. In the Foundry playground settings, which parameter controls creativity versus determinism in the model's responses?

    • ATemperature✓
    • BMax output tokens
    • CDeployment type
    • DModel version

    Answer: Temperature controls creativity versus determinism in the model's output. Max output tokens caps the response length and affects token consumption and throttling, model version selects which release of the model is deployed, and deployment type determines where and how inference is processed.

  4. In the Python chat client shown in the unit, which class is instantiated with base_url and api_key to call the deployed model through the Responses API?

    • AAzureOpenAI from the openai package
    • BAIProjectClient from azure.ai.projects
    • CDefaultAzureCredential from azure.identity
    • DOpenAI from the openai package✓

    Answer: The Foundry lightweight chat client creates an OpenAI client with base_url set to the endpoint plus /openai and api_key read from the environment, then calls the Responses API on it. AIProjectClient is the Foundry project client used with DefaultAzureCredential in the agent sample, not the class that takes base_url and api_key here. AzureOpenAI is a different class in the openai package that the unit's sample does not use.

  5. In the lightweight Python chat client, which method is called to send the system and user messages to the deployed model?

    • Aclient.agents.get()
    • Bclient.get_openai_client()
    • Cclient.responses.create()✓
    • Dclient.chat.completions.create()

    Answer: The Foundry lightweight chat client calls client.responses.create() with the deployment name, an input list of messages, max_output_tokens and temperature, then prints response.output_text. The chat.completions method belongs to the older chat completions interface rather than the Responses API used in the sample. The agents.get() and get_openai_client() calls are methods of the project client in the agent sample, not the way messages are sent to a model.

  6. In the chat client's call to client.responses.create(), which two keyword arguments carry the values the developer captured in the playground for response length and creativity?

    • Amax_output_tokens and temperature✓
    • Bmax_completion_tokens and temperature
    • Cmax_tokens and top_p
    • Dmax_output_tokens and top_p

    Answer: The Foundry chat-client sample passes max_output_tokens=300 to cap the response length and temperature=0.7 to set creativity, mirroring the playground's max output tokens and temperature settings. The names max_tokens and max_completion_tokens are parameters of other OpenAI interfaces and are not used in the Responses API call shown. The top_p parameter is not part of the sample at all.

  7. After calling the Responses API in the chat client, which attribute of the response object does the sample print to display the model's reply?

    • Aresponse.text
    • Bresponse.output
    • Cresponse.output_text✓
    • Dresponse.choices[0].message

    Answer: The Foundry chat-client sample ends with print(response.output_text), which prints the model's generated text. The choices list is the shape of chat completion responses rather than Responses API responses. The output attribute holds the structured list of output items rather than the plain reply, and response.text is not the attribute the unit prints.

  8. According to the unit on using a generative AI model, in which situation would you use a generative AI model on its own rather than an agent?

    • AWhen a goal must be broken into structured steps
    • BWhen the app must call external tools automatically
    • CWhen you want pure inference from a single prompt✓
    • DWhen working memory must persist across a conversation

    Answer: A model on its own is used for pure inference, taking a prompt and generating output, for experimenting in the Playground, or for calling the model through the OpenAI Responses API. Calling external tools automatically, breaking goals into structured steps and maintaining working memory during a conversation are all capabilities the unit attributes to agents, which are packaged, task-oriented workers built on top of a model.

  9. An agent in Microsoft Foundry is described as a packaged, reusable AI component that brings together which three things?

    • AA tenant, a subscription and a region
    • BA knowledge base, a source and retrieval
    • CA model, a deployment and an endpoint
    • DA model, instructions and tools✓

    Answer: A Foundry agent brings together a model for reasoning, instructions in the form of a system prompt that defines its role and behaviour, and tools that define the actions it can take. Tenants, subscriptions and regions are Azure organisational and configuration concepts. A deployment and an endpoint are how a model is served, and a knowledge base, knowledge source and agentic retrieval are the components of a Foundry IQ solution.

  10. In the Foundry portal, a developer has chosen a model and written system instructions for a scheduling assistant. What does the unit say sets an agent apart from using the model alone at this point?

    • ARaising the TPM allocation
    • BChanging the deployment type
    • CAdding tools and knowledge✓
    • DLowering the temperature

    Answer: Adding tools, which let the model act on information, and knowledge, which grounds the model with information, is what sets an agent apart from using the model alone; the unit summarises this as tools equal actions and knowledge equals context. TPM allocation, temperature and deployment type are model and deployment settings that apply whether or not an agent is involved.

  11. A developer wants to call a Foundry agent programmatically through the Project API and needs its agent-id. Where does the unit say this can be found?

    • AIn the Azure AI Search index, in the metadata of the knowledge base
    • BIn the Foundry Tool Catalog, in the entry that lists the agent's tools
    • CIn the Azure portal, in the All resources pane for the Foundry resource
    • DIn the Playground view of the agent, in the code view's .env variables✓

    Answer: The agent-id is found in the Playground view of the agent when you select the code view and open the .env variables. The All resources pane in the Azure portal lists Azure resources rather than agent identifiers, the Foundry Tool Catalog is where tools are discovered and managed, and an Azure AI Search index holds knowledge base content rather than agent IDs.

  12. In the agent client sample, how does the call to openai_client.responses.create() tell Foundry which agent should handle the request?

    • ABy passing the agent's instructions as the system message
    • BBy passing an extra_body argument containing an agent_reference✓
    • CBy passing the agent's name as the model argument
    • DBy passing the agent-id as the api_key argument

    Answer: The sample passes an extra_body argument that references the agent by name with type agent_reference, so that the Responses API routes the input to the named agent. The unit prints this as extra_body={"agent": {"name": agent.name, "type": "agent_reference"}}, and azure-ai-projects 2.0.0b4 and later name the outer key agent_reference instead of agent. The model argument names a deployment rather than an agent, and the agent's instructions are stored with the agent, so they are not resent as a system message. The api_key argument is for key-based authentication, and this sample authenticates with DefaultAzureCredential instead.

  13. In the agent sample, after creating project_client with AIProjectClient, which call obtains the OpenAI-compatible client used to send the request to the agent?

    • AOpenAI(base_url=myEndpoint)
    • Bproject_client.responses.create()
    • Cproject_client.agents.get()
    • Dproject_client.get_openai_client()✓

    Answer: The Foundry agent sample calls openai_client = project_client.get_openai_client() and then uses openai_client.responses.create() to get a response from the agent. The agents.get() call retrieves an existing agent by name rather than returning a client. Constructing OpenAI with a base_url is the pattern of the standalone chat client, and responses.create() is a method of the OpenAI client rather than of the project client.

  14. In a Foundry IQ solution, which component is described as the top-level resource that identifies a collection of related knowledge sources and controls retrieval behaviour?

    • AThe knowledge source
    • BThe Azure AI Search index
    • CThe knowledge base✓
    • DAgentic retrieval

    Answer: The knowledge base is the top-level Foundry IQ resource that groups related knowledge sources and controls how retrieval behaves. A knowledge source is a connection to indexed or remote content such as SharePoint or Azure Storage, and agentic retrieval is the process that plans searches, ranks results and returns a unified response with references. An Azure AI Search index is one kind of content a knowledge source can point to rather than the top-level resource.

  15. When an agent queries a Foundry IQ knowledge base, what does agentic retrieval do?

    • AIt rewrites the question, caches it and returns the previous answer
    • BIt sends the question to one source, then returns the raw document
    • CIt plans subqueries, searches sources in parallel and returns cited results✓
    • DIt retrains the model on the question, then answers from its parameters

    Answer: Agentic retrieval breaks the question into subqueries, searches multiple sources in parallel, ranks the results and returns relevant, citation-backed information while enforcing user permissions. It does not retrain or alter the model, it searches across the configured sources rather than a single one, and it returns ranked content with source references rather than a cached earlier answer.

  16. A team has created a Foundry IQ knowledge base and wants an agent in Foundry Agent Service to use it. How is the knowledge base connected to the agent?

    • AThrough a system prompt listing document URLs
    • BThrough a Key Vault secret for the search service
    • CThrough a Code Interpreter tool on the agent
    • DThrough an MCP tool added to the agent✓

    Answer: Foundry IQ knowledge bases expose a Model Context Protocol (MCP) tool, and the workflow is to add the knowledge base connection to the agent as an MCP tool and then write instructions that say when to use it. Code Interpreter is a tool for data analysis and file handling, not the knowledge base connector. A Key Vault secret stores credentials rather than connecting a knowledge base, and listing document URLs in a system prompt would bypass Foundry IQ's retrieval entirely.

  17. After creating a Foundry IQ resource, what must be done so that knowledge retrieval can be tested in the agent playground?

    • AEnable public web content as a knowledge source in each knowledge base
    • BAdd the project's managed identity to the Search Index Data Reader role✓
    • CDeploy a second chat model in the project and set its TPM to the maximum
    • DStore the Azure AI Search admin key in the agent's system instructions

    Answer: The managed identity used by the Foundry project must be added to the Search Index Data Reader role for the Azure AI Search resource, and the managed identities of production agents need the same role. Putting a search admin key into system instructions exposes a secret and is not how access is granted. Extra model deployments and a higher TPM do not grant the playground identity access to the search resource, and public web content is an optional source rather than a prerequisite for playground testing.

  18. What does the RAG unit say happens to the language model's learned parameters when retrieval-augmented generation is used?

    • AThey are replaced by the weights of the embedding model
    • BThey are fine-tuned on every retrieved document chunk
    • CThey stay unchanged; knowledge is added at run time✓
    • DThey are re-trained whenever the search index is updated

    Answer: RAG does not change the language model's learned parameters; it supplies relevant knowledge at run time when the user makes a request. Because of this, when policies change you update the documentation and the index rather than retraining or fine-tuning the model. The embedding model is a separate model used to create vectors for chunks and queries and never replaces the language model's weights.

  19. A RAG team notices that some retrieved chunks contain a statement but not the heading or definition needed to understand it. According to the unit, what has probably gone wrong with chunking?

    • AThe chunks overlap too much
    • BThe embeddings are too long
    • CThe chunks are too large
    • DThe chunks are too small✓

    Answer: Chunks that are too small can separate a statement from headings, definitions or other information needed to understand it, which is why chunks often overlap slightly. Chunks that are too large cause the opposite problem, pulling in irrelevant information and consuming more of the context window. Overlap exists to preserve context at boundaries, and embedding vectors have a fixed size set by the embedding model rather than being too long.

  20. A support team's RAG index must reliably match exact product codes and identifiers in customer questions. Which search type does the unit say is effective when exact words matter?

    • AMetadata filtering
    • BAgentic retrieval
    • CVector search
    • DKeyword search✓

    Answer: Keyword search is effective when exact words, identifiers or product codes matter. Vector search finds semantically similar content even when the wording differs, so it is not built around exact matches. Agentic retrieval is Foundry IQ's process for planning and coordinating searches, and metadata filtering narrows results by attributes such as region or date rather than matching terms in the text.

  21. Why does the RAG unit say that retrieved documents should be treated as data rather than as instructions when the prompt is augmented?

    • ACitations are generated only for data and never for instructions
    • BSource content could contain text that tries to redirect the model✓
    • CRetrieved chunks are always placed after the system prompt
    • DThe model cannot read instructions inside a JSON request body

    Answer: The prompt must distinguish trusted application instructions from retrieved content because source documents could contain text that attempts to redirect the model. Models can read text wherever it appears in a request body, so that is not the reason. Citations are about letting users verify sources rather than a rule about instructions, and the position of chunks in the prompt does not by itself stop injected text from being followed.

  22. When evaluating a RAG solution, which of these questions belongs to evaluating generation rather than retrieval?

    • ARelevance: do the retrieved chunks address the question?
    • BCoverage: is all the information needed for a complete answer present?
    • CGroundedness: is each claim supported by the retrieved content?✓
    • DRanking: are the strongest results placed ahead of weaker ones?

    Answer: Groundedness, which asks whether each claim is supported by the retrieved content, is a generation consideration alongside relevance of the response, completeness, citation quality and appropriate uncertainty. Coverage, ranking and the relevance of the retrieved chunks are retrieval considerations that ask whether the search system returned useful source content. If the right information is not retrieved, changing the generation prompt alone will not fix the answer.

Start practising Prompts, Chat Apps & Agents in Foundry →